A VS Code extension that turns Chronicle/compact data into a live, visual token-optimization dashboard for GitHub Copilot Chat. It estimates input/output tokens offline (no cloud sync required) and nudges you to act before context bloats.
What the MVP does
Context meter - shows the latest saved promptTokens reported by Copilot for the selected thread, separately from input capacity. Missing or model-mismatched telemetry is shown as unavailable, not a transcript estimate or zero.
Always-on overhead auditor — tokenizes .github/copilot-instructions.md, CLAUDE.md/AGENTS.md, and .vscode/mcp.json so you can see what every turn costs before you type.
Top offenders — flags oversized single turns (e.g. giant terminal/paste dumps) and files read repeatedly within a session.
Recent sessions — estimated token totals per session.
One-click actions — Compact (/compact), Fresh chat, Open mcp.json, Settings.
Rewrite prompt - draft and rewrite a prompt directly in the dashboard, edit the result in the same textbox, then select Send to Chat. Drafts survive dashboard refreshes. Rewriting uses the existing Copilot model logic to preserve intent, clarify targets, and add stopping conditions; it makes a model call only when requested.
Transcript, trend, overhead, and projected-input counts are estimates via gpt-tokenizer (cl100k). The context meter instead uses the latest reported input tokens, including context counted by Copilot. It is not a cumulative billing total or an output-token count, and it can decrease after compaction. Actual billed input/output/cache accounting is not implemented.
How it works
Reads the Copilot chat session store (GitHub.copilot-chat/**/*.db) read-only via VS Code's built-in SQLite, including committed write-ahead log (WAL) updates. Requires VS Code 1.137 or newer.
Auto-detects the most recently modified session DB; override with tokenOptimizer.sessionDbPath.
A directory watcher detects database and WAL changes; a 60-second fallback also refreshes the dashboard. Concurrent refresh requests are coalesced.
Refresh at the top of the dashboard reads the latest local indexed data. It does not trigger cloud sync or force Copilot to index an unfinished turn. The button shows progress and remains available for retries after errors.
Refreshed at is the last successful dashboard read. Session updated is the stored session timestamp. Latest indexed chat is the most recently updated indexed session across workspaces, not necessarily the visible chat; select a thread to pin it.
The meter follows the latest request's matching model telemetry and can lag an unfinished request. It does not reuse older usage when the newest request has no reading. Model capacity comes from saved model metadata, with a labeled configured fallback.
Compaction advice uses a configurable 150,000-token heuristic, capped at 80% of input capacity. It is not Copilot's native compaction trigger or a validated optimum. Manual chat actions target the active Copilot chat, which may differ from the displayed thread.
Run it
cd local-packages/token-optimizer
npm install
npm run compile
Then press F5 in VS Code (or "Run Extension") to launch an Extension Development Host. Open the Token Optimizer view in the Activity Bar.
Settings
Setting
Default
Purpose
tokenOptimizer.sessionDbPath
auto
Override the session DB path
tokenOptimizer.modelContextWindow
128000
Fallback input capacity when model metadata is unavailable
tokenOptimizer.compactAdvisoryTokens
150000
Reported-input advisory, capped at 80% of capacity; 0 disables