Skip to content
| Marketplace
Sign in
Visual Studio Code>Other>DietCode Token OptimizerNew to Visual Studio Code? Get it now.
DietCode Token Optimizer

DietCode Token Optimizer

Abiha Naqvi

|
2 installs
| (0) | Free
Live token-usage dashboard and optimization nudges for GitHub Copilot Chat (input / output / cache).
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

Copilot Token Optimizer (MVP)

A VS Code extension that turns Chronicle/compact data into a live, visual token-optimization dashboard for GitHub Copilot Chat. It estimates input/output tokens offline (no cloud sync required) and nudges you to act before context bloats.

What the MVP does

  • Context meter - shows the latest saved promptTokens reported by Copilot for the selected thread, separately from input capacity. Missing or model-mismatched telemetry is shown as unavailable, not a transcript estimate or zero.
  • Always-on overhead auditor — tokenizes .github/copilot-instructions.md, CLAUDE.md/AGENTS.md, and .vscode/mcp.json so you can see what every turn costs before you type.
  • Top offenders — flags oversized single turns (e.g. giant terminal/paste dumps) and files read repeatedly within a session.
  • Recent sessions — estimated token totals per session.
  • One-click actions — Compact (/compact), Fresh chat, Open mcp.json, Settings.
  • Rewrite prompt - draft and rewrite a prompt directly in the dashboard, edit the result in the same textbox, then select Send to Chat. Drafts survive dashboard refreshes. Rewriting uses the existing Copilot model logic to preserve intent, clarify targets, and add stopping conditions; it makes a model call only when requested.

Transcript, trend, overhead, and projected-input counts are estimates via gpt-tokenizer (cl100k). The context meter instead uses the latest reported input tokens, including context counted by Copilot. It is not a cumulative billing total or an output-token count, and it can decrease after compaction. Actual billed input/output/cache accounting is not implemented.

How it works

  • Reads the Copilot chat session store (GitHub.copilot-chat/**/*.db) read-only via VS Code's built-in SQLite, including committed write-ahead log (WAL) updates. Requires VS Code 1.137 or newer.
  • Auto-detects the most recently modified session DB; override with tokenOptimizer.sessionDbPath.
  • A directory watcher detects database and WAL changes; a 60-second fallback also refreshes the dashboard. Concurrent refresh requests are coalesced.
  • Refresh at the top of the dashboard reads the latest local indexed data. It does not trigger cloud sync or force Copilot to index an unfinished turn. The button shows progress and remains available for retries after errors.
  • Refreshed at is the last successful dashboard read. Session updated is the stored session timestamp. Latest indexed chat is the most recently updated indexed session across workspaces, not necessarily the visible chat; select a thread to pin it.
  • The meter follows the latest request's matching model telemetry and can lag an unfinished request. It does not reuse older usage when the newest request has no reading. Model capacity comes from saved model metadata, with a labeled configured fallback.
  • Compaction advice uses a configurable 150,000-token heuristic, capped at 80% of input capacity. It is not Copilot's native compaction trigger or a validated optimum. Manual chat actions target the active Copilot chat, which may differ from the displayed thread.

Run it

cd local-packages/token-optimizer
npm install
npm run compile

Then press F5 in VS Code (or "Run Extension") to launch an Extension Development Host. Open the Token Optimizer view in the Activity Bar.

Settings

Setting Default Purpose
tokenOptimizer.sessionDbPath auto Override the session DB path
tokenOptimizer.modelContextWindow 128000 Fallback input capacity when model metadata is unavailable
tokenOptimizer.compactAdvisoryTokens 150000 Reported-input advisory, capped at 80% of capacity; 0 disables
tokenOptimizer.oversizedTurnTokens 8000 Threshold to flag a turn as oversized

Roadmap

  • Phase 2 — real-time active-chat meter, cache-break detection (model switch / TTL / MCP change), inline warnings.
  • Phase 3 — cloud DuckDB mode for actual billed input/output/cache tokens; MCP scoping advisor; trend reports wired into .github/hooks telemetry.
  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft