Skip to content
| Marketplace
Sign in
Visual Studio Code>Machine Learning>TokenLens Live CoachNew to Visual Studio Code? Get it now.
TokenLens Live Coach

TokenLens Live Coach

Aditya Singh

|
1 install
| (0) | Free
Live GitHub Copilot Chat token, model, and Premium Request credit usage in VS Code
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

TokenLens Live Coach

TokenLens Live Coach shows a live estimated dollar cost, token usage, and real model for GitHub Copilot Chat directly inside VS Code — always visible in the status bar and broken down per session — automatically, with no SDK, no code changes, no account/API setup, and no configuration needed. It also surfaces evidence-based token-optimization suggestions, a reactive budget alert, and static diagnostics on your source files to flag token-wasteful AI request patterns as you type.

How it works — and an important trade-off to know about

VS Code extensions can't observe another extension's (e.g. Copilot Chat's) network requests directly — that's blocked by design for security. We first tried Copilot Chat's official OpenTelemetry export mechanism, but as of Copilot Chat 0.67.0 it only reports a small internal "utility" call (used for things like chat-title generation) — never your real per-message model, tokens, or cost. That made it useless for this extension's purpose.

Instead, this extension reads VS Code's own local Copilot Chat session transcripts — the same files that power the Chat panel's history — stored per-workspace at:

<VS Code user data dir>/User/workspaceStorage/<hash>/chatSessions/<sessionId>.jsonl

These files contain the real model used per turn (e.g. copilot/claude-opus-4.8), real promptTokens/completionTokens, and GitHub's own AI credit cost for that turn (copilotCredits) — the currency GitHub's usage-based Copilot billing is denominated in (see "Estimated cost" below). This extension only ever reads that small set of numeric/metadata fields; it never reads or stores the actual prompt/response content fields also present in those files.

Trade-off to be upfront about: this file format is an undocumented internal VS Code implementation detail, not a public API. It has already changed between versions we observed (field names/shape) and could change again in a future VS Code release without notice, which could break this extension until it's updated. This is a deliberate trade-off in exchange for real data, made after confirming Copilot Chat's officially documented telemetry mechanism doesn't currently expose what we need.

A second, newer data source: agent-mode sessions

Some Copilot Chat conversations — particularly "Agent mode" turns — are instead handled by a newer agent-host engine that persists its own session logs at:

<home dir>/.copilot/session-state/<uuid>/events.jsonl

alongside a small workspace.yaml recording which folder (cwd) the session belongs to. This extension reads that folder's session.usage_checkpoint events for the same kind of data — real model, real prompt tokens, GitHub's Premium Request quota usage (totalPremiumRequests), and the session's cumulative AI credit cost (totalNanoAiu, see "Estimated cost" below) — while explicitly skipping the session's name/title field, which can contain your literal first message and is therefore treated as content, not metadata.

This is an even less documented, still-evolving mechanism than the Chat panel's own storage above — treat it as at least as likely to change. Because it's a general-purpose local agent engine (not exclusive to VS Code's Chat panel), any tool built on it — including this very AI assistant, if you're using one right now — will show up as a session here too. That's intentional: it represents real usage against your account either way, not a bug.

Estimated cost

The status bar (always visible at the bottom while a workspace is open) and every session in the Recent sessions sidebar list show an estimated USD cost — not a guess from a hardcoded pricing table, but computed from real local numbers using GitHub's own officially documented conversion:

  • The Copilot SDK docs give the exact formula for a session's cumulative totalNanoAiu: aiCredits = totalNanoAiu / 1e9.
  • The Copilot pricing docs state the AI credit's dollar value: 1 AI credit = $0.01 USD.
  • The Chat panel's own copilotCredits field is already expressed in this same AI-credit unit (verified against real data: single-turn values like 25.93 only make sense as ~$0.26 for one agentic turn, not as 26 whole Premium Requests).

So both data sources share one conversion, applied in src/pricing.ts. This is still an estimate, not your actual invoice — it has no way to know whether this usage already falls inside your plan's included monthly AI credit allowance, in which case you may not be billed anything extra for it. Treat the number as a real, locally-computed signal of relative usage, not a guaranteed charge.

No setup required

Unlike the OpenTelemetry approach, there is no setting to enable, no prompt to accept, and no integrated-terminal environment variable to inject — both mechanisms above are already written by default. Install the extension, use Copilot Chat (in any mode), and data appears.

Multiple windows and workspaces

Both data sources are inherently scoped per-workspace — the Chat panel's chatSessions folder is per-workspace-hash, and agent-host sessions are matched by comparing each session's recorded cwd against your open workspace folder. So this extension's data is automatically scoped to the workspace open in the current window — no cross-window bleed, no repository-matching heuristics needed. If you use Copilot across many projects, each window's TokenLens view only shows that project's own usage.

Copilot CLI

We investigated whether a similar local file exists for the standalone GitHub Copilot CLI tool (run outside VS Code, as its own terminal application) and did not find one accessible in the same way; that specific standalone-CLI usage is not currently tracked by this extension. (Note this is distinct from the "agent-host" mechanism above, which does track certain in-VS-Code agent-mode conversations.)

Privacy

  • The only thing ever read from disk is a whitelisted set of metadata fields per turn: model ID, token counts, AI-credit/Premium-Request cost, and timing. Prompt/response content fields (including conversation titles, which can contain your literal message text) are never parsed or stored from either data source.
  • Nothing is sent anywhere — there's no local server, no network call, and no telemetry from this extension at all.
  • No TokenLens cloud service, and no provider/GitHub API key, is involved anywhere.

Token-optimization suggestions

The Suggestions group in the TokenLens sidebar surfaces up to a handful of deterministic, evidence-based recommendations computed directly from your own real local usage data — no guessed pricing, no external catalog, no reading of prompt/response content. Each one cites the exact numbers that triggered it. Current rules:

  • Tool/MCP schema overhead (agent-host and Chat panel sessions): flags when tool/MCP definitions resent on every turn make up ≥30% of the prompt, and suggests disabling unused tool-providing extensions/MCP servers for that workspace.
  • Poor prompt-cache reuse: flags multi-turn agent-host sessions where more tokens were freshly written to cache than reused from it, and suggests avoiding mid-conversation tool/model switches that invalidate the cached prefix.
  • Cost-per-turn outlier: flags a session whose real cost-per-turn is a clear outlier (≥2x) against the median of your own recent sessions in this workspace — not against any external pricing table — and suggests trimming context or splitting long tasks.

These recompute on every poll, so they reflect your current session history as it grows. For a plain-language walkthrough of exactly how each rule works and why (no LLM involved — it's a small, deterministic, local rule engine), see SUGGESTIONS.md in this extension's repository.

Each suggestion's full evidence and suggested action can be hard to read from a hover tooltip alone (tooltips disappear if your mouse moves off the small tree row). To make sure the full text is always reachable: click a suggestion to open it in a modal dialog that stays open until you dismiss it (with a "Copy to clipboard" option), or click the expand arrow next to it to reveal the evidence and suggested action as their own rows directly in the tree.

Budget alert (not a circuit breaker)

We originally set out to build a circuit breaker — something that could stop an over-budget Copilot request before it goes out. After investigating, that isn't possible: VS Code extensions have no API to intercept or block another extension's (Copilot Chat's) outgoing network requests. What this extension does instead is a reactive budget alert: once today's usage in this workspace crosses a threshold you configure, it shows a one-time-per-day warning notification (and the status bar turns amber) — after the fact, not before. It's a genuinely useful early-warning signal, just not a hard stop, and we'd rather be precise about that than overclaim.

Configure thresholds with TokenLens: Configure Budget Threshold, or directly via settings — see tokenlens.budget.warnAtCredits, tokenlens.budget.warnAtPremiumRequests, and tokenlens.budget.warnAtTokens below. Changing a threshold immediately re-arms the alert so you don't have to wait until tomorrow to test it.

Getting started

  1. Install the extension and reload VS Code.
  2. Use Copilot Chat as you normally would, in a workspace with a folder open — Ask mode, Agent mode, or any agent-host-based tool all work.
  3. Open the TokenLens activity-bar view, or check the status bar, to see the estimated cost, tokens, and credit/Premium Request data per session and for today.

Commands

  • TokenLens: Open Live Coach — opens the sidebar view.
  • TokenLens: Configure Budget Threshold — set a credit or token warning threshold for today's usage in this workspace.
  • TokenLens: Start/Stop Local Session — pause/resume the periodic re-check of session files.

Advanced configuration

  • tokenlens.budget.warnAtUsd / tokenlens.budget.warnAtCredits / tokenlens.budget.warnAtPremiumRequests / tokenlens.budget.warnAtTokens: status-bar warning + once-per-day notification thresholds for today's usage in this workspace (estimated USD, AI credits, Premium Requests, and raw token count respectively).
  • tokenlens.pollIntervalMs: how often to re-check Copilot's session files for updates (default 2000ms).
  • tokenlens.diagnostics.enabled / tokenlens.diagnostics.maxOutputTokensThreshold: static source diagnostics for AI request code (works independently of Copilot usage tracking).
  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft