LiteLLM Provider for GitHub Copilot Chat
Use 100+ LLMs in VS Code with GitHub Copilot Chat powered by LiteLLM.
Features
- 100+ LLMs through a unified API (OpenAI, Anthropic, Google, AWS, Azure, Ollama, and any OpenAI-compatible endpoint)
- Multi-server support: connect to multiple LiteLLM servers simultaneously and aggregate models
- Streaming chat completions with SSE
- Function calling (tool use) support
- Multimodal input (vision/image attachments)
- Reasoning/thinking tokens (if the model supports them) — including turning thinking OFF entirely for reasoning models (faster responses, no reasoning tokens burned)
- Dashboard panel for servers, models, and settings
- API key encryption via VS Code SecretStorage
- Per-model pricing display in the model picker
- Status bar context-window indicator — a live fill gauge showing how much of the current model's context window is in use, fed by real backend
usage (when reported) or token estimates. VS Code's own "Session Info" panel stays at 0 for custom providers (no host API to feed it), so this is the extension's own equivalent.
Requirements
- VS Code 1.99.0 or higher, with the GitHub Copilot Chat extension installed and signed in
- LiteLLM proxy running (self-hosted or cloud)
- LiteLLM API key (if required by your setup)
Quick Start
- Install the extension
- Open VS Code's chat interface (
Ctrl+Alt+I / Cmd+Ctrl+I, or the chat icon in the title bar)
- Run the command "LiteLLM: Add Server" from the Command Palette
- Enter a label (e.g.
Local), base URL (e.g. http://localhost:4000), and API key
- Back in chat, pick one of the new models in the model picker and send a message
You can also declare the server in settings.json:
"litellm-vscode-chat.servers": [
{ "label": "Local", "baseUrl": "http://localhost:4000", "apiKey": "sk-..." }
]
The API key is automatically migrated to VS Code's encrypted secret storage on first activation.
Commands
| Command |
Description |
LiteLLM: Add Server |
Add a new LiteLLM server |
LiteLLM: Remove Server |
Remove a configured server |
LiteLLM: Manage Servers |
Quick pick for server management |
LiteLLM: Open Dashboard |
Open the dashboard webview |
LiteLLM: Test Connection |
Test a server connection and discover models |
LiteLLM: Sync Models |
Trigger a model refresh |
LiteLLM: Show Output Log |
Open the extension's output channel |
LiteLLM: Set Reasoning Effort |
Pick a reasoning model and set effort Low/Medium/High/Max — or turn thinking Off entirely (persisted per-model in models.parameters) |
Settings
| Setting |
Default |
Description |
litellm-vscode-chat.servers |
[] |
LiteLLM server entries |
litellm-vscode-chat.discovery.timeout |
30000 |
Timeout (ms) for model discovery |
litellm-vscode-chat.chat.timeout |
300000 |
Timeout (ms) for chat requests |
litellm-vscode-chat.chat.tokenEstimation |
char4 |
Token estimation mode |
litellm-vscode-chat.chat.promptCaching |
false |
Send prompt-cache breakpoints |
litellm-vscode-chat.models.maxOutputTokens |
4096 |
Default max output tokens |
litellm-vscode-chat.models.parameters |
{} |
Per-model request parameter overrides, keyed by glob pattern (e.g. { "zai-org/GLM-5.2-FP8*": { "max_tokens": 50000 } }) — always wins over the request's own options. Disable thinking for a model with "chat_template_kwargs": { "enable_thinking": false } (Qwen3/GLM on vLLM & SGLang) or { "thinking": false } (DeepSeek) |
Architecture
┌─────────────────────────────────────────────────────────────┐
│ VS Code Chat Host │
│ (calls LanguageModelChatProvider interface) │
└──────────────────────────┬──────────────────────────────────┘
│
┌──────────▼──────────┐
│ LiteLLMChatModel │
│ Provider │
│ (src/provider/) │
└──────────┬──────────┘
│
┌───────────────┼───────────────┐
│ │ │
┌──────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
│ Discovery │ │ ChatClient │ │ TokenCount │
│ (/models) │ │ (/chat) │ │ (estim.) │
└──────┬──────┘ └──────┬──────┘ └─────────────┘
│ │
┌──────▼───────────────▼──────┐
│ LiteLLM Proxy Server │
│ (OpenAI-compatible API) │
└────────────────────────────┘
Installation
From VSIX file
Option 1 — VS Code GUI:
- Open VS Code
- Open the Extensions panel (
Ctrl+Shift+X)
- Click the
⋯ menu (top-right) → Install from VSIX...
- Select the
litellm-vscode-chat-0.1.0.vsix file
Option 2 — Command line:
code --install-extension litellm-vscode-chat-0.1.0.vsix
From source (development)
npm install
npm run compile
Press F5 to launch the Extension Development Host.
Packaging
To build a .vsix package from source:
npm run compile
npx vsce package --no-dependencies
This produces litellm-vscode-chat-0.1.0.vsix in the project root.
To rebuild after code changes, re-run the same commands — the .vsix is regenerated with the updated code.
Privacy
Your prompts and completions travel only between VS Code and the LiteLLM servers you configure. No telemetry, no third-party calls.
The extension's output log never contains conversation content: request diagnostics use a content-free fingerprint (message counts, role sequence and a one-way hash of user-turn lengths), and server error messages are reduced to a short stable classification (e.g. rate_limit, authentication) instead of being logged verbatim.
Local usage telemetry (optional, off by default)
Setting litellm-vscode-chat.telemetry.localUsage.enabled (machine scope, default false) records one JSONL event per physical HTTP attempt the provider makes — useful for auditing token spend, retry storms and gateway behaviour. It is purely local: nothing is uploaded, and when the setting is off no file or directory is ever created.
Where: <extension globalStorage>/usage-telemetry/v1/YYYY-MM-DD.jsonl (one file per UTC day; files older than 30 days are removed at startup; schema is append-only and versioned via schemaVersion).
What is recorded (metadata only):
- request identity:
logicalRequestId (per Copilot request), attemptId + httpAttempt (per physical fetch), retry dimensions (truncationAttempt, networkRetry, cause);
- model: requested model, response model, LiteLLM model id/group;
- response:
x-litellm-call-id, response id, HTTP status, outcome (completed / http_error / network_error / timeout / cancelled / stream_error / incomplete_stream / length_limit), finish reason, stable error type;
- timing: start timestamp, duration, time-to-first-byte, time-to-first-token, server-reported durations;
- usage: prompt/completion/total tokens plus
cached_tokens, cache_write_tokens, reasoning_tokens, prediction-audio tokens when the backend reports them (missing values stay null, never 0; reasoning tokens are a subset of completion tokens and are never double-counted);
- cost: only when the LiteLLM gateway reports it via allowlisted
x-litellm-response-cost* headers, with an explicit cost.source — estimates are never presented as measured costs;
- gateway metadata:
x-litellm-attempted-retries / -fallbacks / -max-fallbacks / x-litellm-version;
- validation warnings (e.g. missing terminal usage chunk, mismatched totals).
What is NEVER recorded: prompts, messages, response or reasoning text, tool names/arguments/output, API URLs, server labels, authorization headers or API keys, workspace/user/file paths, raw error bodies, or full header dumps. Each request body also carries only privacy-safe correlation ids (metadata.litellm_vscode_request_id / litellm_vscode_attempt_id / litellm_vscode_client_version) so server-side spend logs can be joined to a client attempt without any content.
Correlating with LiteLLM spend logs: match the event's response.litellmCallId against the server's x-litellm-call-id (preferred), or the request metadata ids above.
License
MIT