Copilot Models
Unlock third-party large language model extensions for GitHub Copilot.
Seamlessly integrate DeepSeek, Zhipu AI, and Qwen LLMs.
One-click switching and native panel compatibility.
Features
- Multi-Model Support: DeepSeek V4, Zhipu AI GLM-5, Qwen 3 series
- Model Routing: Automatic failover and latency-based routing
- Tool Calling: Full Copilot Chat tool calling support
- Thinking Mode: Model reasoning/thinking mode support
- Vision Proxy: Image description proxy for non-vision models
via VS Code LM or custom API
- Circuit Breaker: Automatic failure protection with retry
- Secure Authentication: API keys stored in VS Code SecretStorage;
sensitive values (keys, tokens, URLs) are automatically redacted from logs
- Log Debugging: 4-level logging with hot-reload
- Lightweight: OpenAI SDK replaced with native SSE client code
- Token Plan: Unified prepaid billing for Qwen, DeepSeek, and
GLM token packages via a single endpoint
- Utility Model: Optionally run AI commit-message generation on one of
these models, through VS Code's utility-model settings
Documentation
Quick Start
1. Install Extension
Install from the VS Code Extension Marketplace.
Press Ctrl+Shift+P (macOS: Cmd+Shift+P), run Copilot Models: Set API Key,
select a provider and enter your key.
Note: API keys are stored in VS Code SecretStorage, not as plain settings.
If you use prepaid token packages (e.g., Alibaba DashScope plan),
run Copilot Models: Set Token Plan to configure plan access:
- Press
Ctrl+Shift+P (macOS: Cmd+Shift+P), run Copilot Models: Set Token Plan
- Select a built-in provider preset or enter a custom URL
- The Qwen preset is preconfigured with the endpoint URL and
6 supported models
- Enter the plan API token
- Select the models covered by this plan
The plan token is saved in VS Code SecretStorage.
- Qwen Token Plan —
https://token-plan.cn-beijing.maas.aliyuncs.com/compatible-mode/v1
The Qwen Token Plan preset covers Qwen, DeepSeek, and GLM models
in a single plan.
For other providers, choose "Custom URL" and enter the plan API endpoint.
Run Copilot Models: Clear Token Plan to remove a configured plan.
If you want to use image attachments with models that don't natively support
image input (e.g., GLM-5.3, GLM-5.2), configure a vision proxy to
automatically convert images to text descriptions:
- Press
Ctrl+Shift+P (macOS: Cmd+Shift+P), run
Copilot Models: Set Vision Model
- Select a vision-capable model, or choose "Custom API Endpoint"
- For custom API endpoint, enter the URL, model ID, and API key
(leave the key empty for unauthenticated endpoints)
The vision proxy describes images before sending them to the chat model.
For custom API endpoints, an OpenAI-compatible endpoint is required — enter
either the base URL (https://host/v1) or the full
/chat/completions URL, both are accepted. The API key is stored in
VS Code SecretStorage.
Note: The proxy only applies to models that cannot accept image input
natively. Models with image support receive the original images and are
never routed through the proxy.
Only the most recent message is described. Images from earlier turns — and
any image whose description fails or is empty — are replaced with
[Image: description unavailable], so no image part reaches a model that
cannot accept one.
Run Copilot Models: Clear Vision Model to remove the configuration.
5. Start Using
- Open GitHub Copilot Chat panel
- Click on the model selector
- Select the model to use
- Start chatting
Supported Models
The Thinking Effort selector in the model picker maps onto what each
provider documents. None turns thinking off everywhere: Qwen through
enable_thinking: false — its API defaults thinking to on — and Zhipu AI and
DeepSeek through their thinking toggle. All three document an on/off switch
and no effort level, so low, high and max enable thinking at the
provider's own default effort.
Qwen (Alibaba Cloud)
| Model |
Context |
Output |
Tool Calling |
Image Input |
Thinking Mode |
| Qwen3.8 Max |
1M |
64K |
✅ |
✅ |
✅ |
| Qwen3.8 Flash |
1M |
64K |
✅ |
✅ |
✅ |
| Qwen3.7 Plus |
1M |
64K |
✅ |
✅ |
✅ |
| Qwen3.7 Flash |
1M |
64K |
✅ |
✅ |
✅ |
DeepSeek
| Model |
Context |
Output |
Tool Calling |
Image Input |
Thinking Mode |
| DeepSeek V4.1 Flash |
1M |
384K |
✅ |
✅ |
✅ |
Note: deepseek-v4-flash has been retired and deepseek-v4-pro is being
retired — requests to either ID are served by DeepSeek-V4.1-Flash.
Zhipu AI (BigModel)
| Model |
Context |
Output |
Tool Calling |
Image Input |
Thinking Mode |
| GLM-5.3 |
1M |
128K |
✅ |
❌ |
✅ |
| GLM-5.3-Flash |
1M |
128K |
✅ |
✅ |
✅ |
| GLM-5.2 |
1M |
128K |
✅ |
❌ |
✅ |
| GLM-5.1 |
200K |
128K |
✅ |
❌ |
✅ |
| GLM-5-Turbo |
200K |
128K |
✅ |
❌ |
✅ |
| GLM-5 |
200K |
128K |
✅ |
❌ |
✅ |
| GLM-4.7-Flash |
200K |
128K |
✅ |
❌ |
❌ |
Tip: Models marked with ❌ for Image Input can still handle images
through the Vision Proxy feature (see Quick Start step 4).
Token Plan Coverage
The built-in Qwen Token Plan preset supports the following models through
a single unified endpoint:
| Model |
ID |
| Qwen3.8 Max |
qwen3.8-max |
| Qwen3.8 Flash |
qwen3.8-flash |
| Qwen3.7 Plus |
qwen3.7-plus |
| Qwen3.7 Flash |
qwen3.7-flash |
| GLM-5.2 |
glm-5.2 |
| DeepSeek V4.1 Flash |
deepseek-flash |
Models not listed (e.g. GLM-5-Turbo, kimi-k2.7-code) are still available via direct
provider API access — they are simply not covered by this Token Plan preset.
Note: The IDs above must match the model IDs exposed by the extension
(see the Supported Models tables). deepseek-v4-pro and deepseek-v4-flash
are accepted by the DeepSeek API itself, but the extension only exposes them
as deepseek-flash.
Configuration Options
Available in VS Code settings (search copilot-models):
Provider Settings
| Config |
Description |
Default |
<provider>.enabled |
Enable this provider |
true |
<provider>.baseUrl |
API base URL (e.g. deepseek.baseUrl) |
per provider |
Global Settings
| Config |
Description |
Default |
routingStrategy |
"failover" or "latency" routing |
"failover" |
failoverModels |
Primary model → fallback ID map (chained A→B→C) |
{} |
modelIdOverrides |
Map model IDs to custom API names |
{} |
maxImageSize |
Max image size in bytes (0 = no limit) |
20971520 (20MB) |
timeoutMs |
Request timeout in milliseconds |
60000 |
maxRetries |
Maximum retry attempts |
1 |
showStatusBar |
Show today's token usage in the status bar |
true |
editTools |
File-editing tools to advertise to the editor |
[] |
debugMode |
Log level: minimal / metadata / verbose |
minimal |
On editTools: left empty (the default), the editor tries several edit
tools and picks one itself. Accepted values are find-replace,
multi-find-replace, apply-patch and code-rewrite; fill them in only
when you know which editing tool suits the models, because naming the wrong
one makes their edits worse. The setting applies to every model.
Vision Proxy Settings
| Config |
Description |
Default |
visionModel |
Vision model ID (empty for auto-detect) |
"" |
visionPrompt |
Prompt for vision proxy description |
"Describe all..." |
visionProxy.apiUrl |
Vision proxy API URL (OpenAI-compatible) |
"" |
visionProxy.apiModelId |
Model ID for vision proxy API endpoint |
"" |
visionProxy.timeoutMs |
Vision proxy timeout in milliseconds |
60000 |
visionProxy.maxTokens |
Max tokens for vision proxy response |
1024 |
Note: visionProxy.maxTokens applies only to the vision description
request sent to visionProxy.apiUrl — it does not affect normal chat
requests. Chat requests have no separate max_tokens setting: each model
automatically uses its own maxOutputTokens as the API's max_tokens
parameter. See the "Output" column in the Supported Models tables above.
Token Usage
Every completed request records its token usage locally — including requests
that use a directly configured API key, not only token plan traffic.
- The status bar shows today's tokens and request count, for example
12.3K tok · 18 req. Click it to open the report. The item stays hidden
until the first request is recorded; set copilot-models.showStatusBar to
false to hide it permanently.
- Run
Copilot Models: Show Token Usage for a breakdown by plan and by model.
- Run
Copilot Models: Clear Token Usage to drop the recorded history.
Note: Only the most recent 1000 requests are kept, so the "all time"
figures are a rolling window rather than a lifetime total. Usage is stored in
VS Code global state and never leaves your machine.
Account Balance
The usage report also shows your DeepSeek account balance, queried from the
official GET /user/balance endpoint. It refreshes each time you run the
command, and is shown even before any request has been recorded.
Balance:
deepseek: ¥110.00 (granted ¥10.00 · topped up ¥100.00)
Note: DeepSeek is the only supported provider with a documented balance
API. Zhipu AI, Qwen/DashScope and the Qwen Token Plan endpoint expose none, so
no figure is shown for requests served through them. Providers without a
configured API key are omitted entirely; a configured provider whose lookup
fails reads as unavailable — a balance problem never blocks the rest of the
report. No request is sent when no API key is set.
AI Commit Messages
The sparkle button in the Source Control input box drafts a commit message.
The GitHub Copilot Chat extension provides that button, and by default it runs
on Copilot's own small utility model. It can run on a model from this
extension instead.
Point it at one by setting chat.utilitySmallModel to the <vendor>/<model-id>
value the dropdown stores. Every model from this extension is served by the
Copilot Models router, whose vendor id is copilot-models-router:
"chat.utilitySmallModel": "copilot-models-router/deepseek-flash"
Note: The vendor is the one that registers the models, not the upstream
service. deepseek/deepseek-flash looks right and is silently ignored — the
setting is only read by name <vendor>/<id>, and no model is registered
under the deepseek vendor. Picking from the dropdown avoids the guesswork.
| Setting |
Applies to |
chat.utilitySmallModel |
Short, frequent flows: commit messages |
chat.utilityModel |
Longer utility flows |
chat.byokUtilityModelDefault |
Used when the two above are empty |
chat.byokUtilityModelDefault decides what happens when the model selected
in the chat picker is a BYOK model and neither override is set: copilot
(the default) keeps Copilot's utility models, mainAgent reuses the selected
BYOK model, and none disables utility models.
Notes:
- The model must be selectable, which means its API key or a covering token
plan is configured. Without credentials a model is not offered in these
settings, and a request it does receive fails with
API key not configured.
- Prefer a fast, inexpensive model: these flows run often and the prompt is
mostly the diff.
- Only the Copilot extension's utility aliases follow these settings. VS
Code's own internal flows (chat titles, dictation cleanup, tool risk
assessment) ask for the Copilot utility model directly and ignore them.
- To change what the message says rather than which model writes it, use
github.copilot.chat.commitMessageGeneration.instructions.
Commands
| Command |
Description |
Copilot Models: Set API Key |
Configure API key (select provider first) |
Copilot Models: Clear API Key |
Clear API key (select provider first) |
Copilot Models: Open Settings |
Open extension settings |
Copilot Models: Show Log |
Show log panel |
Copilot Models: Clear Log |
Clear logs |
Copilot Models: Refresh Models |
Refresh model list |
Copilot Models: Show Latency Stats |
Show provider latency statistics |
Copilot Models: Show Token Usage |
Show usage by plan and model |
Copilot Models: Clear Token Usage |
Clear all recorded token usage |
Copilot Models: Set Token Plan |
Configure prepaid token plan |
Copilot Models: Clear Token Plan |
Remove configured token plan |
Copilot Models: Set Vision Model |
Configure vision image proxy |
Copilot Models: Clear Vision Model |
Clear vision proxy configuration |
Debugging
If you encounter issues, check the logs:
- Press
Ctrl+Shift+P, run Copilot Models: Show Log
- Logs appear in the "Copilot Models" output panel
- Set
debugMode to verbose for detailed debug output
- Set to
minimal (default) for warnings and errors only
Log level changes take effect immediately without reloading the extension.
Each request's logs carry a structured prefix
(req=<id> provider=<id> model=<id>), so you can grep for a single req=<id>
to trace one request across routing, provider, and network layers.
API keys, tokens, and URL query strings are automatically redacted from the
logs — secrets never appear in the output panel.
Diagnosing images
Any request that carries an image writes one line describing what became of
it, for example:
[WARN ] [Vision] req=0043ef provider=deepseek model=glm-5.3 \
Image handling: found=1 inMessages=1 described=0 failed=1 unavailable=0 omitted=0
| Field |
Meaning |
found |
Image parts in the request |
inMessages |
Messages holding at least one image |
described |
Images replaced by a description |
failed |
The description request failed or came back empty |
unavailable |
No vision proxy is configured |
omitted |
Images from earlier turns, which are not described again |
bypassed |
Left alone because the model accepts images itself |
model |
Vision model that produced the description |
A request with no images writes no such line. When failed or unavailable
is non-zero the line is logged as a warning, so it appears at the default
minimal level; a request whose images were all handled is logged at info,
which needs metadata or verbose to show up.
License
Apache-2.0