BYOK Models — OpenAI-Compatible Provider
Zero-touch dynamic language model provider for VS Code AI Chat, designed for
local inference — point it at any OpenAI-compatible endpoint you configure.
Unlike a static BYOK entry in chatLanguageModels.json, this extension discovers models
live from /v1/models every time VS Code builds the model picker — when the server adds,
removes, or renames a model, the picker reflects it automatically. No scripts, no manual
syncing.
Enhanced with project context: Your models receive workspace metadata (active files,
git branch, project structure) automatically, enabling more relevant and context-aware
assistance.
History
Originally, I wrote this extension to connect to my Unsloth instance, but then I realized that the extension is quite generic for OpenAI-compatible endpoints. This extension allows you to connect to any OpenAI-compatible API endpoint, making it versatile for various AI model providers that support the OpenAI API format.
Compatible providers
Works with any OpenAI-compatible endpoint — just point it at a base URL ending in /v1
and it will discover models automatically. Designed for local inference, but works with
remote endpoints too. Tested and compatible with:
- Ollama —
http://localhost:11434/v1
- LM Studio —
http://localhost:1234/v1
- vLLM —
http://localhost:8000/v1
- LocalAI —
http://localhost:8080/v1
- llama.cpp server —
http://localhost:8080/v1
- Text Generation WebUI —
http://localhost:5000/v1
- Unsloth — any self-hosted or cloud endpoint
- OpenRouter —
https://openrouter.ai/api/v1
- Any cloud provider with an OpenAI-compatible API (Azure, together.ai, fireworks.ai, etc.)
If it serves /v1/models and /v1/chat/completions, it works.
Install
npx @vscode/vsce package --allow-missing-repository
Then in VS Code: Extensions view → ⋯ → Install from VSIX… → pick the generated
byok-models-<version>.vsix, and reload the window.
Option 1: Native "Add model" (first endpoint)
- Open AI Chat → Manage Models → Add model → BYOK Models
- Enter a Name (e.g. "My Ollama"), Base URL (
http://localhost:11434/v1), API Key (or leave empty for keyless servers like Ollama/LM Studio)
- Done — models appear in the picker.
Option 2: Command Palette (any number of endpoints)
For a second, third, etc. endpoint (VS Code's native UI only stores one per vendor):
- Command Palette → BYOK Models: Add Endpoint
- Enter a display name (e.g. "My Ollama")
- Enter the base URL (e.g.
http://localhost:11434/v1)
- Enter the API key if required (leave blank for keyless servers)
- Repeat for each endpoint.
All endpoints appear in the model picker with their name as a suffix (e.g. llama3 · My Ollama).
Use BYOK Models: Remove Endpoint to delete one.
Use
Open AI Chat → model picker → BYOK Models section. All models reported by every
configured endpoint appear there dynamically. Tool calling is passed through (agent mode
works) when the endpoint supports it; disable via byokModels.enableTools.
Settings
| Setting |
Default |
Purpose |
byokModels.endpoints |
[] |
Array of {baseUrl, apiKey, name?} — primary multi-endpoint config |
byokModels.baseUrl |
(empty) |
Legacy single-endpoint base URL (fallback) |
byokModels.apiKey |
(empty) |
Legacy single-endpoint API key (fallback) |
byokModels.maxInputTokens |
262144 |
Fallback context window (overridden by model's declared limit) |
byokModels.maxOutputTokens |
32768 |
Fallback max output (overridden by model's declared limit) |
byokModels.enableTools |
true |
Advertise + pass through tool calling |
byokModels.injectWorkspaceContext |
false |
Inject workspace metadata (files, git, project structure) into system prompts — disabled by default for local models |
byokModels.maxWorkspaceContextChars |
500 |
Max chars for workspace context (capped at 25% of model's context window) |
byokModels.debugLogging |
false |
Enable verbose logging of tool calls and workspace context |
byokModels.requestTimeoutMs |
300000 |
Request timeout in ms (5 min default) |
Dynamic Context Window Adaptation (v0.2.9+)
The extension now automatically reads the model's context window from /v1/models metadata during discovery. Supported fields (checked in order):
context_window (vLLM)
max_context_length (Ollama)
max_tokens (LM Studio)
n_ctx (llama.cpp)
When a model declares its context window, the extension:
- Sets
maxInputTokens/maxOutputTokens per-model — uses the declared limit (with 25% reserved for output)
- Calculates workspace context budget dynamically — 25% of context window × 4 chars/token, capped by
maxWorkspaceContextChars
- Shows context window in model picker — detail displays e.g., "context: 112,896 tokens"
This prevents "context window exceeded" errors with local models that have smaller limits.
Workspace Context Injection
When byokModels.injectWorkspaceContext is enabled (default: false for local models), the extension automatically
injects workspace metadata into your model's system prompt:
- Active file — current editor, language, line count
- Open files — tabs currently visible in VS Code
- Workspace folders — project structure and root paths
- Git info — current branch, uncommitted changes count
The injected context is dynamically sized based on the model's declared context window (25% budget, capped at maxWorkspaceContextChars). For a model with 112,896 tokens, that's ~112K chars budget (capped at 500 by default).
This allows your models to provide context-aware assistance without requiring you to manually
describe your project. The model understands:
- Which files you're working on
- Your project structure
- Your git branch and uncommitted work
- Your code organization
Example System Prompt with Context
=== Workspace Context ===
Workspace folders: my-app
Root: /home/user/projects/my-app
Active file: src/components/Button.tsx (typescript)
Lines: 45, Modified: yes
Open files:
src/components/Button.tsx (typescript)
src/styles/Button.css (css)
tests/Button.test.tsx (typescript)
Git branch: feature/dark-mode
Changes: 3 modified, 1 staged
Debugging Workspace Context
Use the command BYOK Models: Show Workspace Context (Command Palette → search "BYOK Models")
to see exactly what context will be injected. This is useful for understanding why models respond
the way they do, or for troubleshooting if context isn't being picked up.
The extension logs all tool calls and results for better debugging:
- Tool calls — logged when your model requests a tool execution
- Tool results — logged when VS Code returns tool output back to the model
- Multi-turn flows — full conversation state preserved, enabling iterative problem-solving
Check the BYOK Models output channel (View → Output → select "BYOK Models") to trace
the full interaction flow.
Advanced: MCP Server Integration (Future)
Model Context Protocol (MCP) support is on the roadmap for providing even richer tool access:
- File system operations beyond basic listing
- Custom project-specific tools
- Standardized tool schemas for consistent behavior
For now, tools are limited to what VS Code provides natively. Stay tuned for MCP server support!
Troubleshooting
- Ensure you've added at least one endpoint via Add model (native UI) or BYOK Models: Add Endpoint (Command Palette).
- Check the BYOK Models output channel — it logs the base URL/key source used for discovery (
setting:byokModels.endpoints, secret:..., or setting:...).
- Run BYOK Models: Refresh Model List after changing the configuration.
- If you see duplicate empty entries (e.g.
byok-models 2), delete them via the gear icon next to the provider in Manage Models.
Models aren't receiving workspace context
- Check that
byokModels.injectWorkspaceContext is set to true in settings.
- Run BYOK Models: Show Workspace Context to verify context is being detected.
- Check the BYOK Models output channel for errors or warnings.
- Verify
byokModels.enableTools is true.
- Ensure your endpoint supports OpenAI's tool-calling format.
- Check the BYOK Models output channel for tool call logs and errors.
Enable debug logging
Set byokModels.debugLogging to true in settings for verbose output of all operations.
Notes
- Vendor id is
byok-models; diagnostics go to the BYOK Models output channel.
- Once this extension works, any old static custom endpoint entries in
chatLanguageModels.json can be removed to avoid duplicate model listings.
- Workspace context is injected as a system message, so it counts toward your token limit.