Skip to content
| Marketplace
Sign in
Visual Studio Code>AI>BYOK Models — OpenAI-Compatible ProviderNew to Visual Studio Code? Get it now.
BYOK Models — OpenAI-Compatible Provider

BYOK Models — OpenAI-Compatible Provider

Esysc

|
99 installs
| (0) | Free
Zero-touch dynamic language model provider for local inference. Works with any OpenAI-compatible endpoint — Ollama, LM Studio, vLLM, and more. Models are discovered live from /v1/models.
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

BYOK Models — OpenAI-Compatible Provider

Zero-touch dynamic language model provider for VS Code AI Chat, designed for local inference — point it at any OpenAI-compatible endpoint you configure.

Unlike a static BYOK entry in chatLanguageModels.json, this extension discovers models live from /v1/models every time VS Code builds the model picker — when the server adds, removes, or renames a model, the picker reflects it automatically. No scripts, no manual syncing.

Enhanced with project context: Your models receive workspace metadata (active files, git branch, project structure) automatically, enabling more relevant and context-aware assistance.

History

Originally, I wrote this extension to connect to my Unsloth instance, but then I realized that the extension is quite generic for OpenAI-compatible endpoints. This extension allows you to connect to any OpenAI-compatible API endpoint, making it versatile for various AI model providers that support the OpenAI API format.

Compatible providers

Works with any OpenAI-compatible endpoint — just point it at a base URL ending in /v1 and it will discover models automatically. Designed for local inference, but works with remote endpoints too. Tested and compatible with:

  • Ollama — http://localhost:11434/v1
  • LM Studio — http://localhost:1234/v1
  • vLLM — http://localhost:8000/v1
  • LocalAI — http://localhost:8080/v1
  • llama.cpp server — http://localhost:8080/v1
  • Text Generation WebUI — http://localhost:5000/v1
  • Unsloth — any self-hosted or cloud endpoint
  • OpenRouter — https://openrouter.ai/api/v1
  • Any cloud provider with an OpenAI-compatible API (Azure, together.ai, fireworks.ai, etc.)

If it serves /v1/models and /v1/chat/completions, it works.

Install

npx @vscode/vsce package --allow-missing-repository

Then in VS Code: Extensions view → ⋯ → Install from VSIX… → pick the generated byok-models-<version>.vsix, and reload the window.

Configure

Option 1: Native "Add model" (first endpoint)

  1. Open AI Chat → Manage Models → Add model → BYOK Models
  2. Enter a Name (e.g. "My Ollama"), Base URL (http://localhost:11434/v1), API Key (or leave empty for keyless servers like Ollama/LM Studio)
  3. Done — models appear in the picker.

Option 2: Command Palette (any number of endpoints)

For a second, third, etc. endpoint (VS Code's native UI only stores one per vendor):

  1. Command Palette → BYOK Models: Add Endpoint
  2. Enter a display name (e.g. "My Ollama")
  3. Enter the base URL (e.g. http://localhost:11434/v1)
  4. Enter the API key if required (leave blank for keyless servers)
  5. Repeat for each endpoint.

All endpoints appear in the model picker with their name as a suffix (e.g. llama3 · My Ollama). Use BYOK Models: Remove Endpoint to delete one.

Use

Open AI Chat → model picker → BYOK Models section. All models reported by every configured endpoint appear there dynamically. Tool calling is passed through (agent mode works) when the endpoint supports it; disable via byokModels.enableTools.

Settings

Setting Default Purpose
byokModels.endpoints [] Array of {baseUrl, apiKey, name?} — primary multi-endpoint config
byokModels.baseUrl (empty) Legacy single-endpoint base URL (fallback)
byokModels.apiKey (empty) Legacy single-endpoint API key (fallback)
byokModels.maxInputTokens 262144 Fallback context window (overridden by model's declared limit)
byokModels.maxOutputTokens 32768 Fallback max output (overridden by model's declared limit)
byokModels.enableTools true Advertise + pass through tool calling
byokModels.injectWorkspaceContext false Inject workspace metadata (files, git, project structure) into system prompts — disabled by default for local models
byokModels.maxWorkspaceContextChars 500 Max chars for workspace context (capped at 25% of model's context window)
byokModels.debugLogging false Enable verbose logging of tool calls and workspace context
byokModels.requestTimeoutMs 300000 Request timeout in ms (5 min default)

Dynamic Context Window Adaptation (v0.2.9+)

The extension now automatically reads the model's context window from /v1/models metadata during discovery. Supported fields (checked in order):

  • context_window (vLLM)
  • max_context_length (Ollama)
  • max_tokens (LM Studio)
  • n_ctx (llama.cpp)

When a model declares its context window, the extension:

  1. Sets maxInputTokens/maxOutputTokens per-model — uses the declared limit (with 25% reserved for output)
  2. Calculates workspace context budget dynamically — 25% of context window × 4 chars/token, capped by maxWorkspaceContextChars
  3. Shows context window in model picker — detail displays e.g., "context: 112,896 tokens"

This prevents "context window exceeded" errors with local models that have smaller limits.

Workspace Context Injection

When byokModels.injectWorkspaceContext is enabled (default: false for local models), the extension automatically injects workspace metadata into your model's system prompt:

  • Active file — current editor, language, line count
  • Open files — tabs currently visible in VS Code
  • Workspace folders — project structure and root paths
  • Git info — current branch, uncommitted changes count

The injected context is dynamically sized based on the model's declared context window (25% budget, capped at maxWorkspaceContextChars). For a model with 112,896 tokens, that's ~112K chars budget (capped at 500 by default).

This allows your models to provide context-aware assistance without requiring you to manually describe your project. The model understands:

  • Which files you're working on
  • Your project structure
  • Your git branch and uncommitted work
  • Your code organization

Example System Prompt with Context

=== Workspace Context ===
Workspace folders: my-app
Root: /home/user/projects/my-app
Active file: src/components/Button.tsx (typescript)
  Lines: 45, Modified: yes
Open files:
  src/components/Button.tsx (typescript)
  src/styles/Button.css (css)
  tests/Button.test.tsx (typescript)
Git branch: feature/dark-mode
  Changes: 3 modified, 1 staged

Debugging Workspace Context

Use the command BYOK Models: Show Workspace Context (Command Palette → search "BYOK Models") to see exactly what context will be injected. This is useful for understanding why models respond the way they do, or for troubleshooting if context isn't being picked up.

Tool Calling & Multi-Turn Interactions

The extension logs all tool calls and results for better debugging:

  • Tool calls — logged when your model requests a tool execution
  • Tool results — logged when VS Code returns tool output back to the model
  • Multi-turn flows — full conversation state preserved, enabling iterative problem-solving

Check the BYOK Models output channel (View → Output → select "BYOK Models") to trace the full interaction flow.

Advanced: MCP Server Integration (Future)

Model Context Protocol (MCP) support is on the roadmap for providing even richer tool access:

  • File system operations beyond basic listing
  • Custom project-specific tools
  • Standardized tool schemas for consistent behavior

For now, tools are limited to what VS Code provides natively. Stay tuned for MCP server support!

Troubleshooting

Models don't appear / "Not configured"

  1. Ensure you've added at least one endpoint via Add model (native UI) or BYOK Models: Add Endpoint (Command Palette).
  2. Check the BYOK Models output channel — it logs the base URL/key source used for discovery (setting:byokModels.endpoints, secret:..., or setting:...).
  3. Run BYOK Models: Refresh Model List after changing the configuration.
  4. If you see duplicate empty entries (e.g. byok-models 2), delete them via the gear icon next to the provider in Manage Models.

Models aren't receiving workspace context

  1. Check that byokModels.injectWorkspaceContext is set to true in settings.
  2. Run BYOK Models: Show Workspace Context to verify context is being detected.
  3. Check the BYOK Models output channel for errors or warnings.

Tool calls aren't working

  1. Verify byokModels.enableTools is true.
  2. Ensure your endpoint supports OpenAI's tool-calling format.
  3. Check the BYOK Models output channel for tool call logs and errors.

Enable debug logging

Set byokModels.debugLogging to true in settings for verbose output of all operations.

Notes

  • Vendor id is byok-models; diagnostics go to the BYOK Models output channel.
  • Once this extension works, any old static custom endpoint entries in chatLanguageModels.json can be removed to avoid duplicate model listings.
  • Workspace context is injected as a system message, so it counts toward your token limit.
  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft