Activity Bar Model Settings to add, select, test, and remove models
Native VS Code Chat model picker under Custom AI
OpenAI-compatible streaming (/chat/completions)
Multiple models and multiple endpoints
Per-model API keys (VS Code SecretStorage)
Models stored in extension state (not Settings UI)
Add, fetch, test, and manage models from Command Palette
Image / vision input support
Tool calling support
Azure OpenAI deployment endpoints with configurable api-version
Requirements
VS Code 1.104.0 or higher
An OpenAI-compatible API endpoint
Getting Started
Install the extension from the Marketplace
Open the Custom AI icon in the Activity Bar (Model Settings)
Click Add (or run Custom AI: Add Model from the Command Palette)
Enter:
Model ID
Display name
Base URL
API key (or skip for local/open servers)
Select the model from the dropdown, then use Set key, Test, Fetch, or Remove as needed
In VS Code Chat, pick the same model under Custom AI
Commands
Command
Purpose
Custom AI: Add Model
Add model + base URL + key
Custom AI: Manage Models
Add / fetch / key / test / remove
Custom AI: Fetch Models from API
Load models from /models
Custom AI: Set Model API Key
Set key for one model
Custom AI: Test Connection
Test one model
Custom AI: Remove Model
Remove one model
Custom AI: List Configured Models
List all configured models
Custom AI: Open Chat
Open the Custom AI Model Settings sidebar
API Format
POST {model.baseUrl}/chat/completions
Authorization: Bearer <per-model-key>
Content-Type: application/json
Works with OpenAI-compatible providers such as OpenAI, DeepSeek, local servers (for example Ollama with an OpenAI-compatible endpoint), and similar APIs.
For Azure OpenAI, enter the Azure resource endpoint (for example https://RESOURCE.openai.azure.com) and use the deployment name as the model ID. The extension builds the deployment chat URL and adds the selected api-version.
Privacy
API keys are stored securely with VS Code SecretStorage
Model settings are stored in extension state
Empty keys do not send an authentication header
Remote plaintext HTTP endpoints require an explicit warning confirmation