Feima Copilot
Local & enterprise model endpoints, your own Claude/Codex/Copilot CLI subscriptions natively in VS Code, plus a local proxy to power any tool with Copilot or BYOK models

✨ Key Features
- 🖥️ Local & Enterprise Endpoints - Bring your own Ollama, LM Studio, vLLM, llama.cpp, SGLang, LiteLLM, or an enterprise/private-cloud gateway — one picker for everything you run yourself
- 🧭 Feima Auto - Automatically routes each request to the best available local/enterprise endpoint, with disclosure on every response
- 🧑💻 Agent Participants - Drive the real Claude Code, Codex, and Copilot CLI agents natively in chat with
@claude, @codex, @copilot-cli
- 🔌 Local LLM Proxy - Power any OpenAI- or Anthropic-compatible tool, even outside VS Code, with your Copilot or BYOK models
- 💬 Seamless Integration - Works directly in GitHub Copilot Chat, no interface switching needed
- 🧠 Chain-of-Thought - Full support for reasoning models, solving complex problems effortlessly
Prerequisites
Before you begin, make sure you have:
- ✅ VS Code >= 1.85.0
- ✅ GitHub Copilot Chat extension installed (required)
- ✅ A local runtime (Ollama, LM Studio, etc.), an enterprise gateway, or a Claude Code/Codex/Copilot CLI subscription — no Feima account needed to get started
📦 Installation
Step 1: Install Feima Copilot
- Open VS Code
- Press
Ctrl+Shift+X (or Cmd+Shift+X on Mac) to open Extensions
- Search for "Feima Copilot"
- Click "Install"
Step 2: Verify GitHub Copilot Chat
Make sure you have the GitHub Copilot Chat extension installed. Feima Copilot requires it to function.
- Open Extensions view
- Search for "GitHub Copilot Chat"
- If not installed, click "Install"
🚀 Quick Start
Step 3: Pick a Model Source
If you run a local model (Ollama, LM Studio, vLLM, etc.) or have access to an enterprise gateway:
- Nothing to do for the common local case — on startup, the extension quietly checks well-known local ports and adds anything it finds
- For an enterprise/private-cloud endpoint, press
Ctrl+Shift+P → "Feima Local Models: Add Model Endpoint" and give it a base URL, protocol, and optional API key
Or drive @claude, @codex, or @copilot-cli directly with your own CLI subscription — see below.
Step 4: Select a Model
- Open the Copilot Chat panel (click the chat icon in the sidebar or press
Ctrl+Alt+I)
- Click the model selector at the top of the panel
- Choose Feima Local (any endpoint you registered), Feima Auto (automatic routing), or a Copilot model
Step 5: Start Chatting
- Type your question or coding request in the chat input
- The AI will respond using the selected model
- You can switch models anytime during your session
🖥️ Local & Enterprise Model Endpoints
One picker, every model source. Feima Copilot can surface models from anything you run yourself — a laptop running Ollama or LM Studio, a self-hosted vLLM/llama.cpp/SGLang/LiteLLM instance, an Olla fleet, or an internal enterprise gateway — right in the same Copilot Chat model picker, alongside your Claude/Codex subscriptions (below).
- Auto-discovery: on startup, the extension quietly checks well-known local ports (Ollama, LM Studio, and friends) and adds anything it finds — nothing to configure for the common local case.
- Manual registration: for an enterprise or private-cloud endpoint auto-discovery can't reach, run Feima Local Models: Add Model Endpoint and give it a base URL, protocol, and optional API key.
- Team-shared endpoints: commit a
.feima/endpoints.json (URLs only, never secrets) to your repo so anyone who opens the workspace gets your team's shared gateway offered automatically.
- Refresh on demand: pulled a new local model or changed something? Run Feima Local Models: Refresh Models to re-discover immediately instead of waiting for the cache to expire.
- See what's registered: the Local & Enterprise Models view in the Explorer sidebar lists every registered endpoint, grouped personal/team, with a live health indicator and its discovered models.
Capability metadata (context window, tool-calling support) is read from the endpoint itself when available, and clearly marked as estimated in the picker when it has to be inferred — never presented as fact when it's a guess.
Note on the remote/WSL/SSH case: auto-discovery probes 127.0.0.1, which — when VS Code itself is running remotely (Remote-SSH, Remote-WSL, a dev container, Codespaces) — is the remote machine, not necessarily wherever you're actually running Ollama or LM Studio. In that setup, register the endpoint manually instead.
Feima Auto — pick a model automatically, from among your local/enterprise endpoints
Rather than manually choosing which registered endpoint to use every time, select Feima Auto in the model picker and it routes each request for you, disclosing which model was actually used and why on every response. Choose how it picks via feima.localModels.autoStrategy: local-first, balanced (default), or most-capable.
Feima Auto sticks with the same endpoint across a conversation rather than re-deciding every message, and it never hides the underlying feima-local picker entries — if a routing decision isn't what you wanted, picking a specific model directly always works.
📖 Learn more: Local & Enterprise Model Endpoints Guide
🤖 Agent Participants — Claude Code, Codex & Copilot CLI, right in chat
Feima Copilot isn't just another model provider — it also lets you drive the real Claude Code, Codex, and GitHub Copilot CLI agents from directly inside GitHub Copilot Chat.
Type @claude, @codex, or @copilot-cli and you get that CLI's own agent loop — its own planning, tool calls, and file-edit review — rendered with VS Code's native chat UI: streaming responses, inline diffs, no terminal window, no copy-pasting code back and forth.
Which one is for you?
| If you… |
Try |
Why |
| Already pay for Claude Pro/Max or ChatGPT Plus/Pro |
@claude / @codex, with the CLI's own model selected |
Uses your existing subscription directly — no extra cost, nothing else to configure |
| Don't have (or don't want) a separate Anthropic/OpenAI subscription |
@claude / @codex, with a Copilot or local model selected instead |
Same agent workflow and tool loop, powered by a model you already have access to |
| Just want GitHub Copilot CLI's terminal-automation skills without leaving chat |
@copilot-cli |
Runs Copilot CLI's agent, powered by your Copilot/local model |
| Want to try Claude Code's or Codex's workflow before committing to a subscription |
Either participant, with a Copilot/local model |
Zero new signups — see what the fuss is about first |
What it looks like
You: @claude Refactor parseInvoice() to handle malformed dates and add tests
Claude: I'll take a look at the function first...
⚙ Reading src/billing/parseInvoice.ts
⚙ Reading test/billing/parseInvoice.test.ts
✎ Editing src/billing/parseInvoice.ts
✎ Creating test/billing/parseInvoice.test.ts
Done — added guard clauses for 3 malformed date formats and 5 new test cases.
No terminal, no cd, no copy-paste — the edits land in your open editor exactly like a native Copilot Chat edit would.
Good to know
- Three permission tiers per turn —
/ask (review everything), /acceptEdits (auto-approve file edits, still ask before commands), /fullAuto (hands-off) — or set a persistent default per participant.
- Bring your own model — point any participant at a Copilot or BYOK model through a local, loopback-only Agent Proxy; no separate Anthropic/OpenAI API key required.
- MCP servers — wire your own MCP tools into
@claude and @codex via settings.
- Your CLI, your login — native mode uses the CLI's own subscription and login exactly as it would from a terminal; the extension never touches your Anthropic/OpenAI credentials.
📖 Full guide: Agent Participants · Setup & Troubleshooting · Agent Proxy
🔧 Troubleshooting
Can't find local/enterprise models in the picker
- Confirm the endpoint is actually reachable: run "Feima Local Models: Refresh Models"
- Check the Local & Enterprise Models view in the Explorer sidebar for a health indicator
- Running VS Code remotely (Remote-SSH, Remote-WSL, a dev container, Codespaces)? Auto-discovery probes
127.0.0.1 on the remote machine — register the endpoint manually instead
Agent participant doesn't respond / CLI not detected
- Enable
feima.enableDebugLogging and check the Output panel ("Feima" channel) for the resolved CLI binary path
- See the Setup & Troubleshooting guide
❓ Frequently Asked Questions
VS Code Compatibility
Q: Which VS Code versions are supported?
A: VS Code 1.85.0 and above.
Q: Does it work with VS Code Insiders?
A: Yes, fully compatible with VS Code Insiders builds.
Data Privacy & Security
Q: Is my code stored anywhere?
A: No. For local/enterprise endpoints and agent participants, your code goes only to the endpoint or CLI you've configured — Feima never sees it.
Q: Are conversation histories saved?
A: Conversations are stored locally on your device and never uploaded to Feima's servers.
📚 Documentation
🤝 Feedback & Support
📄 License
This project is licensed under the MIT License.
Feima Copilot - Bring your own models and agents to VS Code
Website | Docs | GitHub