Skip to content
| Marketplace
Sign in
Visual Studio Code>Programming Languages>TierMuxNew to Visual Studio Code? Get it now.
TierMux

TierMux

Mainul

|
132 installs
| (0) | Free
Free agentic AI coding assistant for VS Code. Multiplexes many free-tier LLM providers into one self-healing surface — auto-routes, auto-fails-over, learns your codebase. Zero setup with keyless providers; bring your own keys anytime.
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

TierMux

Stack free. Route smart. Ship faster.

A free, open-source AI coding agent for VS Code that pools 30+ free-tier LLM providers into one self-healing surface.
Works with zero setup. No subscription. No per-token bill.

VS Marketplace VS Marketplace downloads Open VSX Open VSX downloads providers MIT


Every free LLM tier rate-limits, caps out, or goes down. An assistant built on one provider inherits exactly that. TierMux pools them — one 429 and your turn quietly continues on the next provider, mid-sentence.

Install

  1. Extensions (Ctrl+Shift+X) → TierMux → Install (VSCodium / Cursor: the .vsix is on open-vsx.org)
  2. Activity Bar → TierMux → type.

Four providers ship keyless — your first message works with no key, no account, no config. Add free keys later for more headroom → Providers & keys.

A look inside

TierMux chat panel

Activity Bar → TierMux, then type. Pick Ask · Plan · Agent at the bottom left; leave the model on Auto to let routing choose, or click it to pin one. / runs a skill, @ attaches a file, the paperclip adds an image. The footer counts your lifetime tokens and what those turns would have cost on a paid API.

Everything else lives behind the ⚙ gear — six tabs, all optional:

Providers — flip a provider on, paste a free key (some are keyless), and tick only the models you want in the chain. Enabled ones sort to the top; the T V R badges mean tools, vision, reasoning. Skip this entirely and the keyless four still answer.
Custom endpoints (bottom of Providers) — name + base URL + type (OpenAI-style for Ollama / LM Studio / vLLM / LiteLLM, newer for the Responses API, Anthropic-style for Claude-compatible), key optional → Save & fetch models. Local servers get probed for their real context window.
MCP — + Add server or edit settings.json; every tool the server exposes becomes an agent tool. The registry below installs common servers (filesystem, GitHub, Fetch, Puppeteer…) in one click.
Skills — a skill is a saved prompt you run as /name. Install from the catalog (or point tiermux.skillRegistryUrl at your own); they land in .agents/skills/, shared with other agent tools. Read the source before installing — a skill runs with the agent's permissions.
Usage — lifetime tokens, request count and estimated money saved, broken down per model. Tracked locally, cleared only when you ask. Useful for spotting which free tier you actually lean on.
Others — the utility model (chat titles, commit messages), the inline-completion model, Require write confirmation (diff approval before any edit), and Command approval mode (always / safe allowlist / off). Turn on Diagnostic trace when a turn feels slow.

Why people use it

It keeps going A 429, a dead key, a 5xx, a removed model or an empty reply all fail over to the next provider silently. Failing models are cooled down; ones that 404 or reject tools are quarantined. Several keys per provider, rotated.
It finishes A turn that hits its step cap or gets stuck stops visibly and offers Continue with full memory — nothing is redone. After it edits files, your own test/typecheck/build command runs and failures go back to the agent for bounded fix rounds.
Three modes Ask answers from evidence (reads, greps, git log) and never edits. Plan investigates and hands you a plan to approve, edit or discuss. Agent edits and runs commands behind approvals.
Tools that work on weak models ripgrep grep (files-only, context, case-insensitive), paged readFile, editFile with exact failure diagnostics, a shell that Stop really kills, keyless webSearch/fetchUrl, askUser, and every tool from your MCP servers. Old tool output is aged out of the prompt so long turns stay fast.
You stay in control Command approval: ask, safe allowlist, or shell off. Diff approval for writes. Paths can't escape the workspace. Checkpoints with real undo. A Why this model? popover on every reply.
Bring your own Any OpenAI-compatible endpoint — vLLM, LiteLLM, LM Studio, Ollama, llama.cpp, Azure. Local servers are probed for their real context window.

How a turn is routed

Your message is classified (chat · coding · debug · plan · agent · vision — regex, no model call), then a chain is built: your pinned model if any, else the task table's pick, then every enabled model strongest-first — one model per provider first, so the chain spans providers instead of burning one provider's whole list.

the provider did this TierMux does this
429 cools that key → next key → next provider
401 / 402 / 403 tries your other keys, then skips the platform for this turn
404 model removed next provider; skipped for 24 h
400 with tools offered next provider; marked tool-incompatible for 10 min
5xx, timeout, headers then silence next provider
200 but empty next provider — a blank reply is not an answer

Details: Routing.

Providers

689 models across 49 providers, and the catalog updates itself — new free models and whole new providers appear without an extension update.

Keyless — zero setup Kilo Gateway · OpenCode Zen · OVH AI Endpoints · Pollinations · VLM Run
With a free API key Agnes AI · AIHubMix · Aion Labs · Api.Airforce · BAILU AI · BazaarLink · Cerebras · Charm Hyper · ChatAnywhere · Cloudflare Workers AI · Cohere · DreamPrompting · Experiential Labs · FreeInference · Google AI Studio · Groq · Kenari · LLM Gateway · LLM7 · LLMTR · MegaNova · Mistral · ModelScope · NagaAI · Nara Router · Nous Portal · NVIDIA NIM · Ollama Cloud · OpenAdapter · OpenRouter · OrcaRouter · Poolside · Requesty · Router9 · Routeway · SambaNova · Token Harbor · Token Router · Typhoon · unorouter · Vyce AI · xKiro · ZenMux · Zhipu AI
Your own any OpenAI-compatible URL — vLLM, LiteLLM, LM Studio, Ollama, llama.cpp, Azure OpenAI

More keys = fewer walls. Unused keys just rest → Providers & keys.

How it compares

TierMux Copilot Cursor Cline / Kilo
Usable every day at $0 ✓ ✗ 1 ✗ 2 ✗ 3
Works with zero keys or accounts ✓ ✗ ✗ ✗
Pools many providers' free tiers, auto-failover ✓ ✗ ✗ ✗
Many keys per provider, auto-rotated ✓ ✗ ✗ ✗
Any OpenAI-compatible endpoint ✓ ~ ~ ✓
Agent + plan modes, MCP ✓ ~ ✓ ✓
Open source, direct to provider ✓ MIT ✗ ✗ ✓

1. Copilot Free caps at 50 chat requests/month; usage-based billing since June 2026. 2. Proprietary fork; real usage sits behind paid tiers. 3. Free software, but every request bills your own API tokens.

Docs

Providers & keys · Routing · Features & modes · Tips for free models · Contributing

Use as a library

npm install tiermux
import { runAgentStream, setModelSources } from 'tiermux';

setModelSources({ catalog, settings, secrets });
await runAgentStream({
  messages: [{ role: 'user', content: 'Refactor the router' }],
  mode: 'agent', effort: 'medium',
  onChunk: (t) => process.stdout.write(t),
});

Sub-paths: tiermux/router, tiermux/agent, tiermux/providers, tiermux/shared. Headless use needs a small vscode shim — scripts/vscodeMock.cjs is the reference.

Privacy

Keys live in VS Code's encrypted secret storage. Requests go VS Code → provider, directly — no TierMux server in the path. Nothing leaves your machine but the request itself.


MIT — LICENSE · NOTICE

  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft