Stack free. Route smart. Ship faster.
A free, open-source AI coding agent for VS Code that pools 30+ free-tier LLM
providers into one self-healing surface.
Works with zero setup. No subscription. No per-token bill.
Every free LLM tier rate-limits, caps out, or goes down. An assistant built on one
provider inherits exactly that. TierMux pools them — one 429 and your turn quietly continues
on the next provider, mid-sentence.
Install
- Extensions (
Ctrl+Shift+X) → TierMux → Install
(VSCodium / Cursor: the .vsix is on open-vsx.org)
- Activity Bar → TierMux → type.
Four providers ship keyless — your first message works with no key, no account, no
config. Add free keys later for more headroom → Providers & keys.
A look inside
Activity Bar → TierMux, then type. Pick Ask · Plan · Agent at the bottom left; leave the
model on Auto to let routing choose, or click it to pin one. / runs a skill, @ attaches a
file, the paperclip adds an image. The footer counts your lifetime tokens and what those turns
would have cost on a paid API.
Everything else lives behind the ⚙ gear — six tabs, all optional:
|
|
 |
Providers — flip a provider on, paste a free key (some are keyless), and tick only the models you want in the chain. Enabled ones sort to the top; the T V R badges mean tools, vision, reasoning. Skip this entirely and the keyless four still answer. |
 |
Custom endpoints (bottom of Providers) — name + base URL + type (OpenAI-style for Ollama / LM Studio / vLLM / LiteLLM, newer for the Responses API, Anthropic-style for Claude-compatible), key optional → Save & fetch models. Local servers get probed for their real context window. |
 |
MCP — + Add server or edit settings.json; every tool the server exposes becomes an agent tool. The registry below installs common servers (filesystem, GitHub, Fetch, Puppeteer…) in one click. |
 |
Skills — a skill is a saved prompt you run as /name. Install from the catalog (or point tiermux.skillRegistryUrl at your own); they land in .agents/skills/, shared with other agent tools. Read the source before installing — a skill runs with the agent's permissions. |
 |
Usage — lifetime tokens, request count and estimated money saved, broken down per model. Tracked locally, cleared only when you ask. Useful for spotting which free tier you actually lean on. |
 |
Others — the utility model (chat titles, commit messages), the inline-completion model, Require write confirmation (diff approval before any edit), and Command approval mode (always / safe allowlist / off). Turn on Diagnostic trace when a turn feels slow. |
Why people use it
|
|
| It keeps going |
A 429, a dead key, a 5xx, a removed model or an empty reply all fail over to the next provider silently. Failing models are cooled down; ones that 404 or reject tools are quarantined. Several keys per provider, rotated. |
| It finishes |
A turn that hits its step cap or gets stuck stops visibly and offers Continue with full memory — nothing is redone. After it edits files, your own test/typecheck/build command runs and failures go back to the agent for bounded fix rounds. |
| Three modes |
Ask answers from evidence (reads, greps, git log) and never edits. Plan investigates and hands you a plan to approve, edit or discuss. Agent edits and runs commands behind approvals. |
| Tools that work on weak models |
ripgrep grep (files-only, context, case-insensitive), paged readFile, editFile with exact failure diagnostics, a shell that Stop really kills, keyless webSearch/fetchUrl, askUser, and every tool from your MCP servers. Old tool output is aged out of the prompt so long turns stay fast. |
| You stay in control |
Command approval: ask, safe allowlist, or shell off. Diff approval for writes. Paths can't escape the workspace. Checkpoints with real undo. A Why this model? popover on every reply. |
| Bring your own |
Any OpenAI-compatible endpoint — vLLM, LiteLLM, LM Studio, Ollama, llama.cpp, Azure. Local servers are probed for their real context window. |
How a turn is routed
Your message is classified (chat · coding · debug · plan · agent · vision — regex, no model
call), then a chain is built: your pinned model if any, else the task table's pick, then
every enabled model strongest-first — one model per provider first, so the chain spans
providers instead of burning one provider's whole list.
| the provider did this |
TierMux does this |
429 |
cools that key → next key → next provider |
401 / 402 / 403 |
tries your other keys, then skips the platform for this turn |
404 model removed |
next provider; skipped for 24 h |
400 with tools offered |
next provider; marked tool-incompatible for 10 min |
5xx, timeout, headers then silence |
next provider |
200 but empty |
next provider — a blank reply is not an answer |
Details: Routing.
Providers
689 models across 49 providers, and the catalog updates itself —
new free models and whole new providers appear without an extension update.
|
|
| Keyless — zero setup |
Kilo Gateway · OpenCode Zen · OVH AI Endpoints · Pollinations · VLM Run |
| With a free API key |
Agnes AI · AIHubMix · Aion Labs · Api.Airforce · BAILU AI · BazaarLink · Cerebras · Charm Hyper · ChatAnywhere · Cloudflare Workers AI · Cohere · DreamPrompting · Experiential Labs · FreeInference · Google AI Studio · Groq · Kenari · LLM Gateway · LLM7 · LLMTR · MegaNova · Mistral · ModelScope · NagaAI · Nara Router · Nous Portal · NVIDIA NIM · Ollama Cloud · OpenAdapter · OpenRouter · OrcaRouter · Poolside · Requesty · Router9 · Routeway · SambaNova · Token Harbor · Token Router · Typhoon · unorouter · Vyce AI · xKiro · ZenMux · Zhipu AI |
| Your own |
any OpenAI-compatible URL — vLLM, LiteLLM, LM Studio, Ollama, llama.cpp, Azure OpenAI |
More keys = fewer walls. Unused keys just rest → Providers & keys.
How it compares
|
TierMux |
Copilot |
Cursor |
Cline / Kilo |
| Usable every day at $0 |
✓ |
✗ 1 |
✗ 2 |
✗ 3 |
| Works with zero keys or accounts |
✓ |
✗ |
✗ |
✗ |
| Pools many providers' free tiers, auto-failover |
✓ |
✗ |
✗ |
✗ |
| Many keys per provider, auto-rotated |
✓ |
✗ |
✗ |
✗ |
| Any OpenAI-compatible endpoint |
✓ |
~ |
~ |
✓ |
| Agent + plan modes, MCP |
✓ |
~ |
✓ |
✓ |
| Open source, direct to provider |
✓ MIT |
✗ |
✗ |
✓ |
1. Copilot Free caps at 50 chat requests/month; usage-based billing since June 2026. 2. Proprietary fork; real usage sits behind paid tiers. 3. Free software, but every request bills your own API tokens.
Docs
Providers & keys · Routing · Features & modes · Tips for free models · Contributing
Use as a library
npm install tiermux
import { runAgentStream, setModelSources } from 'tiermux';
setModelSources({ catalog, settings, secrets });
await runAgentStream({
messages: [{ role: 'user', content: 'Refactor the router' }],
mode: 'agent', effort: 'medium',
onChunk: (t) => process.stdout.write(t),
});
Sub-paths: tiermux/router, tiermux/agent, tiermux/providers, tiermux/shared. Headless
use needs a small vscode shim — scripts/vscodeMock.cjs is the reference.
Privacy
Keys live in VS Code's encrypted secret storage. Requests go VS Code → provider, directly —
no TierMux server in the path. Nothing leaves your machine but the request itself.
MIT — LICENSE · NOTICE
| |