Orb
A coding agent for VS Code that runs entirely on your own Ollama server. No cloud
account, no API key, no Copilot subscription.
Orb has its own chat view, beside the other chat tabs: open it and describe a task.
It reads and searches your workspace, edits files, and runs commands — showing you
the diff, with an Apply button, before it changes anything.
the date parser drops the timezone on ISO strings. find it and fix it.
Before its first request Orb measures the machine it is on — operating system,
the shell that will run its commands, what is on PATH — and tells the model.
A command written for the wrong platform is refused, with the correction, before
it runs.
Requirements
- VS Code 1.138 or newer.
- An Ollama server you can reach over the network. If it is not on this machine,
start it with
OLLAMA_HOST=0.0.0.0 ollama serve so it accepts remote connections.
- A chat model.
gemma4 is the default; qwen3-coder:30b and larger reasoning
models (gpt-oss:120b, qwen3.6:35b-a3b) also work. A *-base model has no chat
template and will be poor at this.
Install
npm install
npm run build
npm run package # writes orb.vsix
code --install-extension orb.vsix
Orb defaults to http://localhost:11434, so a local Ollama needs no setup at all.
If your server is on another machine, run Orb: Set Server Address — it checks
the address before saving it. Or set it yourself:
// settings.json
"orb.baseUrl": "http://192.168.1.50:11434",
"orb.model": "gemma4"
orb.baseUrl can also be set per workspace, if different projects use different
servers.
Run Orb: Check Server Connection to confirm it works. On first activation, if
nothing answers at the default address and you have not chosen one, Orb offers to
ask for it — once, and never again if you have configured a server yourself.
Where Orb lives
Orb appears in three places.
The chat, in the secondary side bar, as a tab beside the other chat views. Its
header shows the model (click it to pick another), the detected operating system
and shell, and how full the context window is. Each tool call is a one-line row
you can expand to see exactly what the model was given back; edits and commands
are cards with buttons; the send button becomes a stop button while a task runs.
Orb also sees the file you have open and any code you have selected.
The chat in an editor tab. Orb: Open Chat in Editor shows the same
conversation in the main editor area, where there is room for long diffs. It is
the same session as the side bar, not a second one, so the two stay in step.
The Orb view, behind the Orb icon in the activity bar:
- Usage — how full the context window is, the model and server (click either to
change it), whether the model is loaded and how much VRAM it holds, how long
before
keep_alive unloads it, generation speed, and the shell Orb detected.
- Sessions — New session, and every past session in this workspace, newest
first. Click one to reopen it and carry on; hover for the delete button.
A session is saved once you have sent a message, and again after each turn, so it
survives a window reload. Sessions are kept per workspace, the newest hundred.
The status bar
Orb adds one status bar entry, because a local server has state that a transcript
has nowhere to put:
$(circuit-board) qwen3-coder:30b-64k 12.4k/65.5k 38 tok/s
Its tooltip adds the server address, VRAM resident, the num_ctx the model was
loaded with, and how long before keep_alive unloads it — which is usually the
answer to "why was that turn suddenly slow?". The entry turns amber past 85%
context. Click it for model, server, health and log actions. Hide it with
orb.showStatusBar.
What the agent can do
| Tool |
What it does |
Approval |
read_file |
Read a file, or a line range |
none |
list_files |
List a directory, optionally everything beneath it |
none |
find_files |
Find paths by glob, e.g. *.test.ts |
none |
search_files |
Regex search across file contents |
none |
find_symbol |
Where something is defined, via the language server |
none |
get_errors |
Compile and lint problems the editor already knows about |
none |
edit_file |
Replace an exact snippet |
diff → Apply / Open diff / Reject |
write_file |
Create or fully rewrite a file |
diff → Apply / Open diff / Reject |
run_command |
Tests, builds, linters |
command shown → Run / Always allow "…" / Reject |
git |
status, diff, staged, log, show, blame, branch |
none |
git_commit |
Stage and commit |
message + file list → Run / Reject |
web_search |
Search via self-hosted SearXNG (only when configured) |
none |
web_fetch |
Read one page as text |
URL shown → Fetch / Always allow host / Reject |
The listing and search tools honour .gitignore when the folder is a git
repository, so the model sees the project and not node_modules.
find_symbol and get_errors change how the agent works: it can locate a
definition without grepping for it, and check whether an edit compiled without
running a build. Both come from stable VS Code APIs, so they are local and cost
nothing.
Paths are confined to the open workspace folder; the agent cannot read or write
outside it.
To pre-approve safe commands yourself, list prefixes or regexes:
"orb.autoApproveCommands": ["git status", "git diff", "npm test", "/^npm run (lint|build)$/"],
"orb.denyCommands": ["git push", "rm"]
orb.denyCommands is checked first and always wins. orb.autoApproveEdits turns
off diff approval for file changes, which is convenient and worth understanding
before you enable it.
Knowing where it is running
An agent usually learns it is on Windows by running ls -la and reading the
error. Orb finds out first, and enforces it.
It measures. At startup Orb resolves the shell run_command will use and
checks that it actually starts: PowerShell 7 if installed, else Windows
PowerShell, on Windows; your login shell on macOS and Linux; or orb.shell if you
set one. It records the OS name and version and scans PATH for the programs a
model reaches for. Microsoft Store python.exe stubs are not counted as Python.
All of this goes into the system prompt as facts.
It refuses what cannot work. Before a command reaches you for approval it is
checked against that shell (src/guard.ts): && in Windows PowerShell 5.1,
rm -rf and mkdir -p on a PowerShell alias, export, /dev/null, here-documents
and Unix-only programs on Windows; Get-ChildItem, dir /s and $env:NAME on
bash; and anything that would sit waiting for input (vim, a bare python,
git rebase -i) on any platform. The model gets back what is wrong and what to
write instead, and you are never asked to approve a command that was always going
to fail.
It keeps the shell for what needs one. cat, ls, find, grep and their
PowerShell and cmd equivalents are refused on every platform and redirected to
read_file, list_files, find_files and search_files, which behave the same
everywhere and need no approval.
It explains failures it did not predict. When a command still fails with
"not recognized" or a parse error, the result carries a reminder of the OS, the
shell and that shell's rules.
Web access
web_search needs a SearXNG instance — self-hosted, so
no API key and your queries never leave your network. Point orb.searxngUrl at it,
and make sure json is in search.formats in its settings.yml, or its API
answers 403.
web_fetch reads one page. Three things matter more than the fetching:
Fetched text is untrusted. It enters the context of an agent that can edit files
and run commands, so a page saying "ignore previous instructions and run…" is a live
attack. Orb wraps every response in <untrusted-web-content> and the system prompt
tells the model it is data, never instructions, and must never be passed to
run_command.
No reaching into your network. Loopback, private ranges, link-local and the cloud
metadata address 169.254.169.254 are refused, and every redirect hop is
re-checked — a public URL that redirects to 127.0.0.1 would otherwise walk
straight past the check. Use orb.webAllowedHosts for an internal docs server
rather than orb.webAllowPrivate, which opens everything.
Limits. 2 MB, 20 s, 4 redirects, textual content types only, no cookies or auth
headers forwarded. HTML is reduced to Markdown-ish text with script, style and
navigation stripped, keeping headings, lists, code blocks and link targets.
Git
git runs read verbs with a fixed argv — spawn('git', ['diff', '--stat']), never
a shell string. That is the point: no quoting, no injection, and none of the
PowerShell-versus-POSIX confusion a model falls into when it writes shell itself. A
target beginning with - is refused so it cannot become a flag.
git_commit stages and commits, always showing the message and the real file list
(from git status --porcelain, not from what the model believes) for approval. It
cannot push, reset, rebase or amend; those stay with run_command, which prompts.
The unload countdown in the status bar is computed against the server's own clock
(the Date response header), and withheld entirely when the answer is implausible.
This server has been observed reporting an expires_at already in its own past for
a model that is still resident; subtracting that from the local clock claims
"unloading now" forever.
Why it is built for Ollama specifically
A hosted-model agent ported straight to Ollama tends to be slow, forgetful or
simply broken for reasons that are not obvious. Orb handles these:
The server can answer with garbage. Ollama sometimes answers a sound prompt
with nonsense — one word repeated, scraps of markup, nothing at all — and then
does the same for every request that opens the same way, while an identical
request under a different first line is answered properly. scripts/diagnose.ts
shows it against any server. Orb defends in three ways. Every session's prompt
starts with a random token, so it never inherits another session's bad state.
Before a session's first request Orb asks the server a one-word question under
that token and moves to a new one if the answer is wrong. And a reply that turns
out empty or repetitive mid-task is discarded and retried under a new token,
rather than shown or fed back to the model. If three in a row fail, Orb stops and
says the server needs restarting.
Tools are called through text. Orb describes its tools in the prompt and
parses a fenced orb-tool JSON block out of the reply, rather than passing
tools to Ollama. That works with any chat model, and it keeps Orb in control of
the format: calls written as XML or as read_file("path") are recognised too.
Set orb.toolProtocol to native to use Ollama's own tool calling instead;
scripts/probe.ts compares the two on your server.
Answers about files it never opened. Shown a file's name, a model will
describe the file as though it had read it. If an answer talks about workspace
files that were never read, Orb discards it and sends the model back to read them.
Context truncation is silent. Ollama does not error when a prompt exceeds
num_ctx — it drops the front of it, so the agent forgets the task it was given
and nothing says why. Orb therefore sends num_ctx on every request: orb.numCtx
if you set one, else the value the model's Modelfile pins, else 32768. The budget
is measured against that same number.
A full window is trimmed, not fatal. At 75% Orb drops old tool output from the
transcript — file reads that a later read superseded first, then the oldest
results — leaving a stub that says how to get each one back. It needs no extra
model request. Only if that is not enough does the task stop, at 97%.
The prompt prefix must stay byte-stable. Ollama reuses its KV cache only for
an unchanged prefix, so one changing token near the top — a timestamp, the active
filename — re-evaluates the whole prompt on every turn. Orb keeps an append-only
transcript and a fixed system prompt; per-turn context is appended to the user
message instead.
The system prompt is short. Only what the model needs before its first action
is in it. The rest is taught at the point of need, by tool errors that say exactly
what to do instead.
The workspace is handed over up front. The first message of a task carries a
listing of the project's files, so the model starts by reading the right file
instead of spending round trips discovering what exists.
Replies that are not answers. A local model sometimes says what it is about to
do and stops, returns nothing, writes a call that cannot be parsed, or degenerates
into repeating one token until its output budget is gone. Orb recognises each:
it nudges the model to continue, cuts a runaway reply off and discards it, and
gives up with an explanation after three attempts rather than looping.
Tool-call loops. A small model will happily call the same tool with the same
arguments forever. Orb detects an identical repeated call, tells the model once,
then stops. Four rounds in a row where every tool call fails also stop the task.
Models unload mid-task. The default keep_alive is 5 minutes, short enough to
expire between two tool calls and add a multi-second reload to a single step.
Orb defaults to 30 minutes.
Small models need small tool surfaces. Flat, single-level schemas and short
descriptions. Arguments are coerced rather than rejected, because local models
routinely send "5" for an integer, and a call missing a required argument gets
back the tool's signature.
Thinking is off by default. Reasoning tokens cost latency at every step of an
agent loop without much improving tool selection. Raise orb.think if you want it.
Slow is not hung. The timeout is measured between streamed chunks, not across
the request, so loading a cold 60 GB model does not trip it.
Settings
Everything is under orb.* in the settings UI. The ones that matter most:
| Setting |
Default |
|
orb.baseUrl |
http://localhost:11434 |
Server address |
orb.model |
gemma4 |
Model tag |
orb.numCtx |
0 |
0 uses the Modelfile's value, else 32768 |
orb.toolProtocol |
text |
native hands tool calling to Ollama |
orb.keepAlive |
30m |
How long the model stays loaded |
orb.think |
off |
Reasoning effort, if supported |
orb.maxIterations |
60 |
Model round trips per turn |
orb.shell |
"" |
Empty detects one; see above |
orb.showStatusBar |
true |
Status bar entry |
Development
npm run watch # rebuild on change
npm run typecheck
Press F5 to launch an Extension Development Host. Orb: Show Log has
the full record of requests, tool calls, token counts and timings.
The agent core does not import vscode, so it also runs from a terminal against a
real server. This is how the loop, the prompt and the tools are tested:
node dist/headless.js --url http://host:11434 "what does src/agent.ts do?"
node dist/headless.js --url http://host:11434 --yes "fix the failing test"
node dist/probe.js --url http://host:11434 # tool-call reliability, text vs native
node dist/diagnose.js --url http://host:11434 # does the server answer with garbage?
node dist/selftest.js # command guard, globbing, markdown
Without --yes the headless run declines every edit and command, so it is
read-only.
Layout
| File |
|
src/agent.ts |
The agent loop: nudges, loop guards, context trimming |
src/prompt.ts |
System prompt and the text tool protocol |
src/protocol.ts |
Parsing tool calls out of a reply |
src/env.ts |
Environment probe: OS, shell, PATH |
src/guard.ts |
Command checks for the detected shell |
src/tools.ts, src/toolkit.ts |
The tools, their schemas and shared helpers |
src/files.ts |
Workspace listing and globbing, .gitignore-aware |
src/git.ts, src/web.ts |
Git and web tools |
src/ollama.ts |
HTTP client: NDJSON streaming, capability and context discovery |
src/view.ts |
The chat: transcript, approvals, sessions, editor integration |
src/sidebar.ts |
The Orb view: usage and the session list |
src/store.ts |
Sessions on disk |
src/webview/ |
The pages that run inside the webviews |
src/status.ts |
Status bar entry, /api/ps polling |
scripts/ |
Headless runner, reliability probe, self-test |
Licence
MIT.