Skip to content
| Marketplace
Sign in
Visual Studio Code>Programming Languages>OrbNew to Visual Studio Code? Get it now.
Orb

Orb

the-b4rd

|
1 install
| (0) | Free
A coding agent for VS Code, with its own sidebar, that runs entirely on your own Ollama server.
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

Orb

A coding agent for VS Code that runs entirely on your own Ollama server. No cloud account, no API key, no Copilot subscription.

Orb has its own chat view, beside the other chat tabs: open it and describe a task. It reads and searches your workspace, edits files, and runs commands — showing you the diff, with an Apply button, before it changes anything.

the date parser drops the timezone on ISO strings. find it and fix it.

Before its first request Orb measures the machine it is on — operating system, the shell that will run its commands, what is on PATH — and tells the model. A command written for the wrong platform is refused, with the correction, before it runs.

Requirements

  • VS Code 1.138 or newer.
  • An Ollama server you can reach over the network. If it is not on this machine, start it with OLLAMA_HOST=0.0.0.0 ollama serve so it accepts remote connections.
  • A chat model. gemma4 is the default; qwen3-coder:30b and larger reasoning models (gpt-oss:120b, qwen3.6:35b-a3b) also work. A *-base model has no chat template and will be poor at this.

Install

npm install
npm run build
npm run package        # writes orb.vsix
code --install-extension orb.vsix

Orb defaults to http://localhost:11434, so a local Ollama needs no setup at all.

If your server is on another machine, run Orb: Set Server Address — it checks the address before saving it. Or set it yourself:

// settings.json
"orb.baseUrl": "http://192.168.1.50:11434",
"orb.model": "gemma4"

orb.baseUrl can also be set per workspace, if different projects use different servers.

Run Orb: Check Server Connection to confirm it works. On first activation, if nothing answers at the default address and you have not chosen one, Orb offers to ask for it — once, and never again if you have configured a server yourself.

Where Orb lives

Orb appears in three places.

The chat, in the secondary side bar, as a tab beside the other chat views. Its header shows the model (click it to pick another), the detected operating system and shell, and how full the context window is. Each tool call is a one-line row you can expand to see exactly what the model was given back; edits and commands are cards with buttons; the send button becomes a stop button while a task runs. Orb also sees the file you have open and any code you have selected.

The chat in an editor tab. Orb: Open Chat in Editor shows the same conversation in the main editor area, where there is room for long diffs. It is the same session as the side bar, not a second one, so the two stay in step.

The Orb view, behind the Orb icon in the activity bar:

  • Usage — how full the context window is, the model and server (click either to change it), whether the model is loaded and how much VRAM it holds, how long before keep_alive unloads it, generation speed, and the shell Orb detected.
  • Sessions — New session, and every past session in this workspace, newest first. Click one to reopen it and carry on; hover for the delete button.

A session is saved once you have sent a message, and again after each turn, so it survives a window reload. Sessions are kept per workspace, the newest hundred.

The status bar

Orb adds one status bar entry, because a local server has state that a transcript has nowhere to put:

$(circuit-board) qwen3-coder:30b-64k  12.4k/65.5k  38 tok/s

Its tooltip adds the server address, VRAM resident, the num_ctx the model was loaded with, and how long before keep_alive unloads it — which is usually the answer to "why was that turn suddenly slow?". The entry turns amber past 85% context. Click it for model, server, health and log actions. Hide it with orb.showStatusBar.

What the agent can do

Tool What it does Approval
read_file Read a file, or a line range none
list_files List a directory, optionally everything beneath it none
find_files Find paths by glob, e.g. *.test.ts none
search_files Regex search across file contents none
find_symbol Where something is defined, via the language server none
get_errors Compile and lint problems the editor already knows about none
edit_file Replace an exact snippet diff → Apply / Open diff / Reject
write_file Create or fully rewrite a file diff → Apply / Open diff / Reject
run_command Tests, builds, linters command shown → Run / Always allow "…" / Reject
git status, diff, staged, log, show, blame, branch none
git_commit Stage and commit message + file list → Run / Reject
web_search Search via self-hosted SearXNG (only when configured) none
web_fetch Read one page as text URL shown → Fetch / Always allow host / Reject

The listing and search tools honour .gitignore when the folder is a git repository, so the model sees the project and not node_modules.

find_symbol and get_errors change how the agent works: it can locate a definition without grepping for it, and check whether an edit compiled without running a build. Both come from stable VS Code APIs, so they are local and cost nothing.

Paths are confined to the open workspace folder; the agent cannot read or write outside it.

To pre-approve safe commands yourself, list prefixes or regexes:

"orb.autoApproveCommands": ["git status", "git diff", "npm test", "/^npm run (lint|build)$/"],
"orb.denyCommands": ["git push", "rm"]

orb.denyCommands is checked first and always wins. orb.autoApproveEdits turns off diff approval for file changes, which is convenient and worth understanding before you enable it.

Knowing where it is running

An agent usually learns it is on Windows by running ls -la and reading the error. Orb finds out first, and enforces it.

It measures. At startup Orb resolves the shell run_command will use and checks that it actually starts: PowerShell 7 if installed, else Windows PowerShell, on Windows; your login shell on macOS and Linux; or orb.shell if you set one. It records the OS name and version and scans PATH for the programs a model reaches for. Microsoft Store python.exe stubs are not counted as Python. All of this goes into the system prompt as facts.

It refuses what cannot work. Before a command reaches you for approval it is checked against that shell (src/guard.ts): && in Windows PowerShell 5.1, rm -rf and mkdir -p on a PowerShell alias, export, /dev/null, here-documents and Unix-only programs on Windows; Get-ChildItem, dir /s and $env:NAME on bash; and anything that would sit waiting for input (vim, a bare python, git rebase -i) on any platform. The model gets back what is wrong and what to write instead, and you are never asked to approve a command that was always going to fail.

It keeps the shell for what needs one. cat, ls, find, grep and their PowerShell and cmd equivalents are refused on every platform and redirected to read_file, list_files, find_files and search_files, which behave the same everywhere and need no approval.

It explains failures it did not predict. When a command still fails with "not recognized" or a parse error, the result carries a reminder of the OS, the shell and that shell's rules.

Web access

web_search needs a SearXNG instance — self-hosted, so no API key and your queries never leave your network. Point orb.searxngUrl at it, and make sure json is in search.formats in its settings.yml, or its API answers 403.

web_fetch reads one page. Three things matter more than the fetching:

Fetched text is untrusted. It enters the context of an agent that can edit files and run commands, so a page saying "ignore previous instructions and run…" is a live attack. Orb wraps every response in <untrusted-web-content> and the system prompt tells the model it is data, never instructions, and must never be passed to run_command.

No reaching into your network. Loopback, private ranges, link-local and the cloud metadata address 169.254.169.254 are refused, and every redirect hop is re-checked — a public URL that redirects to 127.0.0.1 would otherwise walk straight past the check. Use orb.webAllowedHosts for an internal docs server rather than orb.webAllowPrivate, which opens everything.

Limits. 2 MB, 20 s, 4 redirects, textual content types only, no cookies or auth headers forwarded. HTML is reduced to Markdown-ish text with script, style and navigation stripped, keeping headings, lists, code blocks and link targets.

Git

git runs read verbs with a fixed argv — spawn('git', ['diff', '--stat']), never a shell string. That is the point: no quoting, no injection, and none of the PowerShell-versus-POSIX confusion a model falls into when it writes shell itself. A target beginning with - is refused so it cannot become a flag.

git_commit stages and commits, always showing the message and the real file list (from git status --porcelain, not from what the model believes) for approval. It cannot push, reset, rebase or amend; those stay with run_command, which prompts.

The unload countdown in the status bar is computed against the server's own clock (the Date response header), and withheld entirely when the answer is implausible. This server has been observed reporting an expires_at already in its own past for a model that is still resident; subtracting that from the local clock claims "unloading now" forever.

Why it is built for Ollama specifically

A hosted-model agent ported straight to Ollama tends to be slow, forgetful or simply broken for reasons that are not obvious. Orb handles these:

The server can answer with garbage. Ollama sometimes answers a sound prompt with nonsense — one word repeated, scraps of markup, nothing at all — and then does the same for every request that opens the same way, while an identical request under a different first line is answered properly. scripts/diagnose.ts shows it against any server. Orb defends in three ways. Every session's prompt starts with a random token, so it never inherits another session's bad state. Before a session's first request Orb asks the server a one-word question under that token and moves to a new one if the answer is wrong. And a reply that turns out empty or repetitive mid-task is discarded and retried under a new token, rather than shown or fed back to the model. If three in a row fail, Orb stops and says the server needs restarting.

Tools are called through text. Orb describes its tools in the prompt and parses a fenced orb-tool JSON block out of the reply, rather than passing tools to Ollama. That works with any chat model, and it keeps Orb in control of the format: calls written as XML or as read_file("path") are recognised too. Set orb.toolProtocol to native to use Ollama's own tool calling instead; scripts/probe.ts compares the two on your server.

Answers about files it never opened. Shown a file's name, a model will describe the file as though it had read it. If an answer talks about workspace files that were never read, Orb discards it and sends the model back to read them.

Context truncation is silent. Ollama does not error when a prompt exceeds num_ctx — it drops the front of it, so the agent forgets the task it was given and nothing says why. Orb therefore sends num_ctx on every request: orb.numCtx if you set one, else the value the model's Modelfile pins, else 32768. The budget is measured against that same number.

A full window is trimmed, not fatal. At 75% Orb drops old tool output from the transcript — file reads that a later read superseded first, then the oldest results — leaving a stub that says how to get each one back. It needs no extra model request. Only if that is not enough does the task stop, at 97%.

The prompt prefix must stay byte-stable. Ollama reuses its KV cache only for an unchanged prefix, so one changing token near the top — a timestamp, the active filename — re-evaluates the whole prompt on every turn. Orb keeps an append-only transcript and a fixed system prompt; per-turn context is appended to the user message instead.

The system prompt is short. Only what the model needs before its first action is in it. The rest is taught at the point of need, by tool errors that say exactly what to do instead.

The workspace is handed over up front. The first message of a task carries a listing of the project's files, so the model starts by reading the right file instead of spending round trips discovering what exists.

Replies that are not answers. A local model sometimes says what it is about to do and stops, returns nothing, writes a call that cannot be parsed, or degenerates into repeating one token until its output budget is gone. Orb recognises each: it nudges the model to continue, cuts a runaway reply off and discards it, and gives up with an explanation after three attempts rather than looping.

Tool-call loops. A small model will happily call the same tool with the same arguments forever. Orb detects an identical repeated call, tells the model once, then stops. Four rounds in a row where every tool call fails also stop the task.

Models unload mid-task. The default keep_alive is 5 minutes, short enough to expire between two tool calls and add a multi-second reload to a single step. Orb defaults to 30 minutes.

Small models need small tool surfaces. Flat, single-level schemas and short descriptions. Arguments are coerced rather than rejected, because local models routinely send "5" for an integer, and a call missing a required argument gets back the tool's signature.

Thinking is off by default. Reasoning tokens cost latency at every step of an agent loop without much improving tool selection. Raise orb.think if you want it.

Slow is not hung. The timeout is measured between streamed chunks, not across the request, so loading a cold 60 GB model does not trip it.

Settings

Everything is under orb.* in the settings UI. The ones that matter most:

Setting Default
orb.baseUrl http://localhost:11434 Server address
orb.model gemma4 Model tag
orb.numCtx 0 0 uses the Modelfile's value, else 32768
orb.toolProtocol text native hands tool calling to Ollama
orb.keepAlive 30m How long the model stays loaded
orb.think off Reasoning effort, if supported
orb.maxIterations 60 Model round trips per turn
orb.shell "" Empty detects one; see above
orb.showStatusBar true Status bar entry

Development

npm run watch      # rebuild on change
npm run typecheck

Press F5 to launch an Extension Development Host. Orb: Show Log has the full record of requests, tool calls, token counts and timings.

The agent core does not import vscode, so it also runs from a terminal against a real server. This is how the loop, the prompt and the tools are tested:

node dist/headless.js --url http://host:11434 "what does src/agent.ts do?"
node dist/headless.js --url http://host:11434 --yes "fix the failing test"
node dist/probe.js --url http://host:11434      # tool-call reliability, text vs native
node dist/diagnose.js --url http://host:11434   # does the server answer with garbage?
node dist/selftest.js                           # command guard, globbing, markdown

Without --yes the headless run declines every edit and command, so it is read-only.

Layout

File
src/agent.ts The agent loop: nudges, loop guards, context trimming
src/prompt.ts System prompt and the text tool protocol
src/protocol.ts Parsing tool calls out of a reply
src/env.ts Environment probe: OS, shell, PATH
src/guard.ts Command checks for the detected shell
src/tools.ts, src/toolkit.ts The tools, their schemas and shared helpers
src/files.ts Workspace listing and globbing, .gitignore-aware
src/git.ts, src/web.ts Git and web tools
src/ollama.ts HTTP client: NDJSON streaming, capability and context discovery
src/view.ts The chat: transcript, approvals, sessions, editor integration
src/sidebar.ts The Orb view: usage and the session list
src/store.ts Sessions on disk
src/webview/ The pages that run inside the webviews
src/status.ts Status bar entry, /api/ps polling
scripts/ Headless runner, reliability probe, self-test

Licence

MIT.

  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft