Skip to content
| Marketplace
Sign in
Visual Studio Code>Programming Languages>LocalAITabNew to Visual Studio Code? Get it now.
LocalAITab

LocalAITab

Netajam

|
1 install
| (0) | Free
On-demand local FIM completion and selection refactoring, backed by Ollama.
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

LocalAITab

On-demand code completion and selection refactoring in VS Code, served by a local Ollama. No telemetry, no account, no cloud.

Suggestions are manual by default — nothing is proposed until you ask.

On Windows/Linux the chords are Ctrl+K A and Ctrl+K R.

Not affiliated with Alibaba or the Qwen team; Qwen models are simply the defaults.

Install

Search for LocalAITab in the Extensions view, or:

code --install-extension Netajam.localaitab

Requires

ollama pull qwen2.5-coder:3b-base       # completion: must be a -base tag (FIM)
ollama pull qwen2.5-coder:7b-instruct   # refactor: must be an instruct model

Together they need about 7 GB, which suits a 16 GB machine. With 32 GB or more, a larger instruct model gives noticeably better refactors and chat: ollama pull qwen3.6:35b-a3b-coding (about 24 GB) and pick it with LocalAITab: Select Model....

The two roles need different kinds of model. A -base tag completes text and cannot follow instructions; an instruct tag follows instructions but replies in prose instead of completing code. Using one where the other belongs is the usual cause of "it returns garbage", so the extension warns about both at startup.

Any installed Ollama model can fill either role. LocalAITab: Select Model... asks which one to change (completion, refactor/chat/agent, or chat only), then lists what Ollama has installed, the models that suit that role first, each marked base or instruct from its tag and the capabilities Ollama reports. Picking one of the wrong kind asks before saving it. The choice is written to the setting where it already lives, so a workspace override keeps winning only if that is where you changed it. The "not installed" warning at startup offers the same picker.

The instruct model is the heavy one, and not every machine has it. When Refactor, Chat, Agent or Search planning is invoked and localAITab.refactorModel is not installed, the extension stops and asks rather than failing mid-request: pull the configured model, or run this once with an installed substitute. The substitute is the first installed entry of localAITab.refactorModelFallbacks, or, if that list yields nothing, the largest installed instruct model. A substitute is never used without that confirmation, and the choice is remembered for the rest of the session.

Use

Action Key
Suggest a completion at the cursor Alt+\
Refactor the selection Cmd+K R, or right-click
Open the chat panel Cmd+K Q
Toggle on/off click the status bar item

LocalAITab: Switch Between Manual and Automatic turns on as-you-type suggestions if you ever want them.

Refactoring offers presets (simplify, add types, document, extract functions, handle errors, make idiomatic) plus a free-text instruction, then opens a before/after diff. Nothing is written until you accept it:

In the diff
Accept the $(check) button in the title bar, or Cmd+Enter
Reject the $(close) button, Esc, or just close the tab

An applied refactor is a single undo step, so Cmd+Z puts it back.

If something produces no result, run LocalAITab: Diagnose — it checks the endpoint, both models, the active language, selection state, and runs a live completion, printing everything to the LocalAITab output channel.

Chat

Cmd+K Q opens a panel where you choose the context explicitly. Nothing is gathered implicitly, and the header shows exactly what will be sent, per source, in characters, before you send it.

Source
Pinned lines selections kept with Cmd+K L (or Pin Selection to Chat in the editor menu); they stay attached when you select something else, until /unpin
Selected lines the current selection, every one of them when there are several (Alt+drag, Cmd+D)
Current file the whole active file
Open tabs (same language) every open tab matching the active file's language
Open tabs (all languages) every open tab
Current folder code files beside the active file
Chosen files pick any files from the workspace

Open it from the LocalAITab Chat status bar item, the chat icon in the editor title bar, or Cmd+K Q.

Replies stream, Stop cancels mid-answer, and code blocks get Copy and Insert buttons. Anything that would exceed the attachment cap is listed as skipped rather than silently dropped.

Slash commands

Type / in the composer for a list; Tab or Enter completes, arrows move.

Command
/attach [path or glob] attach files; a folder attaches what is under it, a fragment matches file names, several matches open a picker, no argument opens the full picker
/detach [path] remove attached files whose path contains the argument, or all of them
/unpin [path:lines] drop pinned selections matching the argument (chat.ts, chat.ts:40-52), or all of them
/allow [folder] let agent mode read a folder outside the workspace, see below
/find <task> same as Find context
/chat, /agent, /operator switch mode
/new, /history, /clear, /stats same as the buttons
/skills rescan and list the skills found on disk
/<skill> [args] run a skill, see below
/help list the commands

Commands never reach the model. Start a message with // to send a literal leading slash. Files can also be attached from the explorer or an editor tab's right-click menu with LocalAITab: Attach File to Chat.

Skills

Skills use the same SKILL.md format as Claude Code, Codex and other agent harnesses, so one written for them works here unchanged. A skill is a folder:

.localaitab/skills/review-pr/
  SKILL.md          instructions, with name and description up top
  checklist.md      anything else the instructions refer to
---
name: review-pr
description: Review the attached diff for bugs and missing tests. Use when asked to review a change.
argument-hint: "[focus]"
---

1. Read every changed file before commenting.
2. Report bugs first, then missing tests. Focus on $ARGUMENTS if given.

LocalAITab: New Skill writes that skeleton and opens it. Skills are looked for in the folders listed in localAITab.skillPaths: by default only .localaitab/skills in the workspace, then ~/.localaitab/skills. When two share a name the earlier folder wins. A skill folder is picked up the next time the panel comes into view, or on /skills.

Skills written for other harnesses are not read unless you ask for them, so their descriptions don't crowd a small model's prompt. To use one without copying it, reference it from a skills.json beside a skills folder (.localaitab/skills.json, or ~/.localaitab/skills.json for every workspace):

{
  "skills": [
    "~/.agents/skills/tdd",
    { "path": "~/.claude/skills/my-skill", "model": false }
  ]
}

Each entry is the path of a skill folder, relative to the skills.json or from ~. The skill is read where it lives, lazily like any other: only its name and description are listed to the model, and its body is read when used. "model": false leaves it off the model's list entirely, so it runs only when you type /my-skill. A folder's own skill wins over a referenced one of the same name, and a reference that points at no SKILL.md is reported by /skills.

There are two ways to use one:

  • You invoke it. /review-pr auth sends the skill's instructions as your message, in either mode. $ARGUMENTS in the body becomes auth ($ARGUMENTS[0] is the first word); if the body uses neither, the arguments are appended. The transcript shows what you typed; the saved session keeps what was actually sent.
  • The agent picks it. In agent mode the system prompt lists each skill's name and description, one line each, and the model gets a load_skill tool to read a body, or a bundled file, when a request matches. Only descriptions cost tokens until then, which matters with a local model's window.
Frontmatter
name the /command; lowercase, digits, hyphens. Defaults to the folder name
description what it does and when to use it; the agent decides from this alone
argument-hint shown beside the name in the / list
disable-model-invocation: true only you can run it; not listed to the agent
user-invocable: false only the agent can load it; hidden from /

Other fields (allowed-tools, model, ...) are ignored. LocalAITab has no shell tool, so a skill that asks for a bundled script to be run gets it read, not run. A skill named like a built-in command is hidden by the built-in. Skills that fail to parse are listed by /skills and in the output channel.

Files outside the workspace

There are two kinds of access, on purpose.

Attaching is you choosing a file, so it reaches anywhere. /attach takes an absolute or ~/ path to a file or a folder (a folder opens a picker over what is under it), and every file picker has a Browse outside the workspace button in its title bar that opens the system file dialog. With no workspace open, /attach goes straight to that dialog.

Agent mode reads without asking, so it reaches only folders you granted:

/allow ~/notes until the panel closes; no argument opens a folder dialog
localAITab.extraFolders for good, e.g. ["~/notes", "/srv/shared/specs"]

The system prompt lists the granted folders, and read_file, search (through its path argument) and insert_change accept absolute paths inside them. An edit there is still only staged for your review. With no workspace open, the first granted folder stands in for one, so agent mode works on a notes folder without opening it as a project.

localAITab.extraFolders is read from user settings only. If it could be set in a workspace's .vscode/settings.json, any cloned repository could grant the agent your home directory.

Operator mode does not use these grants: its file writes stay inside the workspace, and its shell already reaches anything you approve a command for. The Find context planner still searches the workspace only.

Find context

Type what you are after, press Find context, and the model writes the ripgrep queries instead of you. Results come back grouped by file with a checkbox per hit and per file. Nothing enters the prompt until you tick it and press Attach.

The division of labour is deliberate. Formulating queries -- trying casings, synonyms, narrowing globs -- is tedious and the model is good at it. Deciding which handful of hits actually matters is the part you are better at, and doing it yourself is what keeps irrelevant code out of the context.

The planner is first shown a cheap local sketch of the codebase: its files and the identifiers it really defines, gathered with ripgrep. This matters a lot. Without it the model guesses naming conventions and misses -- asked about a TypeScript project it proposed max_steps, GetModelInfo and json.Unmarshal and returned zero hits. With the sketch it names the actual functions and found the target file in all five test cases, with 2.6 files to sift on average.

Flags are never taken from the model. Only the pattern, a validated glob and a case-sensitivity bit are honoured, and ripgrep is invoked directly rather than through a shell.

Agent mode

The panel has three modes. Chat is one shot: your question plus the context you ticked, nothing else. Agent gives the model its tools and lets it drive a bounded loop:

Tool
search ripgrep over the workspace or a path in it or in a granted folder, returns file:line:text
read_file one file, truncated past agentMaxReadChars; a folder returns its listing
insert_change stages an edit; never writes
load_skill a skill's instructions or bundled files; only offered when skills exist

The read tools run unattended. insert_change cannot write to disk: it stages one change, which arrives in the same diff with the same Apply/Discard buttons the refactor commands use. A search string matching zero or several places is rejected and handed back as a correctable error, because the dangerous failure is not a missed edit but a confidently misplaced one.

localAITab.agentMaxSteps (default 8) is a hard stop, not a hint. Every step is shown live with its arguments, result and timing, and the run reports steps, tokens and wall time so the modes can be compared. A/B stats totals them.

Operator mode

Operator is the unconstrained mode, kept apart from Agent so Agent stays the staged, safe default. Its own loop and prompt live in src/core/harness/operator/. The model is told to carry a task through: change the code, then build or test to check it.

Tool
search, read_file, load_skill as in agent mode
write_file creates or overwrites a workspace file, immediately
edit_file exact, unique search/replace, applied immediately
run_command a shell command in the workspace root, after your approval

Edits skip the diff, but they are confined to the workspace, so git is the undo. Commands can reach past the workspace, so every one goes through a modal that shows the command line. Run all for this request lifts the prompt until the current request ends, and declining tells the model not to retry. Commands run non-interactively through your login shell. operatorCommandTimeout (default 120s) kills the whole process group. Output is clipped to operatorMaxOutputChars, keeping the head and tail. The full output goes to the LocalAITab output channel. localAITab.operatorMaxSteps defaults to 30.

Saved conversations

Conversations are written to disk as you go, one JSONL file per session, at ~/.localaitab/projects/<folder>/<session id>.jsonl, where <folder> is the workspace path with every character but letters and digits turned into - (/Users/me/my_repo → -Users-me-my-repo). The terminal keeps its conversations in the same place, so one started in the panel can be picked up with /resume in localaitab, and the other way round. Reopening the panel picks up where the last one left off; New session starts clean and History lists past conversations to resume. Show files opens the folder. Sessions from earlier versions, in ~/.localaitab/sessions/<workspace>/ or VS Code's own storage, are converted the first time the panel or localaitab starts; one that can't be is left where it was and logged. LOCALAITAB_HOME moves the whole ~/.localaitab folder.

The files follow Claude Code's transcript shape (~/.claude/projects), so a tool that reads Claude Code's transcripts reads localaitab's by looking in ~/.localaitab instead, and can show their title, goal, branch, token use and tool calls. localaitab never writes into ~/.claude. Each line is appended as it happens and never rewritten, so an interrupted write costs at most the last line, a reader only reads what is new, and a transcript stays greppable:

  • user: a question, with what was typed when a /skill expanded it (display), the context sources attached, the mode and the model. One you typed is marked origin: {kind: "human"}.
  • assistant: each model reply, with its tool calls as tool_use blocks and Ollama's token counts as usage.
  • user with a tool_result: what each tool returned, including a command you declined.
  • ai-title: the session's title, its first question.

While it runs, localaitab also keeps ~/.localaitab/sessions/<pid>.json up to date, in the shape of Claude Code's ~/.claude/sessions/<pid>.json, so a monitor can tell a localaitab agent is running, where, and on which session. The terminal writes one per localaitab process; the panel one per VS Code window while it is open (the pid is the window's extension host). It holds:

Field
pid, procStart the process, and when it started as ps -o lstart= prints it in UTC, so a reused pid is not mistaken for this one
sessionId, name the current transcript and its title; they follow /clear, /resume, New session and History
cwd the workspace
status, statusUpdatedAt busy from sending a prompt until the reply is complete, waiting while an operator command waits for your approval, idle otherwise
kind, entrypoint interactive, and localaitab-tui or localaitab-vscode
tmux session:@window.%pane when the terminal runs in tmux
startedAt, updatedAt, version

It is written whole through a temporary file and a rename, so a reader never sees half of one, and removed when localaitab quits or the window closes. A crash leaves it behind naming a pid that is gone, which is how a reader tells.

Nothing is uploaded. LocalAITab: New Chat Session and LocalAITab: Resume Chat Session... are also in the command palette.

Context meter

A bar tracks usage against the model's context window, updating as you toggle sources and as you type, before anything is sent. It counts the system prompt, attached context, conversation so far, and your draft, and adds localAITab.chatMaxTokens on top, since the reply competes for the same window.

The window comes from /api/ps -- what Ollama actually allocated when it loaded the model, which is what governs truncation. That is often far below the model's advertised maximum: 65,536 against a 262,144 ceiling for qwen3.6 here. When the model is not resident the meter falls back to the maximum and says so.

Attached context is capped at localAITab.chatContextPercent of that window (60% by default), leaving the rest for the conversation and the reply. The cap shows as a tick mark on the bar. Set localAITab.chatContextTokens for an absolute number instead. Everything is counted in tokens against the window: an earlier version also reported a percentage of a fixed character budget, which was a different denominator and read as a contradiction -- 90% of one was 12% of the other.

Ollama exposes no tokenizer endpoint, so counts before sending are estimated from a chars-per-token ratio, then corrected against prompt_eval_count after every reply. Measured against the real tokenizer the starting ratio is within 1-3% on source files; it converges within a turn or two, and each reply also reports its exact prompt and reply token counts.

Context

What the model is given, per command:

Completion (Cmd+K A) Refactor / improve
Model localAITab.model (base/FIM) localAITab.refactorModel (instruct)
Own file 3000 chars before, 1500 after the cursor per localAITab.refactorContext
Other files open tabs, same language, 2000 chars only in openTabs scope

localAITab.refactorContext takes none (selection only, fastest), surrounding (a window either side, the default), file (the whole file with the selection marked), or openTabs (that plus other open tabs of the same language).

To override it for a single run without changing the setting, use LocalAITab: Refactor Selection (choose context)...

Terminal

Chat, agent and operator mode also run without VS Code, as localaitab in a terminal:

npm run install-tui           # compile, test, install to ~/.local/bin/localaitab
localaitab                      # a session in the current directory
localaitab "why does add() subtract?"   # one task, then exit
localaitab --chat               # start in chat mode: answers only, no tools
localaitab -c                   # reopen the latest conversation in this folder
localaitab --resume <id>        # reopen that conversation, by its session id
localaitab --operator -C ~/proj # start in operator mode, elsewhere
localaitab --version            # which build is installed

The install is a copy of the build, not a link to the repo, so it keeps working while you edit or switch branches; rerun npm run install-tui to update it, and npm run uninstall-tui to remove it. It needs Node 18 or newer and ~/.local/bin on your PATH (set LOCALAITAB_PREFIX to install elsewhere).

It runs the same conversation as the chat panel. Chat mode streams an answer from what you attach, with no tools, drawn as markdown while it arrives. In agent mode a staged change is shown as a diff and written only when you answer y. In operator mode, edits land directly and each command is shown before it runs (y, a for the rest of the request, or n). Answers are rendered from markdown: headings, lists, tables, and code blocks with light highlighting.

It takes over the window like a full-screen app: the prompt and a footer (mode, model, and the shortcuts that work right now; F1 lists them all) stay pinned to the bottom, and the conversation scrolls above them on its own, with the mouse wheel or PgUp/PgDn. Scrolled back, the rule above the prompt says how far, and sending a message returns to the bottom.

Each message you send and each reply is a point in the conversation. ⌥↑ and ⌥↓ jump to the previous and next point ([ and ] in vim normal mode), and a scroll bar down the right edge shows where you are and a mark per point: ◆ (cyan) for your messages, • (magenta) for replies, so they differ even without colour. Scrolled back, the rule above the prompt names the point in view, such as ◆ your message 2 of 4. Clicking the bar jumps there, straight to the point when you click a dot. A resumed conversation has its points too. Because the app receives the wheel, select text with ⌥ Option+drag in iTerm2 (Shift+drag in most other terminals). On quit the conversation is printed to the terminal, so it stays in the scrollback.

@path Attach a file, or the source files of a folder. A fuzzy finder opens as you type; Tab opens a folder, Enter attaches. @~/ and @/ browse outside the workspace. Attached context is capped at contextMaxChars (48k).
/ Commands and skills, completed as you type: /mode, /model, /allow, /skills, /clear, /help, /quit, /<skill> args. Their argument completes too: /mode offers chat, agent and operator, /model the models Ollama has installed, /allow folders on disk. Enter on an argument sends.
→ Accept the dimmed suggestion after the cursor: the rest of the highlighted completion, or of your latest earlier message that starts the same way (history is kept in ~/.localaitab/history). Alt+→ takes one word.
Shift+Tab Cycle chat → agent → operator. The thread carries across modes.
Esc Vim normal mode: motions, d/c/y operators, counts, u and Ctrl+R. i/a/A/o go back to typing.
Alt+Enter, Ctrl+J, \ Enter New line. Up/Down walk earlier messages.
Ctrl+C Clear the prompt, or stop a running request. Ctrl+D on an empty prompt quits.

Every conversation is saved as you go, one file per session in ~/.localaitab/projects/, grouped by the folder you started in, in Claude Code's transcript shape (see Saved conversations): the same place the chat panel uses, so its conversations are listed too. /resume lists them by their first question (type the space and the list opens, narrowed as you type), and picking one plays it back and makes it the model's thread again. localaitab -c reopens the latest, and localaitab --resume <id> a given one (the id is the transcript's file name, or sessionId in the live file), which is how a monitor reopens a stopped session; an unknown id is an error. /clear starts a new one.

Piped input (echo "/skills" | localaitab) gets plain lines with no footer.

Defaults come from ~/.localaitab/config.json, keyed by the extension's setting names without localAITab. ({"model": "...", "agentMaxSteps": 12}). Set LOCALAITAB_DEBUG=1 to see the log lines the extension sends to its output channel.

The code is laid out along that split: src/core holds the Ollama client and the tool loops, with no vscode import, behind one entry point, src/core/harness/conversation.ts, which keeps the thread and the skills and runs a turn in either mode. src/hosts/app is the extension and src/hosts/tui the terminal; each supplies how files are read and written, how the user approves a command, and its settings.

Settings

All under localAITab.*. Notable ones: trigger (manual/automatic), model, refactorModel, refactorModelFallbacks, refactorPreview, multiline, debounceMs, neighborFiles.

Develop

npm install
npm run compile
npm test                  # unit tests, no editor or model needed
npm run e2e               # live completions against Ollama
node test/refactor-e2e.js # live refactors, checks for leaked prose and indent drift
npm run e2e:continuity    # live check that both modes carry conversation history
npm run install-all       # compile, test, package, install to every VS Code profile
npm run tui               # the terminal harness, straight from out/
npm run install-tui       # compile, test, install localaitab to ~/.local/bin

Press F5 in VS Code for an Extension Development Host.

Notes

Completion requests set raw: true so Ollama does not wrap the FIM prompt in a chat template. Refactor requests send think: false to models that report the thinking capability — without it a reasoning model spends its whole token budget thinking and returns nothing (and takes ~60s rather than ~1s).

  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft