Skip to content
| Marketplace
Sign in
Visual Studio Code>AI>Boole Code - Offline Coding AgentsNew to Visual Studio Code? Get it now.
Boole Code - Offline Coding Agents

Boole Code - Offline Coding Agents

Preview

Boole Code - Offline Coding Agents

|
21 installs
| (0) | Free
Write code completely offline using your own hardware and the latest models from Boole AI
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

Boole Code - Offline Agents

Write code completely offline using your own hardware and the latest AI models from Boole.

Overview

Boole Code is a VS Code extension that provides an AI-powered terminal coding agent running entirely on your local machine. No API keys, no cloud services, no internet required after initial setup — just you, your hardware, and a capable local language model.

Features

Zero-Setup Model Download

  • Automatic model download on first activation — no manual steps required
  • Downloads Qwen2.5-1.5B-Instruct (~1GB) from Hugging Face
  • Real-time progress tracking with native VS Code progress UI
  • Persistent storage across extension updates

Local Inference

  • Runs models locally using node-llama-cpp (GGUF format)
  • Complete privacy — your code never leaves your machine
  • Works offline after initial model download
  • Supports CUDA/Metal GPU acceleration

Integrated Terminal Agent

  • Launch AI agent in any terminal (automatically installed globally as boole command)
  • Terminal-based UI (TUI) with real-time model display in header
  • Full access to your workspace and tools
  • Works in VS Code's integrated terminal, iTerm, Terminal.app, or any external shell

Flexible Configuration

  • Local backend: On-device inference with auto-downloaded or custom models
  • Remote backend: Optional API fallback (OpenAI-compatible endpoints)
  • Configurable via VS Code settings or environment variables

Command Palette Integration

  • Boole Code: Open Agent Terminal — Launch the AI agent terminal
  • Boole Code: Download Model — Manually download or retry model installation
  • Boole Code: Show Model Status — Check model download status and configuration
  • Boole Code: Restart Agent — Clean restart of the agent terminal
  • Boole Code: Open Settings — Quick access to extension settings

Installation

From VS Code Marketplace

  1. Open VS Code
  2. Go to Extensions (Cmd+Shift+X / Ctrl+Shift+X)
  3. Search for "Boole Code"
  4. Click Install

First Launch

On first activation, the extension will automatically:

  1. Install the boole command globally (available in any terminal)
  2. Start downloading Qwen2.5-1.5B-Instruct (~1GB)
  3. Show progress in a notification (percentage + MB downloaded)
  4. Complete in 1-10 minutes depending on connection speed

Once complete, open any terminal and type boole to start using the agent immediately.

Usage

Basic Workflow

  1. Install the extension — on first activation, it automatically downloads the default model (~1GB)
  2. Run the agent — either:
    • Open Command Palette (Cmd+Shift+P) → Boole Code: Open Agent Terminal
    • Or open any terminal and type boole (command is installed globally)
  3. The TUI launches showing the active model in the header (e.g., "Model: qwen2.5-1.5b-instruct (local)")
  4. Start chatting — type your request and press Enter
  5. Switch models anytime — type /models to browse and download other models from the catalog

Configuration

Settings (UI)

  1. Open Settings: Cmd+, or Ctrl+,
  2. Search for "Boole Code"
  3. Configure backend, model path, or API settings

Settings (JSON)

{
  // Backend: "local" for on-device, "remote" for API
  "booleCode.backend": "local",
  
  // Optional: custom model path (overrides auto-downloaded model)
  "booleCode.modelPath": "/path/to/your/model.gguf",
  
  // Remote backend settings (only if using remote)
  "booleCode.remote.apiEndpoint": "https://api.openai.com/v1/chat/completions",
  "booleCode.remote.apiKey": "",  // Or use BOOLE_API_KEY env var
  
  // Agent loop configuration
  "booleCode.maxIterations": 10
}

Environment Variables

# Backend selection
export BOOLE_BACKEND=local  # or "remote"

# Local backend
export BOOLE_MODEL_PATH=/path/to/model.gguf

# Remote backend
export BOOLE_API_ENDPOINT=https://api.openai.com/v1/chat/completions
export BOOLE_API_KEY=your-api-key

# Or use OpenAI key as fallback
export OPENAI_API_KEY=your-openai-key

Requirements

For Local Backend (On-Device Inference)

  • Disk Space: ~1.5 GB (1 GB for model + extension overhead)
  • RAM: 4 GB minimum, 8 GB recommended
  • CPU: Modern multi-core processor (ARM64 or x86_64)
  • GPU: Optional but recommended (CUDA for NVIDIA, Metal for Apple Silicon)
  • Network: Required only for initial model download

For Remote Backend (API Mode)

  • Network connection
  • API key for OpenAI or compatible service

Model Details

Default Model: Qwen2.5-1.5B-Instruct (Q4_K_M quantization)

  • Size: 1.07 GB
  • Context Length: 4096 tokens
  • License: Apache 2.0
  • Source: Hugging Face

Using a Custom Model

You can use any GGUF-format model:

  1. Download your preferred model (e.g., from Hugging Face)
  2. Set path in settings: booleCode.modelPath
  3. Or set env var: export BOOLE_MODEL_PATH=/path/to/model.gguf

Recommended models:

  • Qwen2.5-1.5B: Fast, lightweight (default)
  • Qwen2.5-3B: Better quality, moderate speed
  • Llama-3.2-3B: Alternative option

Troubleshooting

Model Download Fails

  • Check network: Ensure https://huggingface.co is accessible
  • Retry: Use Boole Code: Download Model command
  • Manual download: Visit model page, download qwen2.5-1.5b-instruct-q4_k_m.gguf, and place in model storage directory
  • Check status: Run Boole Code: Show Model Status

Agent Not Starting

  • Verify model: Run Boole Code: Show Model Status to check if model is ready
  • Check logs: Open Developer Tools (Help → Toggle Developer Tools) and check Console
  • Restart: Use Boole Code: Restart Agent command

Performance Issues

  • GPU acceleration: Ensure CUDA (NVIDIA) or Metal (Apple) is available and configured
  • Reduce context: Lower contextSize if running out of memory
  • Try smaller model: Use a smaller quantization or model size

Extension Not Activating

  • Check VS Code version: Requires VS Code 1.125.0 or later
  • Reload window: Cmd+R / Ctrl+R
  • Reinstall: Uninstall and reinstall the extension

Storage Location

Models are stored in VS Code's global storage (persists across extension updates):

  • macOS: ~/Library/Application Support/Code/User/globalStorage/boole.boole-code/models/
  • Linux: ~/.config/Code/User/globalStorage/boole.boole-code/models/
  • Windows: %APPDATA%\Code\User\globalStorage\boole.boole-code\models\

Privacy & Security

  • Fully offline: Code never leaves your machine when using local backend
  • No telemetry: Extension does not collect usage data
  • Model source: Models downloaded from trusted Hugging Face CDN
  • Open format: GGUF models are transparent, inspectable weights

Development Status

This extension is under active development. Current features:

  • ✅ Auto-download system
  • ✅ Local model inference (node-llama-cpp)
  • ✅ Terminal integration with TUI
  • ✅ Global CLI installation (boole command)
  • ✅ Model switcher with download progress (/models command)
  • ✅ Active model display in TUI header
  • ✅ Command palette commands
  • ✅ VS Code settings integration

Coming soon:

  • [ ] Native webview UI for agent chat
  • [ ] Inline code suggestions
  • [ ] File-based context management
  • [ ] Multi-model comparison mode

Links

  • Homepage: https://booleinference.com
  • Repository: https://github.com/boole-ai/boole-code
  • Issues: https://github.com/boole-ai/boole-code/issues

License

Proprietary - © 2026 Boole AI

  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft