Boole Code - Offline Agents
Write code completely offline using your own hardware and the latest AI models from Boole.
Overview
Boole Code is a VS Code extension that provides an AI-powered terminal coding agent running entirely on your local machine. No API keys, no cloud services, no internet required after initial setup — just you, your hardware, and a capable local language model.
Features
Zero-Setup Model Download
- Automatic model download on first activation — no manual steps required
- Downloads Qwen2.5-1.5B-Instruct (~1GB) from Hugging Face
- Real-time progress tracking with native VS Code progress UI
- Persistent storage across extension updates
Local Inference
- Runs models locally using
node-llama-cpp (GGUF format)
- Complete privacy — your code never leaves your machine
- Works offline after initial model download
- Supports CUDA/Metal GPU acceleration
Integrated Terminal Agent
- Launch AI agent in any terminal (automatically installed globally as
boole command)
- Terminal-based UI (TUI) with real-time model display in header
- Full access to your workspace and tools
- Works in VS Code's integrated terminal, iTerm, Terminal.app, or any external shell
Flexible Configuration
- Local backend: On-device inference with auto-downloaded or custom models
- Remote backend: Optional API fallback (OpenAI-compatible endpoints)
- Configurable via VS Code settings or environment variables
Command Palette Integration
Boole Code: Open Agent Terminal — Launch the AI agent terminal
Boole Code: Download Model — Manually download or retry model installation
Boole Code: Show Model Status — Check model download status and configuration
Boole Code: Restart Agent — Clean restart of the agent terminal
Boole Code: Open Settings — Quick access to extension settings
Installation
From VS Code Marketplace
- Open VS Code
- Go to Extensions (
Cmd+Shift+X / Ctrl+Shift+X)
- Search for "Boole Code"
- Click Install
First Launch
On first activation, the extension will automatically:
- Install the
boole command globally (available in any terminal)
- Start downloading Qwen2.5-1.5B-Instruct (~1GB)
- Show progress in a notification (percentage + MB downloaded)
- Complete in 1-10 minutes depending on connection speed
Once complete, open any terminal and type boole to start using the agent immediately.
Usage
Basic Workflow
- Install the extension — on first activation, it automatically downloads the default model (~1GB)
- Run the agent — either:
- Open Command Palette (
Cmd+Shift+P) → Boole Code: Open Agent Terminal
- Or open any terminal and type
boole (command is installed globally)
- The TUI launches showing the active model in the header (e.g., "Model: qwen2.5-1.5b-instruct (local)")
- Start chatting — type your request and press Enter
- Switch models anytime — type
/models to browse and download other models from the catalog
Configuration
Settings (UI)
- Open Settings:
Cmd+, or Ctrl+,
- Search for "Boole Code"
- Configure backend, model path, or API settings
Settings (JSON)
{
// Backend: "local" for on-device, "remote" for API
"booleCode.backend": "local",
// Optional: custom model path (overrides auto-downloaded model)
"booleCode.modelPath": "/path/to/your/model.gguf",
// Remote backend settings (only if using remote)
"booleCode.remote.apiEndpoint": "https://api.openai.com/v1/chat/completions",
"booleCode.remote.apiKey": "", // Or use BOOLE_API_KEY env var
// Agent loop configuration
"booleCode.maxIterations": 10
}
Environment Variables
# Backend selection
export BOOLE_BACKEND=local # or "remote"
# Local backend
export BOOLE_MODEL_PATH=/path/to/model.gguf
# Remote backend
export BOOLE_API_ENDPOINT=https://api.openai.com/v1/chat/completions
export BOOLE_API_KEY=your-api-key
# Or use OpenAI key as fallback
export OPENAI_API_KEY=your-openai-key
Requirements
For Local Backend (On-Device Inference)
- Disk Space: ~1.5 GB (1 GB for model + extension overhead)
- RAM: 4 GB minimum, 8 GB recommended
- CPU: Modern multi-core processor (ARM64 or x86_64)
- GPU: Optional but recommended (CUDA for NVIDIA, Metal for Apple Silicon)
- Network: Required only for initial model download
For Remote Backend (API Mode)
- Network connection
- API key for OpenAI or compatible service
Model Details
Default Model: Qwen2.5-1.5B-Instruct (Q4_K_M quantization)
- Size: 1.07 GB
- Context Length: 4096 tokens
- License: Apache 2.0
- Source: Hugging Face
Using a Custom Model
You can use any GGUF-format model:
- Download your preferred model (e.g., from Hugging Face)
- Set path in settings:
booleCode.modelPath
- Or set env var:
export BOOLE_MODEL_PATH=/path/to/model.gguf
Recommended models:
- Qwen2.5-1.5B: Fast, lightweight (default)
- Qwen2.5-3B: Better quality, moderate speed
- Llama-3.2-3B: Alternative option
Troubleshooting
Model Download Fails
- Check network: Ensure https://huggingface.co is accessible
- Retry: Use
Boole Code: Download Model command
- Manual download: Visit model page, download
qwen2.5-1.5b-instruct-q4_k_m.gguf, and place in model storage directory
- Check status: Run
Boole Code: Show Model Status
Agent Not Starting
- Verify model: Run
Boole Code: Show Model Status to check if model is ready
- Check logs: Open Developer Tools (
Help → Toggle Developer Tools) and check Console
- Restart: Use
Boole Code: Restart Agent command
- GPU acceleration: Ensure CUDA (NVIDIA) or Metal (Apple) is available and configured
- Reduce context: Lower
contextSize if running out of memory
- Try smaller model: Use a smaller quantization or model size
Extension Not Activating
- Check VS Code version: Requires VS Code 1.125.0 or later
- Reload window:
Cmd+R / Ctrl+R
- Reinstall: Uninstall and reinstall the extension
Storage Location
Models are stored in VS Code's global storage (persists across extension updates):
- macOS:
~/Library/Application Support/Code/User/globalStorage/boole.boole-code/models/
- Linux:
~/.config/Code/User/globalStorage/boole.boole-code/models/
- Windows:
%APPDATA%\Code\User\globalStorage\boole.boole-code\models\
Privacy & Security
- Fully offline: Code never leaves your machine when using local backend
- No telemetry: Extension does not collect usage data
- Model source: Models downloaded from trusted Hugging Face CDN
- Open format: GGUF models are transparent, inspectable weights
Development Status
This extension is under active development. Current features:
- ✅ Auto-download system
- ✅ Local model inference (node-llama-cpp)
- ✅ Terminal integration with TUI
- ✅ Global CLI installation (
boole command)
- ✅ Model switcher with download progress (
/models command)
- ✅ Active model display in TUI header
- ✅ Command palette commands
- ✅ VS Code settings integration
Coming soon:
- [ ] Native webview UI for agent chat
- [ ] Inline code suggestions
- [ ] File-based context management
- [ ] Multi-model comparison mode
Links
License
Proprietary - © 2026 Boole AI