Skip to content
| Marketplace
Sign in
Visual Studio Code>Programming Languages>Evolve AINew to Visual Studio Code? Get it now.
Evolve AI

Evolve AI

Bala Thiyagarajan

|
189 installs
| (0) | Free
Forward-Deployed Engineers Delivery Studio & AI Data Platform: Live Database Introspector, Schema Mapping, dbt Marts, API SDKs & Multi-Cloud Runbooks.
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

Evolve AI — Context-Aware AI Coding Assistant for VS Code

VS Code Marketplace Installs Rating License: MIT

Evolve AI is built by Evolve Mind Solutions Pty Ltd to bring enterprise AI code assistance, autonomous data engineering, and a comprehensive Forward-Deployed Engineers Delivery Studio directly into your editor. It works with Ollama (local/offline), Gemma 4 (Google's multimodal open model), GLM / CodeGeeX (local coding models), Colibri (GLM-5.2 744B running locally), Anthropic Claude, OpenAI-compatible APIs, Google Gemini, GLM (Z.ai), and Hugging Face — so you choose where your code goes.

[!NOTE]

📢 Free Community Edition vs. 💎 Paid Enterprise Edition

You are viewing the Free Community Edition of Evolve AI (MIT License) on the VS Code Marketplace.

The Community Edition provides core AI assistance, local offline LLM execution, and our foundational 4-step Forward-Deployed Engineers delivery studio with limited features.

For corporate data teams, systems integrators, and regulated banking/defense enclaves, we provide the Paid Enterprise Edition — available as both a Zero-Installation Standalone Desktop Application (.exe) and an Enterprise VS Code Suite.

🔗 Explore & Buy Enterprise Edition • Download Standalone Desktop App • Contact Us for Licensing & Demos

📊 Edition Comparison Matrix: Free vs. Paid Enterprise

Capability & Feature Area Free Community Edition (VS Marketplace) Paid Enterprise Edition (Desktop & Studio)
Price & Licensing Free & Open Source (MIT) Commercial Tier (Pro, Standard, Platinum)
Runtime Environment VS Code Extension Only Zero-Install Standalone Desktop App (.exe) + Enterprise VS Code Extension
Air-Gapped & Security Enclave Basic local mode 100% Air-Gapped Enclave, Ed25519 Cryptographic Hardware Machine-Bound Licensing, Offline Patch (.zip) Loader
Telemetry & Outbound Network Zero telemetry Zero external telemetry, strict air-gap compliance guarantee
Core FDE Studio (Phases 1–4) ✅ Included (Ingestion, Schemas, Scaffolder, Runbooks) ✅ Included + Unlimited Client Workspaces
Phase 5: Enterprise Migration Suite ❌ Locked ✅ Oracle & T-SQL to Snowflake/BigQuery SQL Transpiler, Automated Row-Level Security (RLS) Generator, Reverse ETL Sync Workers, Referential Synthetic Data Generator, Mock API Servers
Phase 6: DevOps & Cloud Infra Hub ❌ Locked ✅ Multi-Cloud Terraform Scaffolding, GPU Kubernetes Manifests, Docker Compose, CI/CD Pipelines (GitHub Actions, GitLab CI, Bitbucket, Azure DevOps), VPC Discovery
Data Privacy & Sanitization Standard ✅ Automated PII Masking & Sanitizer, SIEM Event Forwarder (Splunk, Datadog, Sentinel), Great Expectations & Soda Core Drift Gates
Databricks Lakehouse Hub Basic queries ✅ Deep Unity Catalog Table Lineage, Query Cost Analyzer, Impact Analysis Hub
Commercial Support & SLA Community GitHub Issues ✅ Dedicated Enterprise Support, Custom Migration Rules & SLA

Why Evolve AI?

  • Free Community vs. Enterprise Delivery Studio Alignment (v2.20.0) — Clear frontline boundary between the Free Community Edition (4-step Ingestion, API SDK, Pre-Flight, and Runbook factory) and the Paid Enterprise Edition (standalone .exe, 8-step enterprise studio, Oracle/T-SQL transpilation, and site licensing).
  • Display Scale, Zoom Manager & High-DPI Readability (v2.20.0) — Native GPU-accelerated subpixel zoom, High-DPI auto-detection (115%/125%), interactive zoom widget with slider and presets, reading comfort modes, and elevated theme luminance.
  • Updates & Version Management Hub (v2.20.0) — Embedded Settings panel for live GitHub release checking and 1-click air-gapped offline patch bundle management.
  • AI Schema Copilot & Dimensional Mart Discovery (v2.19.1) — domain-agnostic AI schema standardizer with automated PII masking, multi-table foreign-key graph traversal, and natural-language prompt-to-mart modeling.
  • Databricks Lakehouse & Delta Studio Hub (v2.19) — interactive setup wizard (aiForge.databricks.connect), live connection testing with hardware-encrypted secret storage, and dedicated PySpark, Delta Lake, Unity Catalog, and Delta Live Tables (DLT) tools.
  • Multi-Cloud & Pilot Deployment Hub (v2.19) — unified multi-cloud scaffolding for AWS, GCP, Azure, Firebase, and Kubernetes/Docker.
  • Forward-Deployed Engineers Delivery Studio (Beta) (v2.18) — end-to-end frontline delivery toolkit:
    • 1-Click Live Database Introspector: non-destructive INFORMATION_SCHEMA metadata extraction for PostgreSQL, Supabase, Snowflake, BigQuery, MySQL, SQL Server, and SQLite with auto-detected .env connection strings.
    • Enterprise Hardware Encryption (vscode.SecretStorage): DPAPI/Keychain vault security for zero plaintext credentials in code.
    • Cross-Model Dimensional Mart Join Builder: visual SQL join builder emitting production dbt data mart models.
    • Dedicated Git & Bitbucket / GitHub Remote & Terminal Hub: interactive live streaming terminal (🚀 FDE: Git Hub), SSH diagnostics, and remote branch management.
    • Multi-Cloud Connection & Auth Hub (☁️ Cloud Hub & Connect): live status cards and 1-click terminal browser SSO authentication for Google Cloud, AWS, Azure, and Docker.
    • Client API Studio & Handoff Factory: typed TypeScript/Python SDKs, deterministic pre-flight health auditor, and 5-artifact client handoff package with Mermaid architecture diagrams. → docs/FDE_PLAYBOOK.md
  • Free & private — runs fully offline with Ollama or Gemma 4. Your code never leaves your machine.
  • Auto-detecting stack plugins — 18 plugins that activate automatically based on your project: Databricks, Terraform, Docker, Kubernetes, Django, FastAPI, dbt, Airflow, PyTorch, Data Analysis & Reporting, Code Converter, and more.
  • Any AI provider — bring your own model or API key. Switch between local and cloud in one click.
  • Deep context — understands your active file, related files, diagnostics, git state, and cloud platform connections.
  • Connect to GitHub or Bitbucket in one click (v2.0) — wizard handles git install, identity, init/clone, auth (PAT / SSH / VS Code GitHub auth / gh CLI), and verifies the connection. → docs/GIT_CONNECT.md
  • Author CI/CD pipelines with AI (v2.1) — auto-detecting plugin for GitHub Actions / GitLab CI / Jenkins / CircleCI / Azure / Bitbucket, plus a setup wizard that generates a stack-tailored starter pipeline. → docs/CICD.md
  • Convert code between 26 languages (v2.12) — side by side with the original plus a translation fidelity report. → docs/CODE_CONVERSION.md
  • Build data reports instead of receiving them (v2.13) — compose block-by-block charts directly from real CSV/Excel/Parquet data. → docs/DATA_ANALYSIS.md
  • Zero-AI Offline Suite & Air-Gapped Mode (v2.16) — 100% offline developer & data engineering tools: multi-dialect SQL formatter (Databricks, Snowflake, BigQuery, Postgres, DuckDB), dataset profiler & dbt test generator, Cron & Regex workbench, Terraform/Docker security linters, AST codemods, and a strict air-gap network blocker. → docs/OFFLINE_SUITE.md
  • Local (Community) vs. Paid (Enterprise) Management Guide — comprehensive architectural specification on how the Free Community and Commercial Enterprise editions are managed, cryptographically licensed (Ed25519), and tested locally. → docs/LOCAL_VS_PAID_MANAGEMENT.md

Also works in Cursor, VSCodium, and other VS Code forks.


Built for Forward-Deployed Engineers (FDE)

If you are a Forward-Deployed Engineer (FDE) or solutions architect delivering pilots, custom client data integrations, and frontline production deployments, Evolve AI provides an integrated 4-phase delivery system:

  • 1. Ingest, Live Introspect & Mart Join Builder — Connect directly to live databases (Postgres, Snowflake, BigQuery, MySQL, SQL Server, SQLite) or ingest dirty CSVs/dumps. Auto-maps foreign columns with semantic casting rules to generate production dbt staging models, PySpark transformation pipelines, and dimensional data marts.
  • 2. Resilient Client API Studio — Ingest client OpenAPI, Swagger, or cURL specs and scaffold fault-tolerant TypeScript or Python client SDKs equipped with exponential backoff, jitter, Retry-After header parsing, and rate-limiting circuit breakers.
  • 3. Pre-Flight Health Auditor & Pilot Deployment — 100% deterministic local audit detecting dangling backup files (*.bak, *.tmp), secret leakage, and .env parity. Scaffolds multi-target Firebase Hosting, Google Cloud Run / Docker / Kubernetes / Terraform configs, and cross-platform deployment scripts (deploy.sh / deploy.ps1).
  • 4. Client IT Handoff & Runbook Factory — Auto-generates client-ready documentation in docs/: ARCHITECTURE.md with rendered Mermaid system & sequence diagrams, DEPLOYMENT_RUNBOOK.md with rollback/disaster-recovery instructions, DATA_DICTIONARY.md, ENVIRONMENT_CATALOG.md, and CLIENT_HANDOFF_COMPLETE.md.
  • Git Hub & Multi-Cloud Connection — Integrated interactive terminal streaming for Bitbucket & GitHub, and live status / SSO authentication for Google Cloud, AWS, Azure, and Docker.
  • 14-Day Delivery Roadmap & Playbook — Built-in visual guide and step-by-step operational playbook for running client pilots. → docs/FDE_PLAYBOOK.md

100% Out-of-the-Box Built-in Deterministic Engines (Zero AI Required!)

Every single generator, mapper, and scaffolder in the Forward-Deployed Engineers Delivery Studio is powered by built-in deterministic engines.

You do NOT need any external AI API keys, internet connection, or token subscriptions to execute all Studio activities:

Studio Activity How It Works Under the Hood Needs External AI?
Phase 1: Live DB Introspect & Schema Mapper Native INFORMATION_SCHEMA queries + heuristic Levenshtein/synonym field matcher + dbt/PySpark AST compiler. ❌ No (100% Built-in)
Phase 1: Dimensional Mart Join Builder Deterministic SQL compiler generating CTEs, join predicates, and dimensional aggregate metrics. ❌ No (100% Built-in)
Phase 2: Client API Studio Deterministic OpenAPI & cURL parser + typed TypeScript/Python resilient SDK template engine. ❌ No (100% Built-in)
Phase 3: Pre-Flight Health Auditor 100% offline static analyzer (scans for secret leaks, dangling .bak/.tmp files, .env parity). ❌ No (100% Built-in)
Phase 3: Multi-Cloud Scaffolder Deterministic generator producing Terraform HCL, Kubernetes manifests, Docker Compose, & Firebase configs. ❌ No (100% Built-in)
Phase 4: Runbook & Diagram Factory Deterministic Markdown compiler with dynamic Mermaid.js system architecture and sequence diagrams. ❌ No (100% Built-in)
Git & Multi-Cloud Auth Hub Direct OS-level terminal execution via official git, ssh, gcloud, aws, az, docker CLIs. ❌ No (100% Built-in)

[!TIP] Why this matters for Enterprise FDEs:
When working inside client banking, healthcare, or defense VPCs where sending customer data to external LLMs is strictly prohibited, the studio runs in 100% Air-Gapped Mode with zero external network leakage.


Optional AI Augmentation (Inbuilt Local AI or Cloud LLMs)

When you need AI for natural language prompts, custom business logic, or automated refactoring in the sidebar chat, Evolve AI provides full flexibility:

  • Local Inbuilt AI (100% Free & Private): Google Gemma 4 (multimodal via Ollama), Qwen 2.5 Coder, DeepSeek, or Colibri/GLM-5.2 running on your laptop.
  • Enterprise Cloud LLMs (Bring-Your-Own-Key): Anthropic Claude (Claude 3.7 Sonnet), Google Gemini (Gemini 2.0 / 1.5 Pro), OpenAI (GPT-4o), Groq, and custom endpoints.

Get Started in 60 Seconds

# 1. Install Ollama (free, local AI)
# Download from https://ollama.com — or on Linux:
curl -fsSL https://ollama.ai/install.sh | sh

# 2. Pull a model (pick one)
ollama pull gemma4:e4b        # Google Gemma 4 — multimodal, 128K context (recommended)
ollama pull qwen2.5-coder:7b  # Qwen — optimized for code

# 3. Install the extension and start coding
# Ctrl+Shift+A to open chat — Evolve AI detects Ollama automatically

No API key. No account. No data leaving your machine. That's it.

Connecting your repo to GitHub or Bitbucket (new in v2.0)

Open the command palette and run Evolve AI: Connect Git Remote (Wizard), or click the · not connected hint in the status bar. The wizard:

  1. Installs Git if missing (Win / mac / linux instructions)
  2. Sets your user.name / user.email if not already configured
  3. Initialises / clones / links a repo
  4. Authenticates: VS Code GitHub (built-in, recommended for github.com) · PAT · SSH ed25519 · gh auth login — pick what fits you
  5. Optionally creates a new repo on GitHub or Bitbucket via API
  6. Verifies the connection with git ls-remote origin

Tokens are stored in vscode.SecretStorage — they never touch settings.json. Existing credential.helper is never overwritten. Full guide: docs/GIT_CONNECT.md.

Setting up CI/CD (new in v2.1)

Open the command palette and run Evolve AI: CI/CD Setup Wizard, or click Optimize Pipeline on a CodeLens above any existing pipeline. The wizard:

  1. Detects your stack (language, package manager, test framework, git host)
  2. Asks: which CI platform · what kind of pipeline (test only / + deploy / + container build) · which deploy target (npm / PyPI / Docker / AWS ECS / GCP Cloud Run / Azure / k8s)
  3. Generates a starter pipeline tailored to your stack — pinned actions, OIDC where the platform supports it, dependency caching by lockfile hash, concurrency control, timeouts, least-privilege permissions
  4. Writes it to the right path and opens it for review

For existing pipelines, the CI/CD plugin auto-activates on detection of .github/workflows/*.yml, .gitlab-ci.yml, Jenkinsfile, .circleci/config.yml, azure-pipelines.yml, or bitbucket-pipelines.yml. It contributes platform-aware best practices into every AI prompt and adds CodeLens / lightbulb actions like Pin actions to commit SHA, Replace long-lived secrets with OIDC, Convert to matrix strategy. Full guide: docs/CICD.md.


How Does It Compare?

Feature Evolve AI GitHub Copilot Continue.dev Cody
Free local AI (Ollama, Gemma 4) Yes No Yes No
Auto-detecting stack plugins (18) Yes No No No
Code conversion between languages (26, with fidelity report) Yes No No No
Cloud platform integration (AWS, GCP, Azure, Databricks) Yes No No No
Multimodal (images via Gemma 4) Yes No Partial No
Multiple AI providers 9 1 Multiple 1
Offline mode Yes No No No
Open source MIT No Apache 2.0 Apache 2.0
Price Free $10-19/mo Free Free tier

Features

Multi-Provider AI Support

Provider Privacy Setup
Ollama (local) Code never leaves your machine Free, runs locally
Gemma 4 (local) Code never leaves your machine Free, guided setup via Ollama
GLM / CodeGeeX (local) Code never leaves your machine Free, coding model via Ollama (offline)
Colibri — GLM-5.2 (local) Code never leaves your machine Free, no GPU required. Needs ~372 GB disk; hardware check included
Anthropic Claude Cloud API API key required
OpenAI / Compatible Cloud API (Groq, Mistral, Together AI, LM Studio) API key required
Google Gemini Cloud API API key required
GLM (Z.ai) Cloud API API key required — flagship glm-4.6 / glm-4.5
Hugging Face Cloud API API key required
Offline mode Fully offline, pattern-based No setup needed

AI Chat — Sidebar or Editor Tab

  • Two ways to chat. Open in the sidebar (Ctrl+Shift+A), or click the Evolve AI icon in any file's editor title bar to open the chat as a tab to the right of your code — Claude Code-style. Both views share the same conversation in real time.
  • Inline mode pill — pick Chat (ask), Edit (modify the active file), or Create (generate new files) directly above the input box.
  • Inline model pill — switch models within the active provider (e.g., between your installed Ollama models, or between claude-sonnet-4-6 and claude-opus-4-7) without leaving the chat. A More providers… item handles cross-provider changes.
  • Streaming responses with full project context.
  • Understands your active file, related files, diagnostics, and git state.
  • Context budget system ensures efficient token usage.

Smart Code Actions

  • CodeLens hints above every function: Explain | Tests | Refactor
  • Lightbulb actions: "Fix with AI" on any diagnostic
  • Right-click menu: Explain, refactor, fix, document, generate tests
  • Keyboard shortcuts: Quick access to common actions

18 Core Commands

  • Open AI Chat — sidebar (Ctrl+Shift+A)
  • Open AI Chat — editor tab (top-right icon in any file)
  • Generate Code from Description (Ctrl+Alt+G)
  • Fix Current Errors (Ctrl+Alt+F)
  • Explain Selected Code (Ctrl+Alt+E)
  • Generate Commit Message (Ctrl+Alt+M)
  • Refactor Selection, Add Documentation, Generate Tests, Apply Folder Transforms
  • Explain Changes, Generate PR Description, Build Framework, Run & Auto-Fix
  • Switch Provider, Setup Ollama, Gemma 4 Info & Tips, What's New

18 Auto-Detecting Plugins

Plugins activate automatically based on your workspace files. No configuration required.

Plugin Detects Highlights
Databricks databricks.yml, PySpark imports 10+ commands, live workspace API: clusters, jobs, notebooks, Unity Catalog, SQL warehouse, DLT pipelines
AWS serverless.yml, template.yaml, AWS SDK 28+ commands, live API: Lambda, Glue, S3, CloudFormation, Step Functions, DynamoDB, IAM, SAM, CDK
Google Cloud app.yaml, GCP SDK imports 26+ commands, live API: Cloud Functions, Cloud Run, BigQuery, GCS, Pub/Sub, Firestore, Cloud Build
Azure host.json, Azure SDK imports 28+ commands, live API: Functions, Logic Apps, Cosmos DB, Storage, DevOps Pipelines, Bicep, Log Analytics
dbt dbt_project.yml 6 commands: explain models, tests, incremental, docs, optimize
Apache Airflow airflow.cfg, DAG files 6 commands: explain DAGs, TaskFlow, sensors, retry, monitoring
pytest pytest.ini, conftest.py 6 commands: generate tests, fixtures, parametrize, coverage
FastAPI FastAPI imports 6 commands: endpoints, validation, CRUD, auth, tests
Django manage.py 6 commands: models, serializers, admin, views, URLs, tests
Terraform *.tf files 6 commands: explain, variables, tags, modules, outputs, security
Kubernetes K8s YAML manifests 6 commands: explain, probes, resources, security, manifests, network
Docker Dockerfile 6 commands: explain, optimize, healthcheck, security, compose
Jupyter *.ipynb files 5 commands: explain, document, clean, convert, generate
PyTorch PyTorch imports 6 commands: models, training loops, checkpoints, mixed precision
Security Always active 3 commands: scan file, scan workspace, fix findings
Git Always active 4 commands: blame, changelog, commit messages, PR templates
Data Analysis & Reporting .csv, .tsv, .json, .xlsx, .parquet 12 commands: analyze & report, insights in chat, HTML report, notebook/script, profile, analyze from database/cloud source, refine report, report theme, save + run report templates, create + run data pipeline — reports you author block by block, edit in place, and re-run on new data
Code Converter Any recognised source file 7 commands: convert selection / file / folder, choose conversion model, reopen review, check it parses. 26 languages, side-by-side review, fidelity report

Data Analysis & Reporting

Give Evolve AI a data file and an instruction — get insights and a report, PowerBI-style, without leaving your editor.

  • Right-click a data file (.csv / .tsv / .json / .xlsx / .parquet) → Analyze Data & Report, or use the command palette.
  • Deliverables: Insights in chat (Gemini-style narrative analysis inline, with follow-up questions), a self-contained HTML report (KPI tiles, charts, AI insights), a reproducible pandas/plotly notebook or script, or a profiling summary.
  • Not just local files — Analyze Data from Database or Cloud Source pulls a sample from BigQuery, Databricks SQL, Cosmos DB, Azure Log Analytics, DynamoDB, or S3 / GCS / Azure Blob objects, reusing your existing connected-plugin credentials. For Postgres / MySQL / SQLite / Snowflake / SQL Server, it generates a pandas.read_sql script you run with your own connection string (DB_URL env var — no passwords stored).
  • Size-adaptive: small files are analysed directly by the AI; for large files it generates a script that reads the full dataset locally — your data never leaves your machine. If a sample would go to a cloud provider, you're asked first.
  • Output lands next to your data (sales.csv → sales-report.html), as one self-contained file with no network dependencies.
  • Declarative pipelines — define a repeatable analysis once in evolve-data-pipeline.json (steps = source + analysis) and run them all with Run Data Pipeline. A versioned, backend-free "workflow" you own in your repo.

Reports you author, not reports you receive (v2.13)

Most AI report features generate at you: one prompt, one document, take it or run it again. This one is a builder.

  • The design isn't the model's job. Evolve AI owns the stylesheet, the chart styling and the report runtime, and stamps them into the finished document — so every report gets light and dark following the reader's OS, sortable and filterable tables, a consistent chart palette, a responsive layout, and print/PDF styles that never split a card across a page. The look holds no matter which model produced it, local or cloud.
  • Build it block by block. A report is an ordered list of typed blocks — KPI tiles, charts, tables, your own text, insights, recommendations, data quality, relationships. A chart block carries its measure, dimension, aggregation, chart type, top-N and sort; the pickers list your actual columns with their inferred types, because the file has already been read. Anything left on auto is still chosen for you, so you only pin what you care about.
  • Edit the rendered report directly. Hover any section to move, duplicate or delete it, drag it by its grip, or double-click any text — heading, caption, KPI label, table cell — to fix it in place. These are direct edits: instant, no AI call, and incapable of disturbing the sections you didn't touch.
  • Refine one block, not the document. Select a section and describe the change; only that card is sent, so the model structurally cannot rewrite the rest. A block round-trip costs a fraction of a whole-document one.
  • Five report formats — executive summary, deep-dive analysis, data-quality audit, trend/time-series, segment comparison — each changing the sections, chart budget, tone and framing, not just the cosmetics.
  • Prepare the data first. Row filters, derived columns, excluded columns, drop-duplicates and row caps run for real — as generated pandas in the script path — not as a sentence in a prompt. The report discloses the active filters, so a filtered figure is never read as a total.
  • Brand it once — evolve-report-theme.json holds your colours, palette, logo, footer and default report shape, applied to every report from then on.
  • Save it as a template and re-run the same report shape against next month's data. The outline is read back out of the rendered HTML, so a template captures what you actually arranged on screen.
  • Export to PDF via the print styles built into every report.

Full guide: docs/DATA_ANALYSIS.md

Code Conversion (v2.12)

Port code to another language and actually be able to trust the result.

Asking any AI to "convert this to Go" gives you something that looks right — and the dangerous part is what disappears quietly: a retry loop, a Decimal that became a float64, a library call replaced by a function that doesn't exist. The Code Convertor mode is built around that failure.

  • 26 languages, any pairing — Python · TypeScript · JavaScript · Java · C# · Go · Rust · C++ · C · Kotlin · Swift · Scala · Ruby · PHP · Dart · Elixir · R · SQL · Bash · PowerShell · Lua · Perl · VBA · COBOL · MATLAB · SAS. Every one works as source and target, so legacy → modern is a first-class path.
  • A fidelity report, every time — what mapped 1:1, what was approximated, what needs a human, and how each source dependency was mapped. A conversion that comes back without a report is flagged as suspicious rather than quietly trusted.
  • Review before anything is written — the result opens beside the original, tabs per file, report on its own tab. Refine it in plain language ("return errors instead of panicking, drop the third-party HTTP client") and only the affected files are re-emitted. Every round is undoable. A conversion you reject leaves no litter.
  • It checks its own work — runs the target language's own parser over the output when that toolchain is installed (gofmt, javac, node --check, php -l, and a dozen more), and failures can be fed straight back for a repair round.
  • Choose the model for the job — conversion gets its own model choice, separate from what chat runs on. The picker shows each installed model's real context window, read from Ollama itself rather than guessed from the name.
  • Too big? It tells you before it fails — every job is sized against the chosen model first. If it won't fit you get a straight answer ("Too big for one pass — ~28k tokens needed, 8k available") plus the fix: a model you already have that's big enough, or the ollama pull for one that would be. Oversized files are sliced at top-level declarations — never mid-function — converted part by part and stitched back together, with the report naming the seams.
  • How faithful, and how free with dependencies — idiomatic, line-by-line (diffable against the source, for conversions someone has to sign off), or modernise; and stdlib only, well-known packages, or one-for-one with the source's libraries.

Right-click a file or folder, use the lightbulb on any selection, or pick Code Convertor from the mode menu in chat. Full guide: docs/CODE_CONVERSION.md

Cloud Platform Integration

The Databricks, AWS, Google Cloud, and Azure plugins go beyond code assistance. They connect to your actual cloud accounts to:

  • Manage resources — list and inspect Lambda functions, Cloud Run services, Azure Functions, Databricks clusters
  • Execute queries — run SQL on BigQuery, Cosmos DB, Databricks SQL warehouses
  • Browse storage — navigate S3 buckets, GCS objects, Azure Blob containers, Unity Catalog
  • Trigger and monitor jobs — run Glue jobs, Databricks workflows, Step Functions
  • AI-powered diagnostics — analyze failed job runs with AI explanations and fix suggestions
  • Deploy from VS Code — deploy notebooks, upload to S3/GCS/Azure Storage, manage DLT pipelines

Secure by Design

  • API keys stored in VS Code's encrypted SecretStorage — never in plaintext settings
  • Cloud credentials use standard provider SDKs and authentication flows
  • All file edits go through VS Code's undo stack
  • Diff preview before applying AI-generated changes
  • Context budget caps prevent excessive token usage
  • Workspace Trust enforced — in untrusted workspaces, workspace-level overrides of provider host URLs (ollamaHost, openaiBaseUrl, huggingfaceBaseUrl) are ignored so a malicious .vscode/settings.json can't redirect your chat to an attacker-controlled server
  • Remote-host warning — if a provider URL isn't loopback/private, a one-time toast tells you your code is leaving your machine
  • Image uploads validated (10 MB cap, PNG/JPEG/WEBP/GIF only)
  • Ollama minimum 0.12.4 — the smart-setup wizard prompts for upgrades to close known Ollama CVEs

Quick Start

  1. Install the extension from the VS Code Marketplace
  2. Choose your AI provider:
    • For local/private: Install Ollama, pull a model (ollama pull qwen2.5-coder:7b), and you're ready
    • For cloud AI: Run Evolve AI: Switch AI Provider from the command palette, select your provider, and enter your API key when prompted
  3. Start coding: Open the AI Chat sidebar (Ctrl+Shift+A) or use any command from the command palette
  4. Cloud plugins activate automatically when they detect relevant files in your workspace

AI Providers

Ollama (local, recommended)

Run AI completely on your machine — no API key, no cost, no data leaving your network.

# Install Ollama: https://ollama.ai
ollama pull qwen2.5-coder:7b

Set aiForge.provider to ollama (or leave on auto — it detects Ollama automatically).

Also compatible with LM Studio, llama.cpp, and Jan — point aiForge.ollamaHost at your server.

Gemma 4 (local, multimodal)

Google's latest open model with text, image, and audio understanding. Runs locally and privately via Ollama. Apache 2.0 licensed.

One-click setup — run Switch AI Provider → select Gemma 4. Evolve AI:

  1. Asks consent to inspect your system (RAM, GPU, disk, Ollama version) — no data leaves your machine
  2. Recommends the variant that fits your hardware
  3. Shows a single "Install Everything" button that handles Ollama install/upgrade + model download + config
  4. Reports live download progress (MB/total) right in the notification

If your system can't run any variant, you get actionable alternatives instead of a dead end (cloud providers or offline mode).

Or set up manually:

# Install Ollama: https://ollama.com
ollama pull gemma4:e4b    # Recommended for most users (~9.6GB)

Choose your variant in aiForge.gemma4Model:

Variant Params Size Best for
gemma4:e4b 4.5B ~9.6GB Balanced speed & quality (recommended)
gemma4:e2b 2.3B ~7.2GB Fast, lightweight tasks
gemma4:26b 25.2B MoE ~18GB High-quality reasoning (32GB+ RAM)
gemma4:31b 30.7B ~20GB Maximum quality (32GB+ RAM, GPU)

Anthropic Claude

  1. Get an API key from console.anthropic.com
  2. Run command: Switch AI Provider -> select Anthropic
  3. Enter your API key when prompted (stored in VS Code SecretStorage)

OpenAI / Compatible

Works with OpenAI, Groq, Mistral, Together AI, LiteLLM, and any OpenAI-compatible endpoint.

  1. Set aiForge.openaiBaseUrl to your endpoint (default: https://api.openai.com/v1)
  2. Set aiForge.openaiModel to your model name
  3. Run Switch AI Provider -> select OpenAI -> enter API key

Google Gemini

Use Google's Gemini models via the official OpenAI-compatible endpoint.

  1. Get an API key from aistudio.google.com/apikey
  2. Set aiForge.geminiModel (default: gemini-2.5-flash; also gemini-2.5-pro, gemini-2.0-flash)
  3. Run Switch AI Provider -> select Google Gemini -> enter API key

GLM (local, offline)

Run a GLM / CodeGeeX coding model fully offline via Ollama — no API key, no data leaves your machine. Default is codegeex4-all-9b (a coding model built on GLM-4-9B, ~5.5GB, 128K context).

  1. Install Ollama
  2. Run Switch AI Provider -> select GLM (local) -> pick a model (codegeex4-all-9b, glm4:9b, or glm4) -> it offers to download it
  3. Runs locally from then on

Note: the 355B+ GLM-4.5 / GLM-4.6 flagships are too large to run on a normal machine — those are cloud-only (see below).

Colibri — GLM-5.2 (744B) on your own machine

Colibri is a dependency-free C engine that runs frontier Mixture-of-Experts models locally by streaming experts from disk instead of holding them in memory. That makes GLM-5.2 (744B total, ~40B active per token) runnable without a datacenter GPU.

Be aware of the trade-off before you start: the weights are ~372 GB on disk (a hard requirement), and decode speed depends heavily on how much fast memory you have:

Hardware Realistic speed What that feels like
25 GB RAM, CPU only 0.05–0.1 tok/s ~1.5–3 hours for a 500-token answer
Single 24 GB GPU ~1 tok/s ~8 minutes
128 GB desktop, warm ~1.8 tok/s ~5 minutes
Multi-GPU residency 5.8–6.8 tok/s Usable interactively

Evolve AI runs a hardware check when you select Colibri and tells you honestly which tier you land in — offering faster alternatives when the answer is "this will crawl."

  1. Install and build Colibri, then download the GLM-5.2 INT4 weights (setup guide)
  2. Start the server: coli serve
  3. Run Switch AI Provider -> select Colibri — GLM-5.2 (local)
  4. Confirm the endpoint (default http://localhost:8080/v1) and pick a model

Evolve AI does not install or launch Colibri — it only connects to the endpoint you point it at. Colibri also serves kimi-k3, inkling, and olmoe, selectable via aiForge.colibriModel.

Want GLM quality without the 372 GB? An Unsloth GGUF quant via Ollama is far smaller and much faster on the same hardware, and the GLM (Z.ai) cloud provider below needs no local resources at all.

GLM (Z.ai, cloud)

Use Zhipu / Z.ai's flagship GLM models (glm-4.6, glm-4.5) via their OpenAI-compatible cloud API.

  1. Get an API key from z.ai
  2. Run Switch AI Provider -> select GLM (Z.ai) -> enter API key -> pick a model
  3. Set aiForge.zaiModel (default: glm-4.6)

HuggingFace Inference API

Access thousands of open models via the HuggingFace Inference API.

  1. Get a token from huggingface.co/settings/tokens
  2. Set aiForge.huggingfaceModel (default: Qwen/Qwen2.5-Coder-32B-Instruct)
  3. Run Switch AI Provider -> select HuggingFace -> enter token

Built-in Offline AI

Pattern-based code analysis — works instantly with no setup, no network, no LLM.


Settings

Setting Default Description
aiForge.provider auto AI provider: auto, ollama, gemma4, glm, colibri, anthropic, openai, gemini, zai, huggingface, offline
aiForge.ollamaHost http://localhost:11434 Ollama server URL (also LM Studio, llama.cpp)
aiForge.ollamaModel qwen2.5-coder:7b Ollama model name
aiForge.gemma4Model gemma4:e4b Gemma 4 variant: gemma4:e2b, gemma4:e4b, gemma4:26b, gemma4:31b
aiForge.gemma4ThinkingMode false Enable chain-of-thought reasoning (better results, slower)
aiForge.allowHardwareDetection true Allow inspecting system specs (RAM, GPU, disk) to recommend best Gemma 4 variant
aiForge.allowAutoInstall false When true, skip the per-install confirmation dialog. When false (default), the wizard asks before downloading the Ollama installer
aiForge.openaiBaseUrl https://api.openai.com/v1 OpenAI-compatible endpoint
aiForge.openaiModel gpt-4o OpenAI model name
aiForge.anthropicModel claude-sonnet-4-6 Anthropic model name
aiForge.geminiModel gemini-2.5-flash Google Gemini model name
aiForge.glmModel codegeex4-all-9b Local GLM / CodeGeeX model tag (runs offline via Ollama)
aiForge.colibriBaseUrl http://localhost:8080/v1 Colibri server URL (start it yourself with coli serve)
aiForge.colibriModel glm-5.2 Model served by Colibri (glm-5.2, kimi-k3, inkling, olmoe)
aiForge.zaiModel glm-4.6 GLM (Z.ai) cloud model name
aiForge.huggingfaceModel Qwen/Qwen2.5-Coder-32B-Instruct Hugging Face model ID
aiForge.codeLensEnabled true Show CodeLens hints above functions
aiForge.contextBudgetChars 24000 Total character cap for AI context
aiForge.maxContextFiles 5 Max related files in context
aiForge.requestTimeoutMs 0 Idle timeout per request (resets on each streamed chunk). 0 = auto: 5 min for local (Ollama/Gemma/HF), 2 min for cloud. Raise it if a slow local model cold-starts past the limit
aiForge.disabledPlugins [] Plugin IDs to disable (e.g., ["databricks", "aws"])

Code Conversion

Setting Default Description
aiForge.convert.defaultTarget "" Language pre-selected in the converter. Blank = choose every time
aiForge.convert.fidelity idiomatic idiomatic, literal (diffable against the source), or modernise
aiForge.convert.dependencies popular stdlib (nothing third-party), popular, or mirror (one-for-one with the source's libraries)
aiForge.convert.includeTests false Also generate tests in the target's usual framework
aiForge.convert.keepComments true Carry comments across, rewritten in the target's doc style
aiForge.convert.emitManifest true Also emit go.mod / package.json / requirements.txt etc.
aiForge.convert.outputFolder converted Where multi-file conversions land, under a per-language subfolder. Single files go beside the original
aiForge.convert.maxFiles 20 Cap on files queued from one folder
aiForge.convert.maxCharsPerBatch 60000 Upper bound per request. The converter also derives a budget from the chosen model's real context window and uses whichever is smaller

The conversion model is not a setting — it's a per-session choice (Evolve AI: Choose AI Model for Code Conversion), so a model picked for one big port doesn't silently become your default for everything.

Setting Default Description
aiForge.data.reportDensity comfortable Spacing in generated HTML reports. Also adjustable live in the preview's Design tab
aiForge.data.directEdit true Hover a report section to move/duplicate/delete it, and double-click text to edit in place. These edits never call the AI

Report branding isn't a setting either — it lives in evolve-report-theme.json in your workspace (colours, palette, logo, footer, default report shape), so it travels with the project and can be committed. Run Evolve AI: Create Report Theme (branding) to scaffold it.


Keyboard Shortcuts

Action Windows / Linux macOS
Open chat (sidebar) Ctrl+Shift+A Cmd+Shift+A
Open chat (editor tab) Click the Evolve AI icon in the editor title bar Same
Generate code from description Ctrl+Alt+G Cmd+Alt+G
Fix current file errors Ctrl+Alt+F Cmd+Alt+F
Explain selected code Ctrl+Alt+E Cmd+Alt+E
Generate commit message Ctrl+Alt+M Cmd+Alt+M

How Context Works

Every AI call automatically includes:

  1. Active file — full content of your current file (priority budget allocation)
  2. Related files — imported/importing files (remaining budget, capped at maxContextFiles)
  3. Diagnostics — current errors and warnings (if includeErrorsInContext is enabled)
  4. Git diff — unstaged changes (if includeGitDiffInContext is enabled)
  5. Plugin context — domain-specific data from active plugins (e.g., dbt manifest, Terraform state, Databricks cluster info)

Total characters capped by contextBudgetChars (default 24,000). Increase for larger models; decrease for faster/cheaper ones.


Cloud Plugin Setup Guides

Databricks Connected

Connect to your Databricks workspace for live cluster management, job monitoring, notebook deployment, Unity Catalog browsing, and SQL execution.

What you need: A Databricks workspace URL and a Personal Access Token (PAT).

Setup:

  1. Open the command palette (Ctrl+Shift+P)
  2. Run Evolve AI: Databricks: Connect to Workspace
  3. Enter your workspace URL (e.g., https://adb-1234567890.12.azuredatabricks.net)
  4. Enter your Personal Access Token
    • Generate one at: Workspace > User Settings > Developer > Access Tokens > Generate New Token
  5. The status bar will show a green dot with your workspace name when connected

Available commands after connecting:

Command What it does
List Clusters Shows all clusters with status, type, and Spark version
Cluster Details & Optimization AI analyses a cluster's config and suggests optimizations
List Jobs Shows all jobs with schedule and last run status
Run Job Triggers a job run and monitors it
Analyse Failed Job Run Fetches error logs from a failed run — AI diagnoses the root cause
Design Workflow with AI Describe what you need — AI designs a complete Databricks workflow
Browse & Import Notebook Navigate workspace notebooks and open them locally
Deploy Current File as Notebook Push the current file to your Databricks workspace
Explore Unity Catalog Browse catalogs, schemas, and tables with AI-powered data model analysis
AI Query Suggestion for Table Select a table — AI generates useful queries for it
Execute SQL on Warehouse Run SQL against a SQL warehouse and see results
Manage DLT Pipeline View, start, stop, and troubleshoot Delta Live Tables pipelines

AWS Connected

Connect to your AWS account for Lambda management, Glue job monitoring, S3 browsing, CloudFormation analysis, Step Functions design, and DynamoDB exploration.

What you need: An IAM user or role with programmatic access (Access Key ID + Secret Access Key).

Recommended IAM permissions: ReadOnlyAccess for browsing, plus lambda:InvokeFunction, glue:StartJobRun, s3:PutObject, states:StartExecution for execution commands.

Setup:

  1. Open the command palette (Ctrl+Shift+P)
  2. Run Evolve AI: AWS: Connect to Account
  3. Enter your AWS Access Key ID
  4. Enter your AWS Secret Access Key
  5. Enter your AWS Region (e.g., us-east-1, eu-west-1)
  6. The extension tests the connection with STS GetCallerIdentity

Environment variable alternative: Set AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and AWS_DEFAULT_REGION — the plugin picks these up automatically.

Available commands after connecting:

Command What it does
List Lambda Functions Shows all functions with runtime, memory, and timeout
Lambda Function Details Deep-dive into a function's config — AI suggests optimizations
Invoke Lambda Function Run a function with custom payload and see the response
View Lambda Logs Fetch recent CloudWatch logs for a function
Debug Lambda Errors Fetches error logs + config for functions with recent errors — AI diagnoses issues
List Glue Jobs Shows all Glue jobs with type, version, and worker count
Glue Job Details Inspect job config, script location, connections
Run Glue Job Trigger a Glue job with optional arguments
Analyse Glue Job Failure Pick a failed run — AI analyses the error and suggests fixes
Browse Glue Data Catalog Navigate databases and tables with schema details
Browse S3 Drill into buckets and folders, download files to editor
Deploy File to S3 Upload the current file to an S3 bucket
List CloudFormation Stacks Shows stacks with status and drift detection
CloudFormation Stack Details Resources, outputs, events, template — AI explains the architecture
List Step Functions Shows state machines with definition analysis
Design Step Function with AI Describe a workflow — AI generates the complete ASL definition
Explore DynamoDB Browse tables, inspect schemas, sample data — AI suggests access patterns

Google Cloud Connected

Connect to your GCP project for Cloud Functions management, Cloud Run monitoring, BigQuery analysis, GCS browsing, Pub/Sub messaging, and Firestore exploration.

What you need: A GCP service account JSON key file and your project ID.

Setup:

  1. Open the command palette (Ctrl+Shift+P)
  2. Run Evolve AI: Google Cloud: Connect to Project
  3. Select your service account JSON key file (file picker dialog)
    • Create one at: GCP Console > IAM & Admin > Service Accounts > Keys > Add Key > JSON
  4. Enter your GCP project ID
  5. The extension tests the connection by fetching project info

Recommended roles: Viewer for browsing, plus Cloud Functions Invoker, BigQuery User, Storage Object Admin for execution commands.

Available commands after connecting:

Command What it does
List Cloud Functions Shows all functions with runtime, status, and trigger type
Function Details Deep-dive into config — AI suggests optimizations
Invoke Function Call an HTTP function with custom payload
View Function Logs Fetch Cloud Logging entries for a function
Debug Function Errors Scans for functions with errors — AI diagnoses issues
List Cloud Run Services Shows services with URL, revision, and scaling config
Cloud Run Details Inspect config, scaling, traffic routing — AI optimizes
Explore BigQuery Browse datasets and tables with schema details — AI explains data model
Run BigQuery SQL Execute a query and see results — AI analyses the output
Analyse BigQuery Failures Inspect failed BigQuery jobs — AI diagnoses query issues
Browse Cloud Storage Navigate buckets and objects, download to editor
Deploy to Cloud Storage Upload current file to a GCS bucket
List Pub/Sub Topics Shows topics and subscriptions — AI explains messaging architecture
Publish Pub/Sub Message Send a message to a topic
Explore Firestore Browse collections and documents — AI explains data model

Azure Connected

Connect to your Azure subscription for Functions management, Logic Apps monitoring, Cosmos DB querying, Storage browsing, DevOps pipeline analysis, and Log Analytics.

What you need: An Azure service principal (App Registration) with Tenant ID, Client ID, Client Secret, and Subscription ID.

Setup:

  1. Open the command palette (Ctrl+Shift+P)
  2. Run Evolve AI: Azure: Connect to Subscription
  3. Enter your Tenant ID
  4. Enter your Application (Client) ID
  5. Enter your Client Secret
  6. Enter your Subscription ID
  7. The extension tests the connection by fetching subscription info

Creating a service principal:

# Using Azure CLI
az ad sp create-for-rbac --name "Evolve-AI" --role "Reader" \
  --scopes /subscriptions/<your-subscription-id>

This outputs appId (Client ID), password (Client Secret), and tenant (Tenant ID).

Available commands after connecting:

Command What it does
List Function Apps Shows all Azure Functions apps with runtime and status
Function App Details Pick an app — AI analyses config and suggests optimizations
Invoke Function Call a function with custom payload
View Function Logs Fetch recent logs — AI analyses errors
Debug Function Errors AI diagnoses problematic function apps
List Logic Apps Shows Logic Apps with status and workflow info
Analyse Logic App Failure Inspect failed runs — AI diagnoses issues
Explore Cosmos DB Browse accounts, databases, containers — AI explains data model
Query Cosmos DB Run SQL queries against a container
Browse Storage Navigate storage accounts, containers, blobs — download to editor
Deploy to Storage Upload current file to blob storage
List DevOps Pipelines Shows pipelines with recent run status
Analyse Pipeline Failure Pick a failed pipeline run — AI diagnoses the issue
List Web Apps Shows App Service web apps with status
Restart Web App Restart a web app with confirmation
Query Log Analytics Run KQL queries against a Log Analytics workspace
List Active Alerts Shows Azure Monitor alerts — AI explains and suggests remediation

Troubleshooting

Chat shows OFFLINE / No response

Ollama not detected:

  1. Verify Ollama is running: open http://localhost:11434 in your browser — it should say "Ollama is running"
  2. If using Windows and localhost doesn't work, try setting aiForge.ollamaHost to http://127.0.0.1:11434
  3. Make sure you have a model pulled: ollama list should show at least one model
  4. Check the model name matches aiForge.ollamaModel (default: qwen2.5-coder:7b)

Cloud provider not responding:

  1. Check your API key is set: run Evolve AI: Switch AI Provider and re-enter your key
  2. Verify network connectivity to the provider's API endpoint
  3. Check VS Code's Developer Tools console (Help > Toggle Developer Tools) for error messages

Chat input not responding / buttons don't work

  1. Reload the window: Ctrl+Shift+P > "Developer: Reload Window"
  2. If the issue persists, close and reopen the chat panel
  3. Check VS Code's Developer Tools console for JavaScript errors in the webview

Plugin not activating

Plugins activate automatically based on workspace files. If a plugin isn't showing:

  1. Make sure the workspace contains the expected marker files (see the plugin table above)
  2. Check aiForge.disabledPlugins in settings — make sure the plugin ID isn't listed
  3. Reload the window to trigger re-detection

Cloud plugin shows "not connected"

  1. Run the connect command for your provider (e.g., AWS: Connect to Account)
  2. Verify your credentials are correct — the connect command tests the connection
  3. Check that your credentials have sufficient permissions (see setup guides above)
  4. For AWS: ensure your region is correct and your IAM user/role is active
  5. For GCP: ensure the service account JSON key is valid and not expired
  6. For Azure: ensure the client secret hasn't expired
  7. For Databricks: ensure the PAT hasn't expired and your workspace URL is correct

Commands show "command not found"

This happens when a cloud plugin command is triggered but the plugin isn't active. Cloud plugin commands only register when:

  1. The plugin detects matching files in your workspace (e.g., serverless.yml for AWS)
  2. The plugin has activated (connected to the cloud provider)

Fix: Open a workspace that contains files for that cloud platform, then run the connect command.

Code conversion says it's "too big for one pass"

This is the tool doing its job, not an error. A context window is shared between the prompt and the response, and a conversion's response is roughly the size of its input — so a job that would silently truncate is caught before the request goes out. Three ways forward:

  1. Let it split — the offered pass count works; each pass sees what the earlier ones produced. Cross-file consistency is weaker than a single pass, which is why you're offered the alternatives first.
  2. Use a bigger model — Choose AI Model for Code Conversion shows each installed model's real context window and flags which ones handle the job in one pass. qwen2.5-coder:14b (32k) or deepseek-coder-v2 (64k) cover most single-file work.
  3. Convert less at once — a module at a time is easier to review anyway.

Converted files come back truncated or with no fidelity report

Almost always the model, not the conversion. A general-purpose 7B chat model will emit code but frequently drops the structured report — you'll see the "no conversion report was returned" warning. Switch to a coding-tuned model (qwen2.5-coder, deepseek-coder-v2) for the conversion; models under 7B routinely truncate whole files. The offline provider is pattern-based, not an LLM, and cannot convert code at all — the panel says so rather than producing nonsense.

"Check it parses" says the toolchain isn't installed

The check runs the target language's own parser, so converting to Go needs Go on your PATH. The message names exactly what's missing. It's optional — conversion works fine without it; you just don't get the parse verdict. Some languages (C#, Kotlin, Scala, Rust, SQL, COBOL…) have no sound single-file check and say so rather than inventing one.

Slow responses

  1. Ollama: Use a smaller model (e.g., qwen2.5-coder:3b instead of 7b)
  2. Context too large: Reduce aiForge.contextBudgetChars (try 12000) or aiForge.maxContextFiles (try 3)
  3. Cloud context: Connected plugins add live data to context — this adds a small delay on each request

Gemma 4 setup wizard issues

"aiForge.gemma4Model is not a registered configuration" error

  • This happens on v1.4.0 only, when the extension is installed or upgraded into a running VS Code window. VS Code's Configuration Registry hasn't picked up the new settings schema yet.
  • Fix: Reload the window (Ctrl+Shift+P → "Developer: Reload Window"), then run Switch AI Provider → Gemma 4 again. Setup will complete normally.
  • Fixed in v1.4.1+: The wizard now detects this and shows a one-click Reload Window button automatically.

"System cannot run Gemma 4" modal appears

  • Your RAM or free disk space is below the minimum for any variant (8GB RAM, 8GB disk)
  • The modal lists the specific blockers and three alternatives (cloud, offline, free up resources)
  • If you know you have plenty of disk, the check looks at the Ollama models directory (~/.ollama/models on Linux/macOS, %USERPROFILE%\.ollama\models on Windows). Run df -h ~/.ollama (or check disk in Explorer) to confirm

Setup hangs at "Downloading… 0%"

  • Verify Ollama is running: open http://localhost:11434 in your browser
  • Verify internet connectivity: ping ollama.ai
  • If stuck more than 5 minutes, click Cancel in the progress notification and retry

Hardware detection shows "No GPU detected" but you have one

  • NVIDIA: ensure nvidia-smi is on your PATH (nvidia-smi --version in terminal)
  • AMD: ensure rocm-smi is installed (Linux only)
  • Apple Silicon: detection requires system_profiler (built-in on macOS)
  • Intel integrated GPUs are not detected — Gemma 4 won't use them anyway
  • You can manually pick a variant via the "Choose Different Variant" button in the wizard

Ollama upgrade fails during setup

  • The wizard auto-upgrades Ollama when it's older than 0.3.10 (required for Gemma 4)
  • If the upgrade fails, manually download from ollama.com and run the installer
  • Then re-run Evolve AI: Switch AI Provider → Gemma 4

"Could not find gemma4 variant" after setup completes

  • Ollama may still be pulling the model in the background — wait 5-10 minutes
  • Verify with ollama list in your terminal — should show your gemma4:* tag
  • If missing, run ollama pull gemma4:e4b manually and try again

How to disconnect / change credentials

Run the disconnect command for your provider:

  • Evolve AI: AWS: Disconnect
  • Evolve AI: Google Cloud: Disconnect
  • Evolve AI: Azure: Disconnect
  • Evolve AI: Databricks: Disconnect

Then run the connect command again with new credentials.


FAQ

General

Q: Is my code sent to the cloud? A: It depends on your provider. With Ollama, everything stays on your machine — no data leaves your network. With cloud providers (Anthropic, OpenAI, Google Gemini, HuggingFace), your code context is sent to their API. Choose based on your privacy requirements.

Q: Which AI provider should I use? A: For privacy and cost: Gemma 4 or Ollama (free, local, your code never leaves your machine). For best quality: Anthropic Claude, OpenAI GPT-4o, or Google Gemini 2.5 Pro. For speed on a budget: Groq (via OpenAI-compatible endpoint) or Gemini 2.0 Flash. For no setup: the built-in offline mode (limited to pattern-based analysis).

Q: What is Gemma 4 and why should I use it? A: Gemma 4 is Google's latest open-weight AI model (Apache 2.0 license). It runs locally via Ollama with no API key, no cost, and no data leaving your machine. It supports text, image, and audio input with 128K-256K context windows. The E4B variant (~9.6GB) is recommended for most users. Select Gemma 4 in the provider switcher for a guided setup wizard.

Q: How does the Gemma 4 setup wizard pick the right variant for me? A: When you select Gemma 4, Evolve AI asks one-time consent to inspect your system: RAM (os.totalmem()), GPU (NVIDIA via nvidia-smi, AMD via rocm-smi, Apple Silicon via system_profiler), free disk space, and your Ollama version. It scores each variant against your hardware and recommends one — typically E2B for 8GB RAM, E4B for 16GB, 26B MoE for 32GB+, 31B Dense for 32GB+ with a GPU. No data leaves your machine — detection is purely local.

Q: Will the wizard auto-install Ollama or download models without asking? A: No. Each step asks explicit consent before running:

  • "Install Ollama?" — opens the official installer download (only if Ollama isn't installed)
  • "Upgrade Ollama?" — only if your version is older than 0.3.10 (required for Gemma 4)
  • "Download ?" — confirms before pulling the model You see a setup plan listing every step before clicking "Install Everything", and the whole process is cancellable mid-way.

Q: What if my system can't run Gemma 4? A: The wizard shows a modal explaining exactly why (e.g. "only 4GB RAM detected — needs at least 8GB") and offers three actionable alternatives:

  • Switch to a cloud provider (Anthropic Claude, OpenAI, Google Gemini, GLM/Z.ai, HuggingFace) — runs in the cloud, only needs an API key
  • Use Offline mode — pattern-based AI, no LLM required, works instantly
  • Free up resources — disk-space tips if that's the blocker You're never left at a dead end.

Q: Can I disable hardware detection? A: Yes. Set aiForge.allowHardwareDetection to false in settings. The wizard then falls back to showing all 4 variants without inspection — you pick manually. Or decline the one-time consent dialog when it first appears.

Q: Can I use multiple providers? A: You can switch providers at any time via Evolve AI: Switch AI Provider. The extension uses one provider at a time.

Q: What models work with Ollama? A: Any model Ollama supports. Recommended: gemma4:e4b (Google Gemma 4, multimodal, strong coding), qwen2.5-coder:7b (code-optimized), codellama:13b (larger, better quality), deepseek-coder:6.7b. Run ollama list to see installed models.

Q: Does Evolve AI work with LM Studio / llama.cpp / Jan? A: Yes. Set aiForge.ollamaHost to your server's URL (e.g., http://localhost:1234/v1 for LM Studio). These servers implement the same API as Ollama.

Plugins

Q: How do plugins activate? A: Automatically. When you open a workspace, Evolve AI scans for marker files (e.g., Dockerfile for Docker, manage.py for Django). Matching plugins activate silently and start injecting domain knowledge into every AI interaction. The status bar shows active plugins.

Q: Can I disable a plugin? A: Yes. Add the plugin ID to aiForge.disabledPlugins in settings. Example: ["databricks", "docker"]. Plugin IDs: databricks, databricks-connected, aws, aws-connected, gcp, gcp-connected, azure, azure-connected, dbt, airflow, pytest, fastapi, django, terraform, kubernetes, docker, jupyter, pytorch, security, git.

Q: What's the difference between the base and connected versions of cloud plugins? A: The base plugin (e.g., AWS) activates on file detection and injects best-practice knowledge into AI responses — no credentials needed. The connected plugin (e.g., AWS Connected) adds live API access — browse resources, run queries, analyze failures, deploy code. Both can be active simultaneously.

Q: Do cloud plugins cost anything? A: The plugins themselves are free. But they call your cloud provider's APIs, which may incur costs depending on your plan. Read-only operations (listing resources, reading logs) are typically free or low-cost. Execution operations (invoking Lambda, running BigQuery queries) may have associated costs.

Cloud Credentials

Q: Where are my credentials stored? A: In VS Code's encrypted SecretStorage — the same mechanism VS Code uses for its own authentication. Credentials are never written to settings files, .env files, or any plaintext location.

Q: Can I use temporary/session credentials? A: For AWS, yes — you can provide a session token along with your access key and secret key. For Azure, the client secret has an expiry set in Azure AD. For GCP, service account keys don't expire but can be rotated. For Databricks, PATs have configurable expiry.

Q: What permissions do I need? A: At minimum, read-only access to list and inspect resources. For execution features (invoking functions, running jobs, deploying files), you need the corresponding write permissions. See each cloud plugin's setup guide above for specific IAM recommendations.

Q: Is it safe to use in production? A: The extension only performs the actions you explicitly trigger via commands. It never modifies cloud resources automatically. Execution commands (run job, invoke function, deploy) always require your manual action.


Contributing

Contributions are welcome! See CONTRIBUTING.md for the full guide.

Quick start for contributors:

  1. Fork the repo and clone it
  2. npm install && npm run watch
  3. Press F5 to launch the Extension Development Host
  4. Make your changes and test them

Easiest way to contribute — add a new stack plugin:

  1. Read docs/PLUGIN_GUIDE.md for the step-by-step template
  2. Create src/plugins/<name>.ts implementing the IPlugin interface
  3. Register it in src/plugins/index.ts
  4. Add commands to package.json under contributes.commands

Plugin ideas (community contributions welcome):

  • Next.js — App Router, Server Components, API routes
  • Rust — ownership, lifetimes, async patterns
  • Go — goroutines, interfaces, error handling
  • GraphQL — schema, resolvers, queries
  • React Native — Expo, Metro, native modules
  • Spring Boot — Java/Kotlin, dependency injection, JPA

Requirements

  • VS Code 1.85.0 or later
  • For local AI: Ollama with a pulled model
  • For cloud AI: An API key from your chosen provider
  • For cloud plugins: Appropriate credentials (see setup guides above)

License

MIT

  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft