Skip to content
| Marketplace
Sign in
Visual Studio Code>Other>Spec Execution EngineNew to Visual Studio Code? Get it now.
Spec Execution Engine

Spec Execution Engine

Discovery & Quantum

| (0) | Free
Run the Spec Execution Engine from VS Code: launch a run, watch waves live, and merge reviewed work into the run branch without leaving the editor.
Installation
Launch VS Code Quick Open (Ctrl+P), paste the following command, and press enter.
Copied to clipboard
More Info

Spec Execution Engine

Describe a feature. Get pull requests.

Write a spec, and the engine takes it from there: it breaks the work into waves, hands each one to the coding agent, has the change reviewed and built, and stops to ask you before anything merges. You watch it happen in a view beside your code, and you make the call at the end.

Nothing merges without you.


The loop, in four steps

1. Write a spec

In chat:

@see /spec

You are asked a handful of short questions — which repositories the code should land in, where the spec should live, what you are building, and whether the spec arrives as a pull request or goes straight in. A new capability proposes a branch from its name and lets you choose a different one for parallel feature work. Repository names are checked as you type them, so a typo is caught in the interview rather than ten minutes into a run.

For enterprise repositories, the last form also binds each target to an Azure DevOps pipeline. Pipeline discovery reuses the Microsoft account already signed in to VS Code. If the catalogue cannot be listed, the form shows the cause and requires either a positive numeric pipeline id or an explicit none.

Already written a spec? Choose Make an existing spec SEE-compatible instead. Your prose is left exactly as it is; only the front-matter the engine reads is added.

The spec is committed for you, and the settings that point at it are saved to the workspace — not to your user profile, so they do not follow you into unrelated projects.

If a generated L2 spec still has Needs input: questions, authoring pauses before decomposition. Select Answer the questions in chat or the question counter in the status bar to edit the spec in the rich editor. Its find bar opens on Needs input:; use next/previous to scroll to each matching question while keeping focus on the find controls, then click the text to replace the marker with your answer. Enter/Shift+Enter in the find field also moves between matches. Edits save automatically, and the counter follows saved answers. Needs record: bookkeeping blanks do not block this step. Typing refreshes the find highlights and match count without moving your caret or scrolling away from the answer. Use the find controls to navigate explicitly.

Once the questions are answered, select Carry on to resume. Continue anyway still lets you deliberately defer the remaining decisions. Both paths save pending rich-editor edits before continuing; a save failure leaves the flow paused. The Markdown toolbar button opens the same file in the source editor if you prefer raw Markdown. Save source edits to refresh the rich view. The question watch survives a window reload.

The rich editor preserves leading YAML front matter (including target repositories and CI bindings) unchanged and shows a notice instead of rendering it as prose. Use Markdown mode to edit that configuration or paste a complete document containing front matter. The fenced Metadata section in the body remains editable. Unchanged Markdown blocks retain their original source, including comments, reference definitions, indentation and line endings; a small rich edit does not reformat the rest of the spec.

This answer surface is for the generated L2 spec; an equivalent L3 answer step is not yet provided.

For a hierarchy, spec authoring stops twice for review: once after the decomposition plan, and once after the component specs. When either review is merged, the Run view keeps a durable continuation row, a badge, and a warning-colored status action. Dismissing the notification does not dismiss the work. The badge signals pending attention and points to the continuation row or warning status item. Carry on in the notification submits the review request once. Open Chat in the continuation row or status bar prepares the same request without sending it; this option is available even before the notification is dismissed. No additional text is needed. Press Send in Chat to continue. The pending row says Open Chat, then press Send. After Chat opens the draft, a notice says: "Request ready in Chat. Press Send to continue. No additional text is needed." Carry on does not show this notice. Continuing removes the pending Send instruction. The prepared @see /spec includes a visible --see-review=... identifier. Do not edit it. Repeated notification clicks do not submit the same attempt again. Only an accepted request shows Continuing, without another reminder action. An ambiguous manually typed /spec asks for an in-chat Continue this review confirmation. VS Code can still show a queued request; the extension prevents that identified request from executing twice, not from entering the queue. After partial plan progress, Retry spec preparation confirms the exact state written by that attempt. Old attempts do not intercept ordinary /spec after implementation takes ownership. The action clears only after the ledger shows that the flow moved past that exact review, or that the run was reset or replaced. Re-entry state is scoped to the current extension host and ledger, and windows that share that ledger use an exclusive local claim so only one ordinary notification appears. Review checks run immediately at startup and then every 30 seconds, only at a review gate and until that gate is announced. The local ledger refresh remains 1200 ms. Notification focus behavior is unchanged.

After the component-spec review, successful packet preparation records the existing ready phase. The next ledger update removes the review reminder and the Spec row shows committed; implementation still requires its separate confirmation. A failed or cancelled request does not clear the reminder, and an old request cannot advance a replacement run or review. Retry becomes available only after the underlying work finishes, not when cancellation is requested. Attempt identifiers are local to the current extension activation: after a reload, use a fresh Open Chat request. This is not cross-window or crash recovery coordination. A retry can rewrite and commit packet files, so do not retry only to remove a reminder without checking the packet inputs and branch.

In an ai-native-dev source checkout, see docs/see-review-continuation.md for the approved scope, known limits, and possible future work.

2. Implement it

Chat hands you Implement this spec as soon as a spec exists; it also sits in the Run view. The Run view opens when you start, so there is always something to watch.

A run is not a dry run — it dispatches work to the coding agent and opens real pull requests. So it asks first, and tells you what it will build, where the code lands, each target-to-pipeline binding, ADO authentication readiness, and the retry budget it will use. A configured pipeline that cannot be authenticated or reached blocks Start. A target with no pipeline requires the explicit Start Without CI acknowledgement.

Sign-in comes from VS Code's own GitHub and Microsoft accounts. There is no token to paste.

3. Watch the waves

The Run view shows every wave as it moves through dispatch, AI review and CI: which unit it is on, how many retries it has spent, and links straight to the pull request and the build.

When a wave goes red, the engine hands the failure back to the coding agent and lets it try again — up to your retry budget — before it stops and asks you.

4. Make the call

When a wave needs a decision, you get a notification and the wave is marked in the view. The notification offers one thing: Review PR. The decision belongs after reading the code, not instead of it.

Once you have read it, hover the wave for three buttons:

✓ Merge into Run Branch The work is good; it is squash-merged into the run branch the engine owns for this spec — never your default branch — and the engine carries on.
↻ Send Back to the Coding Agent Say what to change, and that becomes the brief — or re-run the review on the work as it stands.
✕ Reject and Close the PR Discard this approach. The pull request is closed unmerged; the branch is kept.

The fourth answer — not yet — is ⏸ Pause: Decide Later in the view's toolbar, where the ■ stop button sits the rest of the time. It is there and not on the row because pausing is not a per-pull-request act: one process holds the whole run, so pausing stops all of it. Decisions are per pull request, because a human judges this code; attachment is per run, because one process holds all of it.

Decide Later is a real answer, not a way of giving up. Nothing merges without an explicit merge decision, and pausing puts no clock on you — take hours or days. It stops the whole run, though, so any other review open at the same time pauses with it. ▶ Resume the Run picks all of them up again — one click, however many are open — after re-checking the latest commits in case a pull request moved on while you were away. Nothing is re-dispatched, so a resume spends no new AI credits on work that already exists.

The first three confirm before they act, and each records the decision on the pull request so the engine has one durable channel. The SEE Run view is the primary UI; manual decision comments remain only as a compatibility and recovery path.

When the row says "engine on standby"

The engine does not hold a process open while you think. It watches an open gate for HIL_POLL_TIMEOUT_SECONDS (an hour by default) — the row says until when while the clock runs:

Waiting for your decision · watching until 19:05

and the pull request comment says the same thing. After that it steps back on its own, posts a note on the pull request, and the row reads:

Waiting for your decision · engine on standby

Nothing has changed: the pull request is still open, nothing was merged or closed, and there is still no deadline. The ⏸ button disappears — the only thing missing is something to act on your answer, and ▶ Resume the Run appears in the view's toolbar to supply it.

Press ✓ ↻ ✕ in the editor as usual. The extension records the decision, resumes the engine automatically, and applies it. If a compatible decision comment was left manually as a recovery action, it is queued, not lost, but it does not act while the engine is on standby. Resume first. If new commits landed, the engine asks again rather than merge code you never saw.

The same is true if you stop the run yourself while a review is open: the engine leaves a note on the pull request saying nobody is watching it and directs the reviewer to the SEE Run view.

Your machine stays awake while a run is in progress

A wave can spend an hour waiting on the coding agent and CI, so a run is normally left alone. The coding agent and CI run on GitHub regardless, but the engine has to be awake to watch them — and on most current laptops, locking the screen puts the machine into Modern Standby within minutes, which throttles the engine and cuts its network.

So the engine holds the machine awake for exactly as long as a run lasts. Two things follow:

  • The screen stays lit. Lock it and you get the lock screen — nothing is exposed — but the display is on. Run long waves on mains; unplugged, this will drain the battery.
  • Only during a run. The hold is taken when the engine starts and released when it stops, crash included. With no run in progress your machine sleeps exactly as it always did.

Getting started

  1. Install the extension.
  2. Open the Spec Execution Engine view in the activity bar.
  3. Click Get started — the walkthrough covers the whole loop in four steps.

There is nothing to install first. The extension ships the engine inside it, and the first time it needs Python it offers to set one up: it looks for a usable interpreter you already have, and if there is none it installs Python 3.11 into its own private folder and builds the environment there. Nothing lands on your PATH and nothing else on the machine changes.

Windows, Linux and macOS are all supported, on Intel and ARM. On Linux and macOS the private interpreter is a relocatable CPython build, verified against the checksums published with it — so no package manager is involved and nothing is installed system-wide. Alpine and other musl-based distributions are out of scope: the Copilot CLI the engine drives has no musl build, so Python alone would not get you a working run.

If a run is already going in a terminal, the view attaches to it automatically.

Command Palette

SEE exposes only these extension commands in the Command Palette, regardless of see.developerMode:

  • Spec Execution Engine: Reset — Forget the Current Run
  • Spec Execution Engine: Create Support Bundle
  • Spec Execution Engine: Get Started

Author through @see /spec, set pipelines with @see /pipelines, and start implementation from the chat button. Stop, Resume, and the review actions remain in the Run view, with their existing availability conditions. Its toolbar still offers Show Engine Log and Refresh, and its menu offers Implement in a Terminal Instead.

VS Code itself also adds Focus on Run View and View: Show Spec Execution Engine for the contributed view. Those navigation entries remain: the extension cannot hide them with Command Palette menu rules without disabling the view.

Other command handlers are retained for existing buttons, walkthrough links and custom keybindings. Diagnose Setup, Check AI Reviewer Sign-in, and Rebuild Python Environment no longer have palette entries; hiding them reduces troubleshooting discoverability. They can still be assigned custom keyboard shortcuts using see.doctor, see.checkCopilotAuth, and see.repairRuntime.


Settings

Most of these are filled in for you when you create a spec. You should rarely need to set one by hand. The walkthrough's Check configuration link walks through the essential ones; Configure is no longer a Command Palette entry.

In Settings (Ctrl+,), search @ext:mdq.spec-execution-engine. Settings are organized into these categories, with advanced options last:

  1. Selected Spec - repository, path, read ref, integration branch and its managed owner identity, and fallback targets.
  2. Workflow - authoring flow, reviewer model, and retry budget.
  3. Authentication - GitHub host and Azure DevOps tenant.
  4. Editor - outline visibility, outline font size, and zoom.
  5. Advanced Runtime - engine checkout, Python interpreter, and download mirror.
  6. Advanced / Diagnostics - dashboard port, environment overrides, and the legacy developer-mode preference.

Grouping changes presentation only. All existing see.* keys, defaults, and user/workspace overrides remain valid. Advanced categories are always available; they are not hidden by see.developerMode.

Setting What it is for
see.specRepo, see.specPath Which spec to implement. Set for you when you create one.
see.specGitRef Branch or tag to read the spec from. Empty means the default branch.
see.integrationBranch Workspace-only engine-owned target branch for a hierarchy capability. Set from placement and bound to the selected spec repo/path; direct identity changes make the old branch ineligible. Empty derives see/<spec-id>.
see.targetRepos Where code lands, when the spec does not say. The spec's own targets: wins.
see.maxVerifyRetries How many times a failing wave may be sent back to the coding agent before it stops and asks you. 0 brings every failure straight to you.
see.engineDir Only for working on the engine itself: run a checkout instead of the copy inside the extension.
see.pythonPath Which Python to use. Detected automatically.
see.reviewModel Which model reviews the code. Empty uses the engine's default.
see.authorModel Which model writes the specs. Defaults to gpt-6-astra so every run is authored by the same model; empty uses whatever the chat request carries.
see.adoTenantId Only if your Azure DevOps organization is in a different tenant than your Microsoft account.
see.env Extra environment variables for the engine. Applied last, so they win.

On upgrade, SEE migrates the old see.specGitRef=see/<slug> placement only for a complete L2 hierarchy selection with no branch owner, explicit empty, operator override, or durable branch. An explicit empty see.integrationBranch opts out.


If something does not work

"No engine found." The extension carries its own copy, so this should not happen — it means the install is damaged, or see.engineDir points somewhere that no longer holds run_SEE.py. Clear that setting, or reinstall.

Python problems. An interpreter is only used if it can actually import what the engine needs, so one that exists but cannot run the engine is skipped instead of tried. If none qualifies, the extension offers to build its own. Diagnose Setup reports what it found — engine, interpreter, sign-ins — and Rebuild Python Environment throws the private one away and starts over, which is the fix if an install was interrupted.

An opaque failure from Azure DevOps. A token from the wrong tenant signs in successfully but is not authorized. Set see.adoTenantId to the tenant your ADO organization lives in.

Anything else. Show Engine Log has the engine's full output, including which identity and which interpreter were chosen. And Implement in a Terminal Instead runs the same engine, the same way, with the environment already set up — the view attaches to that run too.

MCP startup diagnostics

The Spec Execution Engine output channel has [mcp] phase records. activationMs starts at entry to extension activation. Each timed operation has an ID and elapsedMs, measured with a monotonic clock. Python warmup, definition preparation, Python resolution, GitHub host resolution, GitHub identity, ADO identity, and host commands have separate records. ADO and GitHub work can overlap, so do not add their durations to calculate total startup time. Failures record the phase and elapsed time, but omit error payloads that can contain credentials. Launch keys, token hashes, auth output, and environment values are not part of these phase records.

Record What it proves
definitions offered The provider prepared a definition. VS Code might not have added it to its server list yet.
host launch resolution observed VS Code requested launch resolution for the current launch inputs and generation. It does not prove a running process.
start-command ... completed The host command returned. It can also return when no server ID matches.
process=stdio ... event=running The source-built server reached its startup function.
event=transport-connected The server attached its stdio transport, not that the MCP handshake completed.
event=protocol-initialized This server received the client's initialized notification.
metadata selected; live readiness unverified A tool candidate is available for invocation. The metadata can be cached.
phase=tool-invocation ... completed The actual invocation returned. This duration can include the whole interview and user input; it is not startup latency.

The process records use the server PID and processElapsedMs from process start. The extension tails them from the existing scaffold log. Log delivery is delayed by the tail interval; it is not a host readiness event. A completed invocation does not mean a spec was written: the result can describe cancellation or a tool-level failure. All three tool paths (scaffold_spec, set_pipelines, and ask_external_repos) use the same invocation timing wrapper.

Activation and /spec share a pending start request. The first command is sent without waiting for discovery. If it returns before the host knows the server, startup retries the non-destructive start command within a 15-second discovery window. It stops retrying after launch resolution for the current offer. Provider return alone is not sufficient. Concurrent tool requests also share recovery, and one cancelled request does not cancel the other waiters. An ordinary unchanged-input check does not refresh a running interview. The existing refresh barrier and recovery waits remain in place. Repeated offers with identical inputs keep the same launch generation; changed inputs do not accept an old launch callback. A forced unchanged refresh keeps the valid launch observation because the host need not resolve an identical definition again. It still requires a fresh offer and the existing refresh grace period. Only a provider call that began after the refresh event can satisfy that refresh's offer barrier. An older call that finishes later cannot release it.

The refresh input snapshot decides whether to request an offer; it does not pin the identity used by a later offer. The current provider offer and its launch generation are authoritative after the barrier. A snapshot check that finishes after a newer refresh request is ignored. No auth cache is added.

MCP wait limits and cancellation

Each chat request has one 110,600 ms cumulative MCP wait limit. Preparation refresh, tool startup/recovery, the existing reconnect delay, and later reconnect attempts use the same remaining time. The clock runs only while the request waits in these MCP lifecycle operations. It pauses during other chat work and interactive tool invocation. Overlapping waits charge elapsed time once; nested calls cannot reset the limit. This is not a whole-chat wall-clock limit or an expected startup time.

The preparation refresh remains at its existing call site, after runtime setup and sign-in. It does not run again on each tool lookup or interview question. All three tool paths use the same request scope. Settings-triggered refresh keeps its own bounded check. No time limit is added to Python installation, dependency setup, user sign-in, or interactive invocation.

MCP phase Maximum wait
Initial silent refresh inputs 15,000 ms
Fresh provider offer and refresh grace 5,000 ms + 300 ms
Start, commands and discovery together 15,000 ms
Tool metadata after start 5,000 ms after observed launch; otherwise 15,000 ms
Restart command 15,000 ms
Tool metadata after restart 5,000 ms
Forced refresh inputs, offer, and grace 15,000 ms + 5,000 ms + 300 ms
Second start and discovery 15,000 ms
Final tool metadata 5,000 ms

The maximum branch totals 110,600 ms. It retains the earlier polling windows and uses the existing 15-second discovery allowance for the previously unbounded silent refresh inputs and restart command. It is a conservative ceiling, not a performance target. The one-second reconnect delay spends the same remaining total; it does not add a new allowance. Each phase also stops when the request's remaining total expires. Standalone activation has a 15,000 ms start limit. Standalone refresh has a 20,300 ms limit. A standalone tool lookup has a 90,300 ms recovery limit.

Cancellation releases that caller promptly. Other active waiters retain their own deadlines and shared work. After the last waiter leaves, a new caller does not join the abandoned flight. A deadline or cancellation returns no candidate (false for start, stopped for refresh, undefined for tool lookup). Reconnect does not invoke the old candidate when recovery stops; it preserves the original invocation failure. Unchanged launch inputs do not restart a healthy interview.

Refresh has explicit outcomes: unchanged, offered, inapplicable, superseded, offer-pending, and stopped. Only stopped ends preparation. An inapplicable refresh, a superseded input snapshot, or an elapsed offer-wait window can continue into the existing bounded recovery or manual fallback. These outcomes do not claim that refresh completed or that tools are live. Recovery spends only the same request's remaining time. A delayed offer must still satisfy the current epoch, launch generation, grace and pending-restart barriers before metadata can be selected. The existing no-MCP manual repository picker remains available; a stopped preparation request does not open it.

Caller limit, not process cancellation: VS Code host commands, workspace discovery, interpreter probes, and silent identity lookups do not all expose a cancellation API here. The caller stops waiting, but those operations can still finish, update their own internal caches, or fail later. Their terminal handlers remain attached. Stopped caller continuations cannot request a later refresh, restart, or start, or update the current MCP lifecycle state. Provider callbacks remain host-owned; only callbacks from the current registration and matching launch generation can supply launch evidence. Refresh epochs also reject pre-refresh offers. This is not a general provider-completion ordering change.

An already-issued host restart remains a safety barrier until its promise settles, even across provider registration replacement. Later callers must wait within their own remaining allowance; they cannot open an interview that the pending restart could then stop. Cancellation of a refresh after its event does not erase the fresh-offer or grace barrier. No host-owned process is killed. Timers require the extension-host event loop to run, so scheduler delay or synchronous blocking can make the measured return time exceed the nominal limit. Tests check elapsed time and return values with an explicit scheduler allowance, not just the presence of a timer.

A start or restart command timeout stops that recovery attempt, even if some cumulative time remains. The host command can still be active, so timeout does not authorize a blind restart or forced refresh. This differs from an ordinary offer-pending outcome. The 110,600 ms ceiling does not promise that every recovery step runs or that each request waits for the full limit.

Supported-host limit: VS Code 1.101 ignores waitForLiveTools, and the stable API has no live server/tool state event. Tool selection is therefore metadata-only; only the actual invocation establishes invocation completion. No health tool, second MCP client, or extra readiness process is used. The caller limits above include awaited lifecycle prerequisites and commands; they do not prove that an underlying operation stopped or that tools are live. Dependency provisioning remains separate work.

This distinction follows the VS Code 1.101 provider update path and start command: the host adds definitions after the provider returns, and start completes without an error when the server ID is not yet in its list. The 1.101 launch resolver passes the original returned definition object to the provider, which supports the object-to-generation lookup. Live-host and platform-matrix checks remain separate from these source-contract and isolated-process tests.

The ordering tests model separate provider-preparation, host-insertion, and process-initialization steps. The stdio smoke test uses a separate source-build process and no domain tool calls. It does not measure the host-owned MCP server:

npm test
npm run compile
node --test test\mcp-stdio.test.js

The investigation baseline at 03479e24cea35b18476457b96d6eccfe3c071c34 used 12 fresh Code.exe stdio processes: median initialize 102.1 ms, tools available 104.5 ms, and tools/list 2.4 ms, with three tools. These were process-cold runs with normal OS caches, not reboot-cold runs. Cached Python resolution had a separate median of 118.6 ms. These measurements do not include extension activation, identity acquisition, or host discovery.

Create a support bundle

Run Spec Execution Engine: Create Support Bundle from the Command Palette when support needs diagnostics. You choose the destination in a save dialog. The extension creates one local ZIP file and does not upload, attach, share, or open it. Review the ZIP before you attach it to a support case.

The bundle contains a manifest, a safe setup snapshot, the extension's bounded rolling logs, a sanitized scaffold diagnostic log, the current valid status snapshot, and at most three ledger rows. Each entry has a fixed name, a size limit, a final SHA-256 hash, and an included, omitted, truncated, or error state in manifest.json.

The scaffold diagnostic log records form field names, actions, result codes, and failure categories. It never records form answers, titles, prompts, generated documents, argument values, or full scaffolder results. Older raw scaffold.log files are not copied. The sanitized log stays in the extension's global storage across reloads until its 1 MiB one-generation writer truncates it; reloads do not replay its older lines into the visible Output Channel.

The bundle never includes .see.env, checkpoints, source or workspace files, open documents, diffs, prompts or responses, the full process environment, credentials, account names, machine names, SIDs, or home paths. It redacts data before persistent log writes and again before ZIP creation. Redaction is best-effort, so always review the bundle before sharing it.

  • Contact us
  • Jobs
  • Privacy
  • Manage cookies
  • Terms of use
  • Trademarks
  • Your Privacy Choices
  • Consumer Health Privacy
© 2026 Microsoft