Spec Execution EngineDescribe a feature. Get pull requests. Write a spec, and the engine takes it from there: it breaks the work into waves, hands each one to the coding agent, has the change reviewed and built, and stops to ask you before anything merges. You watch it happen in a view beside your code, and you make the call at the end. Nothing merges without you. The loop, in four steps1. Write a specIn chat:
You are asked a handful of short questions — which repositories the code should land in, where the spec should live, what you are building, and whether the spec arrives as a pull request or goes straight in. A new capability proposes a branch from its name and lets you choose a different one for parallel feature work. Repository names are checked as you type them, so a typo is caught in the interview rather than ten minutes into a run. For enterprise repositories, the last form also binds each target to an Azure
DevOps pipeline. Pipeline discovery reuses the Microsoft account already signed
in to VS Code. If the catalogue cannot be listed, the form shows the cause and
requires either a positive numeric pipeline id or an explicit Already written a spec? Choose Make an existing spec SEE-compatible instead. Your prose is left exactly as it is; only the front-matter the engine reads is added. The spec is committed for you, and the settings that point at it are saved to the workspace — not to your user profile, so they do not follow you into unrelated projects. If a generated L2 spec still has Once the questions are answered, select Carry on to resume. Continue anyway still lets you deliberately defer the remaining decisions. Both paths save pending rich-editor edits before continuing; a save failure leaves the flow paused. The Markdown toolbar button opens the same file in the source editor if you prefer raw Markdown. Save source edits to refresh the rich view. The question watch survives a window reload. The rich editor preserves leading YAML front matter (including target repositories and CI bindings) unchanged and shows a notice instead of rendering it as prose. Use Markdown mode to edit that configuration or paste a complete document containing front matter. The fenced Metadata section in the body remains editable. Unchanged Markdown blocks retain their original source, including comments, reference definitions, indentation and line endings; a small rich edit does not reformat the rest of the spec. This answer surface is for the generated L2 spec; an equivalent L3 answer step is not yet provided. For a hierarchy, spec authoring stops twice for review: once after the
decomposition plan, and once after the component specs. When either review is
merged, the Run view keeps a durable continuation row, a badge, and a
warning-colored status action. Dismissing the notification does not dismiss the
work. The badge signals pending attention and points to the continuation row or
warning status item. Carry on in the notification submits the review request
once. Open Chat in the continuation row or status bar prepares the same
request without sending it; this option is available even before the notification
is dismissed. No additional text is needed. Press Send in Chat to continue.
The pending row says Open Chat, then press Send. After Chat opens the draft,
a notice says: "Request ready in Chat. Press Send to continue. No additional text
is needed." Carry on does not show this notice. Continuing removes the
pending Send instruction.
The prepared After the component-spec review, successful packet preparation records the
existing In an 2. Implement itChat hands you Implement this spec as soon as a spec exists; it also sits in the Run view. The Run view opens when you start, so there is always something to watch. A run is not a dry run — it dispatches work to the coding agent and opens real pull requests. So it asks first, and tells you what it will build, where the code lands, each target-to-pipeline binding, ADO authentication readiness, and the retry budget it will use. A configured pipeline that cannot be authenticated or reached blocks Start. A target with no pipeline requires the explicit Start Without CI acknowledgement. Sign-in comes from VS Code's own GitHub and Microsoft accounts. There is no token to paste. 3. Watch the wavesThe Run view shows every wave as it moves through dispatch, AI review and CI: which unit it is on, how many retries it has spent, and links straight to the pull request and the build. When a wave goes red, the engine hands the failure back to the coding agent and lets it try again — up to your retry budget — before it stops and asks you. 4. Make the callWhen a wave needs a decision, you get a notification and the wave is marked in the view. The notification offers one thing: Review PR. The decision belongs after reading the code, not instead of it. Once you have read it, hover the wave for three buttons:
The fourth answer — not yet — is ⏸ Pause: Decide Later in the view's toolbar, where the ■ stop button sits the rest of the time. It is there and not on the row because pausing is not a per-pull-request act: one process holds the whole run, so pausing stops all of it. Decisions are per pull request, because a human judges this code; attachment is per run, because one process holds all of it. Decide Later is a real answer, not a way of giving up. Nothing merges without an explicit merge decision, and pausing puts no clock on you — take hours or days. It stops the whole run, though, so any other review open at the same time pauses with it. ▶ Resume the Run picks all of them up again — one click, however many are open — after re-checking the latest commits in case a pull request moved on while you were away. Nothing is re-dispatched, so a resume spends no new AI credits on work that already exists. The first three confirm before they act, and each records the decision on the pull request so the engine has one durable channel. The SEE Run view is the primary UI; manual decision comments remain only as a compatibility and recovery path. When the row says "engine on standby"The engine does not hold a process open while you think. It watches an open gate
for
and the pull request comment says the same thing. After that it steps back on its own, posts a note on the pull request, and the row reads:
Nothing has changed: the pull request is still open, nothing was merged or closed, and there is still no deadline. The ⏸ button disappears — the only thing missing is something to act on your answer, and ▶ Resume the Run appears in the view's toolbar to supply it. Press ✓ ↻ ✕ in the editor as usual. The extension records the decision, resumes the engine automatically, and applies it. If a compatible decision comment was left manually as a recovery action, it is queued, not lost, but it does not act while the engine is on standby. Resume first. If new commits landed, the engine asks again rather than merge code you never saw. The same is true if you stop the run yourself while a review is open: the engine leaves a note on the pull request saying nobody is watching it and directs the reviewer to the SEE Run view. Your machine stays awake while a run is in progressA wave can spend an hour waiting on the coding agent and CI, so a run is normally left alone. The coding agent and CI run on GitHub regardless, but the engine has to be awake to watch them — and on most current laptops, locking the screen puts the machine into Modern Standby within minutes, which throttles the engine and cuts its network. So the engine holds the machine awake for exactly as long as a run lasts. Two things follow:
Getting started
There is nothing to install first. The extension ships the engine inside it, and
the first time it needs Python it offers to set one up: it looks for a usable
interpreter you already have, and if there is none it installs Python 3.11 into
its own private folder and builds the environment there. Nothing lands on your
Windows, Linux and macOS are all supported, on Intel and ARM. On Linux and macOS the private interpreter is a relocatable CPython build, verified against the checksums published with it — so no package manager is involved and nothing is installed system-wide. Alpine and other musl-based distributions are out of scope: the Copilot CLI the engine drives has no musl build, so Python alone would not get you a working run. If a run is already going in a terminal, the view attaches to it automatically. Command PaletteSEE exposes only these extension commands in the Command Palette, regardless
of
Author through VS Code itself also adds Focus on Run View and View: Show Spec Execution Engine for the contributed view. Those navigation entries remain: the extension cannot hide them with Command Palette menu rules without disabling the view. Other command handlers are retained for existing buttons, walkthrough links and
custom keybindings. Diagnose Setup, Check AI Reviewer Sign-in, and
Rebuild Python Environment no longer have palette entries; hiding them reduces
troubleshooting discoverability. They can still be assigned custom keyboard
shortcuts using SettingsMost of these are filled in for you when you create a spec. You should rarely need to set one by hand. The walkthrough's Check configuration link walks through the essential ones; Configure is no longer a Command Palette entry. In Settings (
Grouping changes presentation only. All existing
On upgrade, SEE migrates the old If something does not work"No engine found." The extension carries its own copy, so this should not
happen — it means the install is damaged, or Python problems. An interpreter is only used if it can actually import what the engine needs, so one that exists but cannot run the engine is skipped instead of tried. If none qualifies, the extension offers to build its own. Diagnose Setup reports what it found — engine, interpreter, sign-ins — and Rebuild Python Environment throws the private one away and starts over, which is the fix if an install was interrupted. An opaque failure from Azure DevOps. A token from the wrong tenant signs in
successfully but is not authorized. Set Anything else. Show Engine Log has the engine's full output, including which identity and which interpreter were chosen. And Implement in a Terminal Instead runs the same engine, the same way, with the environment already set up — the view attaches to that run too. MCP startup diagnosticsThe Spec Execution Engine output channel has
The process records use the server PID and Activation and The refresh input snapshot decides whether to request an offer; it does not pin the identity used by a later offer. The current provider offer and its launch generation are authoritative after the barrier. A snapshot check that finishes after a newer refresh request is ignored. No auth cache is added. MCP wait limits and cancellationEach chat request has one 110,600 ms cumulative MCP wait limit. Preparation refresh, tool startup/recovery, the existing reconnect delay, and later reconnect attempts use the same remaining time. The clock runs only while the request waits in these MCP lifecycle operations. It pauses during other chat work and interactive tool invocation. Overlapping waits charge elapsed time once; nested calls cannot reset the limit. This is not a whole-chat wall-clock limit or an expected startup time. The preparation refresh remains at its existing call site, after runtime setup and sign-in. It does not run again on each tool lookup or interview question. All three tool paths use the same request scope. Settings-triggered refresh keeps its own bounded check. No time limit is added to Python installation, dependency setup, user sign-in, or interactive invocation.
The maximum branch totals 110,600 ms. It retains the earlier polling windows and uses the existing 15-second discovery allowance for the previously unbounded silent refresh inputs and restart command. It is a conservative ceiling, not a performance target. The one-second reconnect delay spends the same remaining total; it does not add a new allowance. Each phase also stops when the request's remaining total expires. Standalone activation has a 15,000 ms start limit. Standalone refresh has a 20,300 ms limit. A standalone tool lookup has a 90,300 ms recovery limit. Cancellation releases that caller promptly. Other active waiters retain their
own deadlines and shared work. After the last waiter leaves, a new caller does
not join the abandoned flight. A deadline or cancellation returns no candidate
( Refresh has explicit outcomes: Caller limit, not process cancellation: VS Code host commands, workspace discovery, interpreter probes, and silent identity lookups do not all expose a cancellation API here. The caller stops waiting, but those operations can still finish, update their own internal caches, or fail later. Their terminal handlers remain attached. Stopped caller continuations cannot request a later refresh, restart, or start, or update the current MCP lifecycle state. Provider callbacks remain host-owned; only callbacks from the current registration and matching launch generation can supply launch evidence. Refresh epochs also reject pre-refresh offers. This is not a general provider-completion ordering change. An already-issued host restart remains a safety barrier until its promise settles, even across provider registration replacement. Later callers must wait within their own remaining allowance; they cannot open an interview that the pending restart could then stop. Cancellation of a refresh after its event does not erase the fresh-offer or grace barrier. No host-owned process is killed. Timers require the extension-host event loop to run, so scheduler delay or synchronous blocking can make the measured return time exceed the nominal limit. Tests check elapsed time and return values with an explicit scheduler allowance, not just the presence of a timer. A start or restart command timeout stops that recovery attempt, even if
some cumulative time remains. The host command can still be active, so timeout
does not authorize a blind restart or forced refresh. This differs from an
ordinary Supported-host limit: VS Code 1.101 ignores This distinction follows the VS Code 1.101 provider update path and start command: the host adds definitions after the provider returns, and start completes without an error when the server ID is not yet in its list. The 1.101 launch resolver passes the original returned definition object to the provider, which supports the object-to-generation lookup. Live-host and platform-matrix checks remain separate from these source-contract and isolated-process tests. The ordering tests model separate provider-preparation, host-insertion, and process-initialization steps. The stdio smoke test uses a separate source-build process and no domain tool calls. It does not measure the host-owned MCP server:
The investigation baseline at Create a support bundleRun Spec Execution Engine: Create Support Bundle from the Command Palette when support needs diagnostics. You choose the destination in a save dialog. The extension creates one local ZIP file and does not upload, attach, share, or open it. Review the ZIP before you attach it to a support case. The bundle contains a manifest, a safe setup snapshot, the extension's bounded
rolling logs, a sanitized scaffold diagnostic log, the current valid status
snapshot, and at most three ledger rows. Each entry has a fixed name, a size
limit, a final SHA-256 hash, and an included, omitted, truncated, or error state
in The scaffold diagnostic log records form field names, actions, result codes, and
failure categories. It never records form answers, titles, prompts, generated
documents, argument values, or full scaffolder results. Older raw
The bundle never includes |