Markdown TTS
Hear when Claude Code needs you, and listen to your docs and code instead of reading them.
A text-to-speech extension for VS Code that gives your AI coding agent a voice and reads everything else aloud:
- Claude Code, out loud. A spoken heads-up, not just a chime, when Claude Code asks you a question, hits an error, or is waiting on a permission prompt. Or hear every reply. No API key and no extra AI call, since Claude Code already wrote the words. One-time setup
- GitHub Copilot can speak too. A
#speak tool that Copilot's agent mode can call to say a short summary aloud, when you ask it to. Setup
- Read markdown, code and docs aloud. Offline voices on Windows and macOS (SAPI /
say), or free Microsoft Edge neural voices, which are also how it reads aloud on Linux. Pause, skip between headings, change speed, or export to MP3.
- AI explanations, spoken. Hear a plain-English explanation of a file or selection, a summary of your git changes, or a short audio digest MP3 to listen to later. Bring your own OpenAI, Anthropic, Groq or Gemini key (Groq and Gemini have free tiers).
- Dictate into Copilot Chat with your OS's built-in voice typing (Windows and macOS).
Free, with no account, and core reading needs no API key. See what leaves your machine.
FAQ
Is this free? Yes, entirely. Core reading (file/selection/clipboard, offline SAPI/say, or online Edge neural voices) has no cost, no account, and no API key, and speaking Claude Code's or Copilot's replies needs no API key either. AI narration is optional and only runs when you invoke an AI command (Explain, Narrate Git Changes, or Export Audio Digest).
Do I need to pay for an API key to use AI narration? No — Groq and Google Gemini both have genuinely free tiers with no credit card required. Get a Groq key at console.groq.com/keys or a Gemini key at aistudio.google.com/apikey, and you're narrating in under a minute.
What does "Explain This File Aloud" actually do? It's not text-to-speech reading your code line by line — it sends the file to an LLM, gets back a plain-English explanation of what it does and why, and speaks that. Useful for unfamiliar code or dense docs you'd rather listen to than read.
Can it read Claude Code's or GitHub Copilot's replies aloud? Yes, with no API key. For Claude Code, turn on markdownTts.claudeCode.enabled and add the Stop hook that Markdown TTS: Show Claude Code Voice Hook Setup generates; by default it only speaks when Claude Code asks you something or is blocked. For Copilot, tell agent mode to call the #speak tool, inline or in .github/copilot-instructions.md. See Claude Code Voice Setup and GitHub Copilot Voice Setup.
What does "Narrate Git Changes Aloud" do, and is my code sent anywhere? It runs git diff locally, then sends only the diff text (not your whole repo, and capped at 12,000 characters) to whichever AI provider you've configured, and speaks back a summary of what changed. For Uncommitted changes, brand-new untracked files are included as "new file" diffs, so their full contents are sent too (up to 20 files; anything matched by .gitignore is skipped). Your code only goes to the AI provider when you explicitly run an AI command.
What leaves my machine? It depends on the voice engine and the command:
- Local engines (SAPI /
say): nothing. Speech is generated on your machine.
- Edge engine: whatever is being spoken or exported to MP3 (including AI explanations and Claude Code/Copilot replies) is sent to Microsoft's online text-to-speech service to generate the audio.
- AI commands (Explain, Narrate Git Changes, Export Audio Digest): the file, selection, or diff goes to the AI provider you configured, using your own key.
- Claude Code / Copilot voice: this extension makes no AI call of its own. Claude Code's
Stop hook writes the turn's hook data (including the reply text) to a file in your OS temp folder, which the extension reads locally; Copilot hands its text straight to the extension's #speak tool. The text only leaves your machine if you use the Edge engine.
- Voice input: handled entirely by your OS's own dictation (Windows Voice Typing / macOS Dictation), under the OS's privacy settings. This extension only starts it.
Does this work offline? Core reading (SAPI on Windows, say on macOS) is fully offline. Edge neural voices and all AI narration features require internet, and so does reading aloud on Linux, which uses the Edge voices.
Windows, macOS, or Linux? All three can read aloud. Windows and macOS also have an offline engine. Linux has no offline engine: set markdownTts.engine to "edge" (online) and have a command-line audio player installed (paplay, ffplay or mpg123). Linux playback has only been tested on Ubuntu 24.04 under WSL2, and voice input isn't supported there. See Platform Support below.
Requirements
- VS Code 1.95.0 or later (raised from 1.80.0 for the GitHub Copilot voice tool's Language Model Tool API, finalized in 1.95 — most users are on a newer version already via VS Code's own auto-update)
- Internet connection — only for the Edge Neural TTS engine and the AI commands
- Node.js on your PATH — only for the Claude Code voice hook, whose command runs
node
- Windows 10/11 with Voice Typing enabled
- Enable: Settings → Privacy & Security → Speech → Online speech recognition → On
- Test manually: press
Win+H anywhere — the Voice Typing toolbar should appear
- macOS with Dictation enabled
- Enable: System Settings → Keyboard → Dictation → On
- First run: macOS will prompt for Accessibility permission for VS Code — grant it
Linux
- Set
markdownTts.engine to "edge" (requires internet). The offline engine is Windows/macOS only
- Install one of these MP3-capable players:
paplay (pulseaudio-utils package), ffplay (ffmpeg package) or mpg123. The first one found on your PATH is used, in that order. If one fails (for example, a paplay too old to read MP3), the next is tried. If none is installed or none can play, you get an error saying which and why, rather than silence
paplay and mpg123 (run with -o pulse) play through a PulseAudio sound server. On PipeWire systems that is PipeWire's PulseAudio-compatible service, which hasn't been tested with this extension
- Export to Audio File and Export Audio Digest only save an MP3, so they need no player
- Voice Input is not supported
Install
Search "Markdown TTS" in the VS Code Extensions panel, or:
code --install-extension AbhishekShr.markdown-tts
Quick Start
- Open any markdown file
- Click the ▶ Read File button in the editor title bar — the file is read aloud (pause/resume and stop buttons appear there while reading)
- Or open the Command Palette (
Ctrl+Shift+P) and run any "Markdown TTS" command
- For voice input: run Markdown TTS: Voice Input to Chat → speak into Copilot Chat
That's it. For Edge neural voices, change markdownTts.engine to "edge" in Settings. Prefer keyboard shortcuts? See Running Commands to bind your own.
Features
Text-to-Speech
- Read File Aloud — speak the entire file, or from cursor position if placed mid-file
- Read from Cursor — explicitly read from cursor to end of file
- Read Selection Aloud — speak only highlighted text
- Read Clipboard Aloud — speak copied text (great for chat responses, browser content, etc.)
- Heading Navigation — skip to next/previous heading while reading
- Pause / Resume — temporarily pause and continue speech
- Stop Reading — stop the current speech immediately
- Edge Neural TTS — optional high-quality voices via Microsoft Edge's online service. Long files stream chunk-by-chunk so audio starts in seconds, not after a 2-minute wait
- Live chunk progress — status bar shows
Reading 1.2x 3/10 while streaming, plus a Buffering N/M… spinner when waiting on the next chunk
- Click-to-change reading speed — clickable
↑ / ↓ status bar buttons adjust speed live, with the new rate applied to upcoming chunks
- Export to Audio File — save any markdown file or selection as MP3 using Edge TTS
- List Edge Voices — browse 400+ voices in a searchable picker
- Reading time estimate — shows word count and estimated duration before reading long files
AI Narration (Bring Your Own Key)
- Explain This File Aloud (✨ button in the editor title bar) — instead of reading the file verbatim, an LLM explains what it does in plain English, then narrates the explanation. Great for unfamiliar code or dense docs.
- Explain Selection Aloud — right-click a highlighted block to explain just that part (falls back to the whole file if nothing is selected).
- Narrate Git Changes Aloud (git-compare button in the Source Control title bar) — runs
git diff locally and speaks an AI summary of your uncommitted changes (including brand-new untracked files, so a file an AI agent just created isn't skipped), last commit, or branch vs main. Perfect for a pre-commit review by ear.
- Export Audio Digest — turns a file or selection into a short, conversational spoken-word digest (not a line-by-line reading) and saves it as an MP3 you can take with you — notes on a commute, docs at the gym, whatever you'd rather listen to later than read now. Requires the Edge TTS engine (
markdownTts.engine: "edge"), since that's what makes the export.
- Speak Claude Code's replies aloud (
markdownTts.claudeCode.enabled, off by default) — hear Claude Code's own response after each turn, with no AI key needed since Claude Code already wrote the text. Run Markdown TTS: Show Claude Code Voice Hook Setup for the exact snippet to add to your Claude Code settings.json (a Stop hook) — this extension never writes to that file for you. Step-by-step in Claude Code Voice Setup. Only speaks in the VS Code window whose open folder matches the session that replied (including when that window is opened at an outer folder and the session runs in a project within it), so multiple windows don't all talk over each other.
markdownTts.claudeCode.mode — "smart" (default) only speaks when Claude Code is asking a question or flags it needs you (errors, blockers, permission, confirmation); routine progress stays silent. Set to "always" to hear every reply. Either way, whatever does get spoken is capped by markdownTts.voiceSummaryMaxChars (500 by default, at a sentence boundary) — smart mode decides whether to speak, this decides how much, so a long detailed reply that happens to end in one question doesn't get read in full just because that question triggered narration. Raise the setting (or set it very high) if you'd rather hear longer or full-length replies.
markdownTts.claudeCode.alerts (default on) — a short spoken alert via Claude Code's separate Notification hook for when it's waiting on you specifically: a permission prompt, an idle timeout, or similar. This is independent of claudeCode.mode above — a permission prompt isn't new reply text for the smart/always filter to judge, it's Claude Code telling you directly that it's stuck. Same one-time setup as claudeCode.enabled.
- Speak GitHub Copilot's replies aloud — Copilot has no hook system like Claude Code's, so instead this extension registers a Speak Aloud tool (
#speak) that Copilot's agent mode can call. It won't call it on its own; tell it to, either inline ("...and speak a summary aloud when you're done") or once in .github/copilot-instructions.md (e.g. "When you finish a task, call the speak tool with a short summary."; a ready-to-paste version is in GitHub Copilot Voice Setup). Same no-API-key, no-extra-AI-call design as the Claude Code path, and the same markdownTts.voiceSummaryMaxChars sentence-aware cap (500 by default) — Copilot also decides what text to send, and its tool description tells it to send a brief summary rather than the full response, but the cap is what's actually enforced regardless. Requires VS Code 1.95+ (the Language Model Tool API); on older versions this is silently skipped rather than breaking the rest of the extension.
- Bring your own key — works with OpenAI, Anthropic (Claude), Groq, or Google Gemini (Groq and Gemini both have free tiers — no credit card). Your key is stored in VS Code's encrypted SecretStorage, never in settings. Set it with Markdown TTS: Set AI API Key.
- If narration suddenly stops working: free-tier providers retire model names often (this happened to Groq's default in mid-2026). Set
markdownTts.ai.model to a current model ID for your provider — see Groq's models or Gemini's models — no reinstall needed.
- Read-along output — the generated explanation is also written to the Markdown TTS – AI output panel so you can follow the text while it speaks.
- ✨ Sparkle button in the editor title bar for one-click file explanation.
Voice Input (Speech-to-Text)
- Voice Input to Chat — triggers OS-native dictation directly into Copilot Chat
- Run Markdown TTS: Voice Input to Chat → Copilot Chat opens → dictation starts → speak → text streams into chat
- Windows: launches Win+H Voice Typing (press
Esc to stop)
- macOS: triggers Edit → Start Dictation (press
Fn to stop)
- Pre-warmed PowerShell process for near-instant launch on Windows
Smart Markdown Processing
- Strips YAML frontmatter, HTML tags, code blocks, footnotes
- Converts headings to "Heading level N: ..."
- Reads image alt text as "Image: alt text"
- Tables read as comma-separated values (wide tables truncated to 5 columns)
- Removes bold/italic/link syntax, keeps the text
- Pronunciation dictionary — custom replacements for tech terms TTS engines mispronounce
- Works in markdown preview mode and with unsaved/untitled files
Running Commands
Every command is available from several places — pick whichever is fastest:
- Command Palette —
Ctrl+Shift+P (Cmd+Shift+P on macOS) → type "Markdown TTS" to see them all.
- Editor title bar buttons — ▶ Read File, ✨ Explain File, 📋 Read Clipboard, plus pause/resume/stop and heading navigation while reading.
- Right-click menu — Read Selection, Explain Selection, Read/Explain File, Export to Audio, and more.
- Source Control title bar — the git-compare button runs Narrate Git Changes Aloud.
- Status bar — while reading, click the
↑ / ↓ buttons to change speed live, or the reading indicator to pause/resume.
Want keyboard shortcuts?
This extension ships no default keyboard shortcuts on purpose: the natural choice, Ctrl+Alt+<letter>, collides with AltGr on many non-US keyboard layouts (Indian, German, French, Spanish, Nordic, and others), where the OS types an accented character instead of running the command. Rather than hijack keys that break for a large share of users, you bind exactly the ones you want:
- Open Keyboard Shortcuts:
Ctrl+K Ctrl+S (Cmd+K Cmd+S on macOS).
- Search for "Markdown TTS".
- Click the
+ next to a command and press your preferred key combo.
For layout-safe shortcuts, Ctrl+K chords (e.g. Ctrl+K R) or plain Ctrl+Shift+<letter> combos avoid the AltGr problem.
Settings
{
"markdownTts.engine": "sapi",
"markdownTts.rate": 2,
"markdownTts.voice": "",
"markdownTts.edgeVoice": "en-US-AriaNeural",
"markdownTts.replacements": null,
"markdownTts.ai.provider": "openai",
"markdownTts.ai.model": "",
"markdownTts.claudeCode.enabled": false,
"markdownTts.claudeCode.mode": "smart",
"markdownTts.claudeCode.alerts": true,
"markdownTts.voiceSummaryMaxChars": 500
}
engine — "sapi" (local system voice, offline, default) or "edge" (Microsoft Edge neural voices, online)
rate — Speech rate: -10 to 10. Default 2. Maps to SAPI rate on Windows, WPM on macOS, percentage on Edge.
voice — Voice name. Windows: SAPI voice (e.g. "Microsoft David Desktop"). macOS: say voice (e.g. "Samantha"). Leave blank for system default.
edgeVoice — Edge voice name. Default "en-US-AriaNeural". Run Markdown TTS: List Edge Voices to browse.
replacements — Custom pronunciation dictionary. Object mapping text → replacement (e.g. {"npm": "en pee em"}). Applied before speaking.
ai.provider — AI narration provider: "openai", "anthropic", "groq", or "gemini". Default "openai".
ai.model — Override the model. Leave blank for a fast, low-cost default (gpt-4o-mini for OpenAI, claude-haiku-4-5 for Anthropic, openai/gpt-oss-20b for Groq, gemini-2.5-flash for Gemini).
claudeCode.enabled — Speak Claude Code's replies aloud via its Stop hook (no AI key needed). Default false. Needs a one-time hook setup, see Claude Code Voice Setup.
claudeCode.mode — "smart" (default) only speaks when Claude Code asks a question or flags it needs you (errors, blockers, permission, confirmation); "always" speaks every reply. Either way, spoken text is capped by voiceSummaryMaxChars below. Only applies when claudeCode.enabled is on.
claudeCode.alerts — Default true. Speaks a short fixed alert via Claude Code's Notification hook when it's waiting on you (permission prompt, idle timeout, etc.) — independent of claudeCode.mode, since this isn't reply text for that filter to judge. Same one-time hook setup as claudeCode.enabled.
voiceSummaryMaxChars — Default 500. How long a spoken summary can be, in characters, for Claude Code's replies and Copilot's #speak tool alike — cut at a sentence boundary where possible. Short is this extension's own preference, not a hard rule: raise it (or set it very high) if you'd rather hear longer or full-length replies.
AI Narration Setup (Optional)
The AI commands (Explain, Narrate Git Changes, Export Audio Digest) need an API key — your own, so there's no subscription. Two free options, no credit card:
- Get a free key at console.groq.com/keys (Groq) or aistudio.google.com/apikey (Gemini).
- In Settings, set
markdownTts.ai.provider to "groq" or "gemini" to match.
- Run Markdown TTS: Set AI API Key (Command Palette) and paste the key. It's stored in VS Code SecretStorage — never in your settings file.
- Open any file and click the ✨ button in the editor title bar (or run Markdown TTS: Explain This File Aloud from the Command Palette).
Prefer OpenAI or Anthropic? Set the provider accordingly and paste that key instead. Only those AI commands ever contact the AI provider; nothing else does (see What leaves my machine? above).
Claude Code Voice Setup
Hear Claude Code's replies spoken aloud, and get a short alert when it's waiting on you. There's no API key and no AI call: Claude Code already wrote the text, and this extension just reads it out. One-time setup:
- In VS Code Settings, turn on
markdownTts.claudeCode.enabled (speaks replies), markdownTts.claudeCode.alerts (speaks "waiting on you" alerts, on by default), or both. Both take effect immediately, no reload needed.
- Run Markdown TTS: Show Claude Code Voice Hook Setup from the Command Palette. It opens an unsaved JSONC tab with ready-made
hooks → Stop and Notification blocks, with the hook command already filled in for this machine — Stop covers replies, Notification covers the "waiting on you" alerts.
- Open your Claude Code settings file, either
~/.claude/settings.json (all projects) or .claude/settings.json (one project), and merge in that hooks block. If the file already has a "hooks" key, add the Stop and Notification entries to it rather than replacing your existing hooks. This extension never edits that file for you.
- Start a new Claude Code session (restart any that were already running) so the hooks are picked up.
- Run Claude Code from the folder you have open in VS Code, or any folder inside it. Only a VS Code window with that folder (or a parent of it) open will speak, so unrelated windows stay quiet.
In the default "smart" mode you'll only hear replies where Claude Code asks you a question or flags that it needs you (an error, a blocker, a permission or confirmation request). Routine progress stays silent. To check the setup works, temporarily set markdownTts.claudeCode.mode to "always" so every reply is spoken. Either way, spoken text is capped by markdownTts.voiceSummaryMaxChars (500 by default, cut at a sentence boundary where possible) — raise it if you want longer replies.
Separately, claudeCode.alerts (default on) speaks a short fixed phrase whenever Claude Code's own Notification hook fires — a permission prompt, an idle timeout, or similar — regardless of what claudeCode.mode is set to, since there's no reply text involved for that filter to judge in the first place. This is the one case Stop-based narration structurally can't catch on its own.
Good to know:
- The hook command runs
node, so Node.js must be on your PATH.
- All the hook does is write the turn's hook data (including the reply text, for
Stop) to markdown-tts-claude-code-message.json in your OS temp folder, and the extension watches for that file. The snippet contains this machine's temp-folder path, so generate it on each machine rather than copying it between machines.
- To stop narration, turn
markdownTts.claudeCode.enabled and/or .alerts off. The hooks keep writing their temp file each turn until you also remove them from your Claude Code settings.
- On Linux, set
markdownTts.engine to "edge" and install one of the audio players listed under Requirements → Linux.
GitHub Copilot Voice Setup
Copilot has no hook system, so this extension gives Copilot's agent mode a Speak Aloud tool (#speak) to call instead. Copilot won't call it unprompted; you tell it when:
- Use VS Code 1.95 or later with GitHub Copilot Chat in agent mode.
- Ask inline, e.g. "...and #speak a one-sentence summary when you're done", or make it standing behaviour by adding an instruction like this to
.github/copilot-instructions.md in your repo:
## Spoken summaries
When you finish a task in agent mode, call the Markdown TTS speak tool
(`markdown_tts_speak`) with a one- or two-sentence plain-English summary of
what you did and anything you need from me. Never read code, file paths, or
your full response aloud, and don't call it for routine progress updates.
Whatever text Copilot sends is capped by markdownTts.voiceSummaryMaxChars (500 by default, cut at a sentence boundary where possible), so even if it ignores the "short" part you won't hear an essay — unless you've deliberately raised that setting. As with Claude Code, there's no API key and no AI call on this extension's side.
| Feature |
Windows |
macOS |
Linux |
| Local TTS engine |
SAPI via PowerShell |
say command |
— |
| Edge Neural TTS (live playback) |
✅ |
✅ |
✅ ¹ |
| Pause / Resume |
✅ |
✅ |
✅ ¹, long pauses: see ² |
| Voice Input |
Win+H Voice Typing |
Edit → Start Dictation |
— |
| Heading Navigation |
✅ |
✅ |
✅ ¹ |
| Reading Time Estimate |
✅ |
✅ |
✅ ¹ |
| AI narration (Explain / Narrate Git Changes) |
✅ |
✅ |
✅ ¹ |
Speak Claude Code's replies (Stop hook) |
✅ |
✅ |
✅ ¹ |
Speak Copilot's replies (#speak tool) |
✅ |
✅ |
✅ ¹ |
| Audio Export (MP3) |
✅ (Edge) |
✅ (Edge) |
✅ (Edge) |
| Export Audio Digest (AI → MP3) |
✅ (Edge) |
✅ (Edge) |
✅ (Edge) |
¹ Linux needs the Edge engine (markdownTts.engine: "edge", online) and one of paplay, ffplay or mpg123 installed (see Requirements → Linux). Tested on Ubuntu 24.04 under WSL2, through WSLg's PulseAudio server, with each of the three players. Playback, pause/resume, stop, the reading-time prompt, the Claude Code hook and the Copilot #speak tool were each run end to end there, by driving the extension's own code outside VS Code. Heading navigation and AI narration play through the same path but weren't run on Linux themselves. Not yet run inside VS Code on Linux, on other distributions, or with PipeWire or plain ALSA. Saving an MP3 needs no player, which is why both export commands work everywhere.
² On Linux, Pause freezes the player (SIGSTOP). Under WSL2/WSLg, a 30-second pause resumed normally with each of the three players. After a 90-second pause, though, the resumed audio stalled, and WSLg's audio sink logged queue overruns. Press Stop if that happens, then read again from the cursor. The same 90-second pause resumed normally on a regular PulseAudio sink, so this looks specific to WSLg. Long pauses on a Linux desktop haven't been tested.
Architecture
extension.js — single-file entry point (~2,000 lines)
- Windows SAPI: spawns PowerShell with
System.Speech.Synthesis.SpeechSynthesizer
- macOS say: spawns
say command, pause/resume via Unix signals
- Edge engine: splits text into ~1.5 KB chunks; prefetches 2 chunks ahead via parallel WebSockets; plays chunks sequentially through platform-native player (
afplay on macOS, WPF MediaPlayer on Windows, the first installed of paplay / ffplay / mpg123 on Linux, moving on to the next one if a player fails). 30 s per-chunk timeout with one retry.
- Claude Code voice: the
Stop hook runs a one-line node -e script that writes the hook's JSON input to markdown-tts-claude-code-message.json in the OS temp folder. The extension fs.watches that folder, matches the payload's cwd against this window's workspace folders, applies the smart/always filter and the voiceSummaryMaxChars cap, then speaks it.
- Copilot voice: a Language Model Tool (
markdown_tts_speak, referenced as #speak) registered with vscode.lm.registerTool. It's feature-detected, so a VS Code build without that API just skips it. The text Copilot passes is capped and spoken through the same pipeline.
- Voice Input: pre-warmed PowerShell sends Win+H via
keybd_event (Windows) or AppleScript triggers Dictation menu (macOS)
- Pause/resume/stop: control file on Windows, signals (
SIGSTOP/SIGCONT, and SIGKILL to stop) on both macOS and Linux
- Title bar buttons and status bar indicator for visual control
Troubleshooting
| Problem |
Solution |
| Voice Input: nothing happens |
Enable OS dictation (see Requirements above) |
| Edge TTS: "Timed out" |
A single chunk failed twice; check internet. Long files split into ~1.5 KB chunks (30 s timeout each) so total reads no longer block on a 2 min ceiling. |
| Edge TTS: "Cannot find module" |
Reinstall VSIX — dependencies may not have been bundled |
| No sound on Windows |
Check default audio output device; try SAPI engine |
| No sound on macOS |
Try say "hello" in Terminal; check volume |
| No sound on Linux |
Set markdownTts.engine to "edge". If the error says no audio player was found, install paplay (pulseaudio-utils package), ffplay (ffmpeg package) or mpg123. If it lists players that failed, check your sound server is running (pactl info should answer). |
| Claude Code voice: nothing is spoken |
Check markdownTts.claudeCode.enabled is on and the Stop hook is in your Claude Code settings. Run Claude Code from the folder open in VS Code (or one inside it), and make sure node is on your PATH. In "smart" mode routine replies are silent on purpose, so set markdownTts.claudeCode.mode to "always" to test. |
| Copilot never speaks |
It only calls #speak when asked. Use agent mode and add the instruction from GitHub Copilot Voice Setup. |
Limitations
- Linux: no offline engine, so reading aloud needs the Edge engine (internet) plus
paplay, ffplay or mpg123. Only tested on Ubuntu 24.04 under WSL2, where resuming after a long pause can stall (see Platform Support). No voice input
- Edge engine requires internet — text is sent to Microsoft servers
- Voice Input requires OS-level dictation to be enabled
- Copilot voice is up to Copilot: you can tell it when to call
#speak, but the model decides whether it actually does
- Markdown stripping is regex-based, not a full parser
- Mid-stream speed change has ~1 min delay (Edge engine). The currently-playing chunk continues at its original rate; the new rate kicks in when the next chunk starts. With ~1.5 KB chunks (about 60–90 s of audio), expect up to a minute before the change is audible. SAPI/macOS engines apply the new rate only on the next read.
| |