WrengleWrengle
Dictation and voice control

Dictation and voice control

Beta

Wrengle has two voice modes, selected in Settings → Dictation:

  • Quick Dictation (Recommended) treats ordinary utterances as text for the focused note or assistant draft. Trigger words do not run commands in this mode; only exact reserved session and correction phrases are controls.
  • Voice Control dictates bare speech and uses explicit trigger words for editor, workspace, configured-agent, and terminal commands.

Voice Control remains the persisted mode default. Dictation cleanup defaults to Light edit, and response speed defaults to Fast; choose Quick Dictation when you want a dictation-only path.

Speech recognition and voice analysis both default to Auto. When a connected provider key is confirmed available in the OS keychain, Auto uses a cloud route without requiring a second opt-in; without one, it falls back to the local route, which may ask you to download the required model. This availability check does not contact the provider, so invalid, revoked, or quota-limited keys can still fail when a request starts. Local speech recognition and local analysis keep audio and analysis context on your machine. Cloud dictation sends microphone audio to OpenAI, Deepgram, or ElevenLabs according to the route-specific boundaries described below. Cloud analysis sends transcribed utterance text plus bounded workspace context: the active note title, up to 40 folder paths, up to 15 recent note paths, and—for cleanup—the previous dictated sentence. It never sends microphone audio to the analysis provider.

Auto does not override an explicit choice. Selecting Local, a specific provider, or turning Cloud voice features off remains sticky when keys are added, replaced, or removed. Resetting the speech or analysis preference returns that surface to Auto.

Starting a voice session

Start or stop the selected mode from the Dictate or Voice control in the status bar, or with ⌘⇧V (macOS) / Ctrl+Shift+V on other platforms. For push-to-talk-style capture, hold ⌘⇧D / Ctrl+Shift+D and release it when you finish. Hold to Dictate always starts Quick Dictation for that hold, regardless of the saved mode.

While a session is active, the heads-up panel shows its state, normalized microphone level, device errors, live transcript, recent results, and the current destination, such as a note name or Assistant draft. If it says No note focused, focus an editable note or the assistant composer before speaking.

Stopping with the button, toggle shortcut, or Hold to Dictate key release commits the current partial and pending text even when live note preview is off. Wrengle briefly stays in a stopping/tidying state while it finishes the selected cleanup and inserts the words. In either mode, a whole-utterance spoken stop command ends the session without committing an unfinalized live partial, so use an explicit button or shortcut stop when preserving the latest partial matters.

If insertion is blocked because a confirmation, agent preview, missing note focus, assistant composer focus change, or dictation target change owns the destination, the voice panel keeps the protected pending text in an expandable recovery card. From there:

  • Insert here inserts it into the currently focused writable destination.
  • Return to original reopens the original note when the pending owner was an editor; focus its editor, then choose Insert here.
  • Copy places the pending text on the clipboard.
  • Discard clears the protected text without inserting it.

Pending recovery text is scoped to its vault and is discarded when you switch vaults, so it cannot be displayed or inserted into a same-named note in another vault.

After successful note dictation, Undo last phrase removes the most recently inserted dictated phrase when that note is still the valid undo target.

Voice and meeting recording share the same capture path, so a voice session cannot start while a meeting is recording or finalizing, and vice versa.

Speech recognition

In-app dictation has its own speech setting under Settings → Dictation → Dictation speech engine. It is separate from the meeting transcription model.

ChoiceCurrent dictation modelData boundary
Auto (default)Most recently connected or successfully verified speech key confirmed available in the OS keychain; otherwise OpenAI → Deepgram → ElevenLabs → Local WhisperDepends on the resolved route shown in Settings
Local WhisperAuto (fastest installed) or the selected installed Whisper rungMicrophone audio stays on the device
OpenAIRealtime WhisperSpeech-gated microphone audio streams to OpenAI
DeepgramNova-3 streamingMicrophone audio streams to Deepgram
ElevenLabsScribe v2 RealtimeMicrophone audio streams to ElevenLabs

Auto first uses the speech provider whose key was most recently connected or successfully verified as readable in the OS keychain. If that provider is unavailable, or an older installation has no preference history, it checks connected keys in the fixed order OpenAI, Deepgram, then ElevenLabs before falling back to Local Whisper. An unknown or inaccessible key does not qualify.

Cloud choices require the matching transcription key under Settings → Dictation → Speech keys and Cloud voice features access in Settings → Privacy & Data. Cloud voice access defaults on when no previous preference exists, so saving a connected key is enough for Auto to use it. A previously explicit Off remains off and forces automatic speech and analysis routes local; a specific cloud selection remains selected but unavailable until access is restored.

Speech keys are separate from AI-chat keys. Wrengle requests Deepgram's model-improvement opt-out, which Deepgram says can affect pricing. ElevenLabs realtime dictation uses the logging and retention behavior available to your ElevenLabs account; Wrengle does not claim enterprise Zero Retention Mode for ordinary BYOK accounts. Provider retention, billing, and account policies may still apply.

Turning cloud access off stops active capture promptly; audio or analysis already captured and in flight may finish during bounded finalization.

Deepgram and ElevenLabs receive microphone audio while their cloud dictation streams are active. OpenAI dictation remains speech-gated on the device: Wrengle suppresses leading ambient silence, keeps up to 240 ms of audio before an on-device speech detection so the first sound is not clipped, sends that bounded pre-roll with the detected speech window, and commits windows manually. Response-speed tuning does not turn the OpenAI route into a full ambient-session stream.

When a cloud dictation stream disconnects, Wrengle keeps the session open, leaves the last live ghost text visible, shows a reconnecting state, buffers recent microphone audio, and retries automatically. If the buffer overflows, Wrengle warns that some speech may not be transcribed.

Dictation language defaults to English and offers Auto detection, English, English (US), English (UK), Spanish, French, German, Italian, Portuguese, Portuguese (Brazil), Dutch, Polish, Japanese, Korean, Chinese, and Hindi. The setting controls speech recognition; Wrengle's UI, built-in control phrases, and command examples remain English in this release. An English-only local Whisper rung cannot be used for a non-English or Auto selection, so install a multilingual rung for those choices.

Personal vocabulary accepts names, acronyms, and specialist terms as bounded recognition hints. It improves the speech engine's context but is not a guaranteed replacement rule. Local Whisper uses the hints on your machine; Deepgram and ElevenLabs receive their supported key-term form when selected. ElevenLabs' current keyterm guide allows up to 50 realtime keyterms of at most 20 characters and lists an additional keyterm charge; Wrengle sends only entries that fit that provider limit. OpenAI Realtime Whisper does not receive personal-vocabulary hints in this release.

Voice analysis

Voice analysis & polish model also defaults to Auto. It prefers the AI provider whose key was most recently connected or successfully verified as readable in the OS keychain. If that route is unavailable, Auto checks OpenAI, then Anthropic, using that provider's catalog default model—GPT-5.6 Luna for OpenAI or Claude Sonnet 5 for Anthropic—and finally falls back to a selectable Wrengle Local model. Current and Previous model groups remain available for an explicit cloud choice; an existing explicit choice is not replaced by the new OpenAI default.

Choosing Local or a specific provider disables Auto for voice analysis and remains sticky. An explicit cloud choice with a missing key reports the missing credential instead of silently asking for a local-model download. Quick Dictation with Verbatim cleanup does not use this setting; Voice Control needs it for commands, and Light edit needs it for cleanup.

Microphone and voice setup

Settings → Dictation → Microphone and voice setup combines the selected speech engine's readiness with microphone controls:

  • Dictation microphone chooses a listed input or System default microphone. Use Refresh microphones after connecting or removing hardware.
  • Test microphone starts a bounded, model-free five-second level test on the selected input. It does not start speech recognition, analysis, or a cloud provider request. The normalized meter and device label show whether Wrengle is receiving input; choose Stop test to finish early.
  • Speech engine ready means the selected or Auto-resolved local or cloud recognition path is ready. Speech engine needs setup points back to the missing local model, provider key, or disabled cloud voice access.

A test and a dictation or meeting recording cannot own capture at the same time. If a specifically selected microphone disappears, Wrengle reports the device error instead of silently switching inputs; reconnect it, choose another input, or return to System default microphone. Because the current cross-platform audio layer does not expose stable hardware IDs, identically named microphones are omitted from persistent selection rather than risk silently choosing the wrong one; use the system default or leave only one of those devices connected.

Quick Dictation

Quick Dictation routes ordinary recognized speech directly to the current destination. editor, workspace, agent, and terminal are ordinary dictated words, and no command-classification step can turn bare prose into an action. This is the simplest path for note-taking and assistant prompt drafting.

Exact whole-utterance stop listening, stop dictating, and stop voice phrases end the session without trigger or LLM command analysis. Embedded occurrences remain literal. Two additional exact correction phrases, scratch that and undo that, remove the latest eligible note dictation. If there is nothing eligible to undo, Wrengle reports that instead of deleting other content.

Spoken punctuation is an optional Quick Dictation setting and is off by default. When enabled, the exact standalone phrases comma, period, full stop, question mark, colon, semicolon (or semi colon), open quote, close quote, and new paragraph insert that punctuation or break. Wrengle does not scan for these phrases inside a longer utterance, so discussing “a question mark” in a sentence does not silently rewrite the sentence.

Voice Control

Every command requires a trigger word — a word that tells Wrengle which area of the app you are addressing. Speech without a trigger is treated as plain prose dictation for the active voice target.

What you sayWhat happens
Trigger word + commandRuns the command in that area
Bare speech (no trigger)Dictated as plain prose into the open note, or into the assistant composer when that composer is focused
"stop listening", "stop dictating", or "stop voice"Ends the Voice Control session when spoken as the whole utterance
A confirmation replyResolves a pending confirmation without a trigger

While a confirmation card is visible, affirmative replies are yes, yeah, yep, confirm, do it, ok, and okay. Negative replies are no, nope, cancel, and never mind. A choice card also accepts first through fourth or a visible candidate name. These are whole confirmation replies, not keywords pulled out of ordinary dictation.

Shell-command and hosted/external-agent cards are stricter: approval requires clicking the visible button (or activating it from the keyboard). Spoken approval is rejected so the microphone cannot approve its own boundary-crossing request; spoken cancel still works.

When the assistant composer is focused, dictation drafts text in the composer and never sends it automatically. You can review or edit the dictated prompt, then send it with Enter or the Send button. In that focused composer context, trigger words become drafting prefixes instead of commands: saying "agent summarize this" drafts "summarize this", and saying "terminal run tests" drafts "run tests" instead of running a terminal command. Say "literally" before a trigger word if you want the trigger word itself in the draft.

Voice commands

Speak a trigger word, then describe the action. The four trigger areas are editor, workspace, agent, and terminal. Trigger words are editable under Settings → Dictation → Voice trigger words. Alongside the global trigger words, you can set context-specific trigger words there too — for example, a different word for the editor than for the terminal — so each area can answer to the wording that feels natural to you.

editor

Structures the open note without typing.

  • Headings: "editor heading two Project Goals" → level-2 heading containing "Project Goals".
  • Lists: "editor bullet list milk eggs bread" → three-item bullet list. Items split on natural pauses, "and", or "next"; "editor numbered list" and "editor checklist" work the same way.
  • Blocks: "editor quote", "editor divider", "editor code block", "editor callout", "editor task", and more — the same blocks available in the slash menu.
  • Breaks: "editor new line" for a line break; "editor new paragraph" for a new block.
  • Shape: "editor indent" / "editor outdent" to nest; "editor make this a heading two" to convert the current block; "editor delete block"; "editor undo".

Format and refine text

Apply inline formatting and make small corrections to what you have written, all without leaving voice.

  • Format what you just dictated: "editor bold that" applies bold to the text you most recently spoke into the note. "editor italic that", "editor strikethrough that", and "editor code that" work the same way.
  • Format named words: "editor bold the word budget" emphasizes a single word; "editor italic from quarterly to review" formats the run of words between two anchors you name.
  • Clear formatting: "editor clear formatting" removes bold, italic, strikethrough, and inline code from the current selection or the words you name.
  • Refine the last few words: "editor delete last word" removes the word you just spoke; "editor delete last sentence" removes the sentence.

Every formatting and refinement command is reversible — say "editor undo" (or use the editor's own undo) to step back. The same bold, italic, strikethrough, and inline-code styles are also available as buttons in the editor's formatting toolbar when you would rather format by hand.

To dictate a trigger word as literal text rather than a command, say "literally" before it — for example, "literally editor".

workspace

Controls notes, folders, navigation, and app settings.

  • Notes: "workspace new note called Budget" · "workspace open my budget note" · "workspace rename this note to Q3 Budget" · "workspace delete the meeting note" (asks to confirm).
  • Folders: "workspace new folder called Projects" · "workspace delete the archive folder" (asks to confirm).
  • Search: "workspace search quarterly review" · "workspace search tax documents".
  • App actions: "workspace toggle terminal" · "workspace toggle assistant" · "workspace focus mode" · "workspace open settings" · "workspace theme dark" (also light, system).

agent

Hands a request to the assistant backend currently selected or resolved by Auto in Wrengle. That backend can be Wrengle Local, a built-in provider-backed assistant, or a configured external ACP agent; its normal context, permission, provider, authentication, and billing boundaries still apply.

Outside the focused assistant composer, the agent trigger uses the resolved assistant backend. A Wrengle Local request can run immediately; a built-in cloud or external agent shows the complete request in a confirmation card and requires button or keyboard approval before it is sent. When the assistant composer itself is focused, the same trigger is treated as a dictation prefix for the draft, so you can speak naturally into the prompt before sending.

  • "agent summarize this note"
  • "agent write a packing list for the trip"
  • "agent rewrite this section more concisely"

terminal

Controls the built-in terminal. A run command requires an already-running terminal, rejects hidden/control characters, and shows the exact command in a confirmation card. Only button or keyboard approval writes it to the shell; spoken approval cannot execute it.

  • Shell writes (button/keyboard confirmation required): "terminal run npm test" · "terminal run git status" · "terminal clear".
  • Session control: "terminal open" · "terminal close" · "terminal toggle".
  • Tab navigation: "terminal next tab" · "terminal previous tab".

Safety tiers

Voice actions have different safety boundaries:

  • Read-only and view actions (open, search, toggle panels, change theme) run immediately.
  • Creating and renaming notes or folders runs immediately and shows an Undo action so a mistaken create or rename is one step away from being reversed. You can also say "undo" to reverse the last voice action.
  • Deleting a note or folder always asks for confirmation first. Wrengle shows a confirmation card naming the target, and you confirm by saying "yes" or "cancel" (or by clicking). A pending confirmation must be resolved before another one can replace it.
  • Agent requests go immediately only when the resolved assistant backend is Wrengle Local. Built-in cloud and external agents require confirmation and may receive request and note context according to that backend's normal privacy boundary; voice Undo cannot recall a sent request.
  • Terminal run commands require an already-running terminal and confirmation, then execute in the embedded local shell with normal OS permissions. They can change files, contact services, or cause other side effects and are not made reversible by voice Undo.

Cleanup and response speed

Dictation behavior is configured under Settings → Dictation:

  • Dictation cleanup → Verbatim keeps the recognized wording and only removes speech/pause artifacts.
  • Dictation cleanup → Light edit removes fillers and false starts and fixes casing and punctuation while preserving the intended wording. Light edit uses the selected voice analysis model, so a cloud analysis choice sends transcribed text to that provider.
  • Response speed → Fast, Balanced, or Deliberate controls end-to-end phrase timing for every supported dictation engine: Local Whisper, OpenAI, Deepgram, and ElevenLabs. Fast is the default and commits after shorter pauses, so it can split speech more readily; Balanced allows a little more thinking time, and Deliberate leaves the longest pauses within a thought.
  • Dictation language and Personal vocabulary guide speech recognition; Spoken punctuation enables exact Quick Dictation punctuation controls.
  • Show live dictation preview controls whether words appear as grey ghost text at the note caret. When off, the volatile preview remains in the voice panel instead.

Quick Dictation with Verbatim cleanup bypasses the analysis model completely: it does not load, warm, or call a local or cloud LLM. Voice Control still needs analysis for commands, and Light edit needs analysis for cleanup.

Grey note preview text is not saved note content until it settles. Complete thoughts can settle sooner than dangling fragments such as a phrase ending in “and” or “because,” so partial thoughts can continue naturally.

On an explicit button or shortcut stop, Wrengle freezes the pending words, applies the selected cleanup, and inserts the result before the voice session returns to idle. If cleanup takes too long or fails, Wrengle inserts a de-artifacted raw version instead of dropping the words.

If you switch notes, move focus to another editor, or leave one assistant composer for another before the live preview settles, Wrengle does not insert that pending preview into the newly active surface.

Cleanup does not reshape your words into layout, and markdown-looking dictation stays visible as literal text. Saying - item, # heading, bold, italic, code, or label without a trigger inserts those characters as prose instead of creating a list, heading, bold or italic text, inline code, or a link. To structure the document or apply formatting — new paragraphs, headings, lists, bold, italic, inline code — use voice commands with a trigger word. With Quick Dictation's spoken-punctuation option off (the default), saying "new paragraph" types those words; when the option is on, that exact standalone phrase creates a break. In Voice Control, saying "editor new paragraph" creates one. Saying "bold budget" as plain prose types those words; saying "editor bold budget" applies inline bold formatting.

Privacy and requirements

The speech-recognition and analysis choices are independent:

Dictation speechVoice analysisRequirementsWhat leaves the device
LocalLocalMicrophone access, an installed local Whisper model, and an installed selectable Wrengle Local modelNothing for speech recognition or analysis
LocalCloudMicrophone access, local Whisper, a supported AI-provider key/model, and cloud voice accessTranscribed utterance text and bounded workspace context go to the AI provider
CloudLocalMicrophone access, the selected speech-provider key, cloud voice access, and an installed selectable Wrengle Local modelMicrophone audio goes to the speech provider
CloudCloudMicrophone access, separate speech and AI-provider keys/models, and cloud voice accessMicrophone audio goes to the speech provider; transcribed utterance text and bounded workspace context go to the AI provider

For cloud analysis, bounded workspace context means the active note title, at most 40 folder paths, at most 15 recent note paths, and the previous dictated sentence when Light edit needs cross-sentence cleanup. Note bodies are not part of this voice-analysis context.

Quick Dictation with Verbatim cleanup is the exception to the analysis column: it needs only the selected local or cloud speech path and does not require an analysis model. Voice Control requires analysis even with Verbatim cleanup.

Auto can resolve each column independently. A connected ElevenLabs transcription key and a connected OpenAI AI key, for example, route microphone audio to ElevenLabs and the resulting bounded text context to OpenAI. Cloud voice features Off forces both automatic voice routes local but does not disable the separately resolved assistant. Meeting live captions and meeting transcription keep their existing local-first defaults; cloud meeting final-pass transcription remains a separate opt-in.

Wrengle keeps a bounded, process-memory-only set of current-session performance marks for capture ready, first transcript, speech-end-to-final, insertion, and stopping timing. Those marks contain only a stage and elapsed time. They do not contain audio, transcript or utterance text, note paths, note titles, provider hosts, or keys, are not sent to remote telemetry, and disappear when the desktop process exits.

Voice sessions do not use meeting recovery storage. The strict temporary .app/live/ records described in the meeting docs are created only for meeting recording, finalization, saving, report generation, and recovery. They contain completed transcript records and exact save/report bindings, not raw audio or the memory-only prepared report.

A request you hand off to a hosted assistant follows that assistant's own privacy boundaries, the same as typing the request into the assistant panel yourself.

Setup and troubleshooting

  • Open Settings → Dictation and choose the mode, speech engine, language, cleanup level, response speed, personal vocabulary, spoken-punctuation behavior, and trigger words you want. Speech and analysis start on Auto unless you make an explicit selection.
  • For local speech, install the selected Whisper model when the start gate offers the download. For cloud speech, save and verify the matching OpenAI, Deepgram, or ElevenLabs transcription key. Auto uses it immediately unless Cloud voice features was explicitly turned off in Privacy.
  • A specific local Whisper selection uses that exact downloaded, language-compatible rung rather than silently substituting another. Auto chooses the fastest installed compatible rung. After local dictation, Wrengle keeps that speech model warm for roughly 60 seconds of inactivity to make a near-term restart faster.
  • Voice Control or Light edit also needs the selected or Auto-resolved analysis path. Install a Wrengle Local model for local analysis, or configure a supported cloud AI model and connected key. A machine below the local model's RAM threshold can still use Quick Dictation with Verbatim cleanup or a fully configured cloud analysis path.
  • On macOS, Wrengle asks for microphone access during startup. If it was denied or later revoked, use Settings → Privacy & Data → Desktop access to open the operating-system settings, then explicitly retry or restart as directed. The current unpackaged Windows build instead reports native permission as Not required without probing the device or global desktop-app switch. If Windows or the selected device blocks capture, the later device-open attempt fails prompt-free; the Privacy page always offers a Windows microphone Settings link. Test microphone and the dictation shortcuts never open a permission prompt. Choose Dictation microphone, then use Test microphone to confirm its input level. If a saved device is unavailable, choose Refresh microphones and select a connected input or System default microphone.
  • Focus an editable note or assistant composer and confirm its name in the destination indicator. A target change while text is pending is blocked instead of inserting into a different surface.
  • Stop or finish any meeting recording or finalization before starting dictation. Meeting and voice capture cannot own the microphone path at the same time.
  • If cloud dictation shows Reconnecting, keep the session open while it retries. A long outage can overflow the bounded audio buffer; any affected speech is reported rather than silently claimed as transcribed.

Limitations

Recognition quality depends on your microphone, audio environment, selected language, vocabulary hints, and transcription model. Selecting a recognition language does not localize Wrengle's interface or its built-in English control phrases. In Voice Control, exact trigger matching separates commands from bare dictation; unrecognized triggered commands surface in the panel instead of being silently treated as successful. Quick Dictation never routes trigger words as commands. Delete actions, shell commands, and non-local agent requests are confirmation-gated, but actions whose external state changes later are not generally reversible. Voice capture is macOS-first and follows the same platform support as meeting capture.

docs / voice-controlAll documentation