On-device transcription
Whisperstream transcribes your speech entirely on your computer using NVIDIA's open-weight Parakeet model running on your CPU. Your audio is never uploaded, streamed, or stored on a server, and you do not need an account. After the one-time model download, transcription works with no internet connection at all.
Push-to-talk hotkey
Dictation is push-to-talk by default: hold your hotkey, speak, and release, so nothing is ever listening in the background. The default is Right Shift, and you can remap it to any key from the Controls tab. Prefer hands-free? Switch to toggle mode, where one press starts recording and another stops it. No admin rights are needed to register the shortcut.
Text delivery across apps
Whisperstream delivers transcribed text to the focused field by clipboard paste or optional simulated keystrokes. It works in many applications, including Outlook, Word, Slack, VS Code, browsers, and terminals, without an add-in. Application, security, privilege, administrator, and remote-session policies can still block or limit either method.
Transcript history
Whisperstream can keep a private history of completed dictations, so you can find a past transcript, replay saved audio, and copy the result again later. History is enabled by default, can be disabled, and is encrypted at rest on your device rather than synced by Whisperstream. Search by keyword, play saved audio beside its text, and delete records you no longer want.
Audio file import
Beyond live dictation, Whisperstream transcribes audio files you select on your device without uploading or modifying the source audio. If cloud enhancement is enabled, the resulting transcript can be sent to the provider you configured. When transcript saving is on, the finished result lands in history, ready to search, replay when audio was saved, and copy.
39 languages
Whisperstream transcribes 39 languages, including English, Spanish, French, German, Chinese, Japanese, Korean, Arabic, and Hindi. NVIDIA Parakeet covers 25 languages and Qwen3 ASR adds 14 more, including Cantonese, Indonesian, Malay, Persian, Thai, Turkish, and Vietnamese. Pick your language and Whisperstream selects the on-device model, downloading a matching model when needed.
Works offline
After its required speech model is downloaded, Whisperstream can record and transcribe without a connection. Updates, model downloads, license activation and periodic validation, user-configured cloud enhancement, and the separate Claude Code workflow used by Agent Speak can still use the network. Core dictation remains available when your network is down, subject to license state.
System-volume ducking
Whisperstream can lower or mute your system volume while you dictate, so music or a video does not bleed into your recording or distract you. Choose off, reduce to 20 percent, or full mute in the Audio settings. Volume returns to its previous level the moment you release the hotkey.