Under the hood
Whisper, Parakeet, Ultra, Claude, GPT-5, Gemini, and more. Pick the model that fits the job and your hardware.
How to read this page
Superwhisper runs two models back to back. A speech recognition model turns your voice into text. Then an optional language model rewrites that text in Super Mode, so an email comes out like an email and a code prompt stays technical.
You can mix and match. On-device Parakeet into cloud Claude Sonnet is a common setup. So is S1-Voice into nothing at all for raw transcription. The transcription figures below are measured numbers from our own benchmark. The language model dots are relative ratings, because we do not benchmark those ourselves.
Speech recognition
This is the first model in the chain. Run it in the cloud when you want top accuracy or the lowest latency, or run it on your own machine when the audio has to stay local. The Fast, Nano, and Standard Whisper models come with the free tier. Want to try one on a real recording first? Transcribe an audio file and check the output.
Audio goes to the Superwhisper proxy, gets transcribed, and comes back. Best for top accuracy and the lowest cloud latency.
| Model | Provider | Error rate | Speed | Languages | Tier |
|---|---|---|---|---|---|
| S1-Voice | 6.6% | 26× | 100+ | Pro | |
| Scribe V2 | 7.5% | 14× | 99 | Pro | |
| Nova 3 | 10.1% | 23× | 36 | Pro | |
| Ultra | 10.6% | 24× | 100+ | Pro | |
| Nova 2 | — | — | 36 | Pro | |
| Nova Medical | — | — | English | Pro |
These run locally on your Mac, PC, or iPhone. Audio never leaves the device and no internet connection is needed. Fast, Nano, and Standard are free. The faster Parakeets and larger Whisper variants ship with Pro.
| Model | Provider | Error rate | Speed | Languages | Size | Tier |
|---|---|---|---|---|---|---|
| Cohere Transcribe | 8.4% | 24× | 100+ | 1.3 GB | Pro | |
| Parakeet V2 | 10.4% | 133× | English | 476 MB | Pro | |
| Parakeet V3 | 10.7% | 118× | 24 | 494 MB | Pro | |
| Ultra V3 Turbo | 10.9% | 9.4× | 100+ | 1.6 GB | Pro | |
| Ultra V3 | 13.0% | 6.1× | 100+ | 3.0 GB | Pro | |
| Standard | 15.6% | 30× | 100+ | 500 MB | Free | |
| Pro | 15.8% | 11× | 100+ | 1.5 GB | Pro | |
| Nano | — | — | 100+ | 150 MB | Free | |
| Fast | 20.3% | 126× | 100+ | 75 MB | Free |
Super Mode
The second model is optional. Super Mode hands your transcript to a language model that can rewrite it, translate it, or reshape it for the app you're typing into. Cloud requests run through the Superwhisper proxy, so providers never see your account and nothing is kept for training.
The widest range of intelligence and the longest context windows. Pick a fast one for quick replies, a smarter one for careful rewrites.
| Model | Provider | Speed / Intelligence | Context | Tier |
|---|---|---|---|---|
| Claude Sonnet 4.6 | 1M | Pro | ||
| Claude Sonnet 4.5 | 200k | Pro | ||
| Claude Haiku 4.5 | 200k | Pro | ||
| GPT-5.4 mini | 400k | Pro | ||
| GPT-5.4 nano | 400k | Pro | ||
| GPT-5.3 Instant | 128k | Pro | ||
| GPT-5.2 | 400k | Pro | ||
| GPT-5.1 | 400k | Pro | ||
| GPT-5 mini | 400k | Pro | ||
| GPT-5 nano | 400k | Pro | ||
| Gemini 3 Flash | 1M | Pro | ||
| Gemini 3.1 Flash Lite | 1M | Pro | ||
| Grok 4.1 Fast | 2M | Pro | ||
| S1-Language | 128k | Pro | ||
| Llama 3.1 8B | 128k | Pro |
These run locally through llama.cpp on Apple Silicon or Windows. No internet needed. Size is the download on disk. Included with Pro.
| Model | Provider | Speed / Intelligence | Size | Tier |
|---|---|---|---|---|
| GPT OSS 20B | 14 GB | Pro | ||
| DeepSeek R1 Distill | 5.4 GB | Pro | ||
| Ministral 3 8B | 5.2 GB | Pro | ||
| Llama 3 8B | 4.9 GB | Pro | ||
| Mistral 7B v0.2 | 4.4 GB | Pro | ||
| Llama 3.2 3B | 1.9 GB | Pro | ||
| Phi-2 3B | 1.8 GB | Pro |
Methodology
The transcription numbers are ours. Every error rate and speed figure in the tables above is an average over the same eight datasets you can read in full on the benchmarks page, covering read speech, spontaneous speech, meetings and specialist vocabulary. Click a model name for its dataset-by-dataset breakdown and the runs behind it.
Error rate is the share of words that differ from a human transcript, so lower is better. Speed is a multiple of real time, so 20x means a minute of speech comes back in three seconds. On-device figures are measured on an Apple M4 and a slower machine will be slower. A dash means we have not run that model over the full set yet, not that it scored badly.
Language model scores are anchored to Artificial Analysis, because we do not run a benchmark for rewriting. Those dots are relative to the other language models in Superwhisper: a 5 is the best in its class, not a claim that every 5-dot model is equivalent.
Privacy
Every on-device model in this list runs locally. Your microphone input never leaves the machine, we don't log audio, and nothing you dictate is stored on our servers. Work on a plane or inside a locked-down environment and the models behave the same way.
Cloud models are there when you want them. Pick one and your audio and text pass through the Superwhisper proxy to the provider and back. Providers see a proxy request, not your account, and nothing is retained for training. Enterprise customers can swap in their own API keys or host compatible models behind a VPC. Superwhisper is SOC 2 Type II certified and HIPAA compliant.
Drop in an audio file
.mp3, .wav, .aac...
Transcribe an audio file
Drop in an MP3 or voice memo and see how the models read it.
Test your microphone
Check your mic and levels before you record anything important.
Support
The free tier runs the core Whisper models on-device. Pro adds Parakeet, Cohere, S1-Voice, and every cloud model.
Download free