Under the hood

AI models in Superwhisper

Whisper, Parakeet, Ultra, Claude, GPT-5, Gemini, and more. Pick the model that fits the job and your hardware.

How to read this page

Two jobs, two model types

Superwhisper runs two models back to back. A speech recognition model turns your voice into text. Then an optional language model rewrites that text in Super Mode, so an email comes out like an email and a code prompt stays technical.

You can mix and match. On-device Parakeet into cloud Claude Sonnet is a common setup. So is S1-Voice into nothing at all for raw transcription. The transcription figures below are measured numbers from our own benchmark. The language model dots are relative ratings, because we do not benchmark those ourselves.

Speech recognition

Models that turn your voice into text

This is the first model in the chain. Run it in the cloud when you want top accuracy or the lowest latency, or run it on your own machine when the audio has to stay local. The Fast, Nano, and Standard Whisper models come with the free tier. Want to try one on a real recording first? Transcribe an audio file and check the output.

Cloud transcription

Audio goes to the Superwhisper proxy, gets transcribed, and comes back. Best for top accuracy and the lowest cloud latency.

ModelProviderError rateSpeedLanguagesTier
S1-Voice
Superwhisper
6.6%26×100+Pro
Scribe V2
ElevenLabs
7.5%14×99Pro
Nova 3
Deepgram
10.1%23×36Pro
Ultra
Superwhisper
10.6%24×100+Pro
Nova 2
Deepgram
36Pro
Nova Medical
Deepgram
EnglishPro

On-device transcription

These run locally on your Mac, PC, or iPhone. Audio never leaves the device and no internet connection is needed. Fast, Nano, and Standard are free. The faster Parakeets and larger Whisper variants ship with Pro.

ModelProviderError rateSpeedLanguagesSizeTier
Cohere Transcribe
Cohere
8.4%24×100+1.3 GBPro
Parakeet V2
NVIDIA
10.4%133×English476 MBPro
Parakeet V3
NVIDIA
10.7%118×24494 MBPro
Ultra V3 Turbo
OpenAI Whisper
10.9%9.4×100+1.6 GBPro
Ultra V3
OpenAI Whisper
13.0%6.1×100+3.0 GBPro
Standard
OpenAI Whisper
15.6%30×100+500 MBFree
Pro
OpenAI Whisper
15.8%11×100+1.5 GBPro
Nano
OpenAI Whisper
100+150 MBFree
Fast
OpenAI Whisper
20.3%126×100+75 MBFree

Super Mode

Language models that clean up and reformat

The second model is optional. Super Mode hands your transcript to a language model that can rewrite it, translate it, or reshape it for the app you're typing into. Cloud requests run through the Superwhisper proxy, so providers never see your account and nothing is kept for training.

Cloud language models

The widest range of intelligence and the longest context windows. Pick a fast one for quick replies, a smarter one for careful rewrites.

ModelProviderSpeed / IntelligenceContextTier
Claude Sonnet 4.6
Anthropic
1MPro
Claude Sonnet 4.5
Anthropic
200kPro
Claude Haiku 4.5
Anthropic
200kPro
GPT-5.4 mini
OpenAI
400kPro
GPT-5.4 nano
OpenAI
400kPro
GPT-5.3 Instant
OpenAI
128kPro
GPT-5.2
OpenAI
400kPro
GPT-5.1
OpenAI
400kPro
GPT-5 mini
OpenAI
400kPro
GPT-5 nano
OpenAI
400kPro
Gemini 3 Flash
Google
1MPro
Gemini 3.1 Flash Lite
Google
1MPro
Grok 4.1 Fast
xAI
2MPro
S1-Language
Superwhisper
128kPro
Llama 3.1 8B
Meta / Groq
128kPro

On-device language models

These run locally through llama.cpp on Apple Silicon or Windows. No internet needed. Size is the download on disk. Included with Pro.

ModelProviderSpeed / IntelligenceSizeTier
GPT OSS 20B
OpenAI
14 GBPro
DeepSeek R1 Distill
DeepSeek
5.4 GBPro
Ministral 3 8B
Mistral
5.2 GBPro
Llama 3 8B
Meta
4.9 GBPro
Mistral 7B v0.2
Mistral
4.4 GBPro
Llama 3.2 3B
Meta
1.9 GBPro
Phi-2 3B
Microsoft
1.8 GBPro

Methodology

Where these numbers come from

The transcription numbers are ours. Every error rate and speed figure in the tables above is an average over the same eight datasets you can read in full on the benchmarks page, covering read speech, spontaneous speech, meetings and specialist vocabulary. Click a model name for its dataset-by-dataset breakdown and the runs behind it.

Error rate is the share of words that differ from a human transcript, so lower is better. Speed is a multiple of real time, so 20x means a minute of speech comes back in three seconds. On-device figures are measured on an Apple M4 and a slower machine will be slower. A dash means we have not run that model over the full set yet, not that it scored badly.

Language model scores are anchored to Artificial Analysis, because we do not run a benchmark for rewriting. Those dots are relative to the other language models in Superwhisper: a 5 is the best in its class, not a claim that every 5-dot model is equivalent.

Privacy

On-device by default, cloud by choice

Every on-device model in this list runs locally. Your microphone input never leaves the machine, we don't log audio, and nothing you dictate is stored on our servers. Work on a plane or inside a locked-down environment and the models behave the same way.

Cloud models are there when you want them. Pick one and your audio and text pass through the Superwhisper proxy to the provider and back. Providers see a proxy request, not your account, and nothing is retained for training. Enterprise customers can swap in their own API keys or host compatible models behind a VPC. Superwhisper is SOC 2 Type II certified and HIPAA compliant.

Keep exploring

Support

Frequently asked questions

One app, every model

The free tier runs the core Whisper models on-device. Pro adds Parakeet, Cohere, S1-Voice, and every cloud model.

Download free