Under the hood

AI models in Superwhisper

Whisper, Parakeet, Ultra, Claude, GPT-5, Gemini, and more. Pick the model that fits the job and your hardware.

How to read this page

Two jobs, two model types

Superwhisper runs two models back to back. A speech recognition model turns your voice into text. Then an optional language model rewrites that text in Super Mode, so an email comes out like an email and a code prompt stays technical.

You can mix and match. On-device Parakeet into cloud Claude Sonnet is a common setup. So is Ultra into nothing at all for raw transcription. The speed and accuracy dots rank models relative to each other inside Superwhisper, not against the whole industry.

Speech recognition

Models that turn your voice into text

This is the first model in the chain. Run it in the cloud when you want top accuracy or the lowest latency, or run it on your own machine when the audio has to stay local. The Fast, Nano, and Standard Whisper models come with the free tier. Want to try one on a real recording first? Transcribe an audio file and check the output.

Cloud transcription

Audio goes to the Superwhisper proxy, gets transcribed, and comes back. Best for top accuracy and the lowest cloud latency.

ModelProviderSpeed / AccuracyLanguagesTier
UltraSuperwhisper
100+Pro
S1-VoiceSuperwhisper
100+Pro
Scribe V2ElevenLabs
99Pro
Nova 3Deepgram
36Pro
Nova 2Deepgram
36Pro
Nova MedicalDeepgram
EnglishPro

On-device transcription

These run locally on your Mac, PC, or iPhone. Audio never leaves the device and no internet connection is needed. Fast, Nano, and Standard are free. The faster Parakeets and larger Whisper variants ship with Pro.

ModelProviderSpeed / AccuracyLanguagesSizeTier
Parakeet V2NVIDIA
English476 MBPro
Parakeet V3NVIDIA
24494 MBPro
Ultra V3OpenAI Whisper
100+3.0 GBPro
Ultra V3 TurboOpenAI Whisper
100+1.6 GBPro
ProOpenAI Whisper
100+1.5 GBPro
StandardOpenAI Whisper
100+500 MBFree
NanoOpenAI Whisper
100+150 MBFree
FastOpenAI Whisper
100+75 MBFree

Super Mode

Language models that clean up and reformat

The second model is optional. Super Mode hands your transcript to a language model that can rewrite it, translate it, or reshape it for the app you're typing into. Cloud requests run through the Superwhisper proxy, so providers never see your account and nothing is kept for training.

Cloud language models

The widest range of intelligence and the longest context windows. Pick a fast one for quick replies, a smarter one for careful rewrites.

ModelProviderSpeed / IntelligenceContextTier
Claude Sonnet 4.6Anthropic
1MPro
Claude Sonnet 4.5Anthropic
200kPro
Claude Haiku 4.5Anthropic
200kPro
GPT-5.4 miniOpenAI
400kPro
GPT-5.4 nanoOpenAI
400kPro
GPT-5.3 InstantOpenAI
128kPro
GPT-5.2OpenAI
400kPro
GPT-5.1OpenAI
400kPro
GPT-5 miniOpenAI
400kPro
GPT-5 nanoOpenAI
400kPro
Gemini 3 FlashGoogle
1MPro
Gemini 3.1 Flash LiteGoogle
1MPro
Grok 4.1 FastxAI
2MPro
S1-LanguageSuperwhisper
128kPro
Llama 3.1 8BMeta / Groq
128kPro

On-device language models

These run locally through llama.cpp on Apple Silicon or Windows. No internet needed. Size is the download on disk. Included with Pro.

ModelProviderSpeed / IntelligenceSizeTier
GPT OSS 20BOpenAI
14 GBPro
DeepSeek R1 DistillDeepSeek
5.4 GBPro
Ministral 3 8BMistral
5.2 GBPro
Llama 3 8BMeta
4.9 GBPro
Mistral 7B v0.2Mistral
4.4 GBPro
Llama 3.2 3BMeta
1.9 GBPro
Phi-2 3BMicrosoft
1.8 GBPro

Methodology

Where the ratings come from

Transcription scores are anchored to the Hugging Face Open ASR Leaderboard. Parakeet V2 leads on real-time factor with an industry-best 6.05% word error rate for English. Whisper Large V3 scores a little higher on accuracy but runs much slower, which is why Ultra V3 earns a 2 on speed and a 5 on accuracy while the Turbo variant evens out at 4 and 4.

Language model scores are anchored to Artificial Analysis. Claude Sonnet 4.6 and the GPT-5.4 series sit at the top of the intelligence index next to GPT-5.1 and GPT-5.2. Gemini 3.1 Flash Lite pushes past 200 tokens per second, so it takes the speed crown on the lite side. Groq's Llama 3.1 8B is faster still, but its smaller parameter count holds it at a 2 on intelligence.

One thing worth repeating. The dots are relative to the other models in Superwhisper. A 5 is the best in its class. A 1 is there for people who want a tiny download and can live with a rougher transcript.

Privacy

On-device by default, cloud by choice

Every on-device model in this list runs locally. Your microphone input never leaves the machine, we don't log audio, and nothing you dictate is stored on our servers. Work on a plane or inside a locked-down environment and the models behave the same way.

Cloud models are there when you want them. Pick one and your audio and text pass through the Superwhisper proxy to the provider and back. Providers see a proxy request, not your account, and nothing is retained for training. Enterprise customers can swap in their own API keys or host compatible models behind a VPC. Superwhisper is SOC 2 Type II certified and HIPAA compliant.

Keep exploring

Support

Frequently asked questions

One app, every model

The free tier runs the core Whisper models on-device. Pro adds Parakeet, Ultra, and every cloud model.

Download free