Benchmarks

Speech to text benchmarks

We run 16 speech models over 8 datasets and publish what comes out. Our own models and the ones we compete with, on the same audio and the same machines.

Word error rate, speed, and how well each one handles names it has never heard. Including the rows where somebody else wins.

Leaderboard

Every model, measured the same way

Accuracy and response time in one number, weighted three to one toward accuracy. The best 7 of 7, best on the left. Columns start at zero, so a close field draws columns of a similar height.Higher is better

7 results

Model
Superwhisper S1 VoiceSuperwhisper
836.8%32×74.8%80.2%
ElevenLabs Scribe v2ElevenLabs
817.5%14×91.0%91.4%
OpenAI GPT TranscribeOpenAI
807.8%13×90.9%93.6%
Wispr Flow (raw)Wispr
769.7%68.5%77.1%
Deepgram Nova 3Deepgram
7510.1%23×83.3%86.3%
Gemini 3.5 TranscribeGoogle
748.2%7.1×87.6%90.1%
Superwhisper UltraSuperwhisper
7310.6%24×85.1%87.8%

Updated August 27, 2026 from bench commit 5344cab. Pick a model to see it dataset by dataset.

Word error rate is the share of words a model gets wrong against a human transcript. Response time is how long you wait for the text after you stop talking, measured on dictation-length clips. The blended score puts both into one number out of 100, weighted three to one in favour of accuracy.

Cloud and offline are listed separately. A hosted model and one running on your laptop are answering different questions, and a combined ranking would mostly measure which of the two you happened to pick.

Select the datasets you care about and the table re-averages. Each dataset counts once regardless of how many clips it holds, so a large easy corpus cannot paper over a small hard one.

Scoring

What the blended score is

Accuracy is scored against a fixed 30% word error rate. At roughly one word in three wrong, a transcript is faster to retype than to fix, so that is where the accuracy half of the score reaches zero.

The other half is response time: how long you sit waiting for words after you stop talking. Anything under a second scores full marks, because you cannot beat instant, and the score reaches zero at eight seconds.

This is measured, not estimated from throughput. We take the median time to a finished transcript over clips of 8 to 15 seconds, the length of an ordinary dictation, from the same runs the accuracy comes from. A dictation app reports its own end to end latency instead of a wall clock and goes in the same column, because both are the gap between the end of speech and the text arriving.

Accuracy carries three quarters of the weight, because a transcript you have to correct is not a fast transcript. Both anchors are fixed numbers rather than a ranking against the field, so a model's score does not move when another model joins or leaves the benchmark.

Method

How these numbers are produced

Local models are measured by driving the real Superwhisper app, not a reimplementation of it, so the result comes from the same engines and decode settings you get in the product. Cloud models are called over their public APIs.

Competitor apps that have no API get driven as apps. Wispr Flow is measured by running the real product over the same clips and reading back what it types, with its cleanup pass switched off so the transcript can be scored against a word-for-word reference.

Every run writes its clip-level output to the benchmark repository along with the machine it ran on and the commit that produced it. Nothing on this page is a number we typed in by hand.

Datasets

What the models are listening to

Eight sets, chosen to cover the range between an audiobook read in a quiet room and four people talking over each other in a meeting.

AMI Meetings
Recorded multi-party meetings with overlapping speech and far-field microphones, via the End-to-end Speech Benchmark.
Common Voice, Australian English
Volunteer-recorded sentences from Australian speakers in Mozilla Common Voice 24.
Common Voice, spontaneous English
Unscripted English speech from Common Voice, with the pauses and restarts of ordinary talking.
Earnings 22
Earnings calls in a mix of global English accents, heavy on company names and financial jargon.
Earnings 22 (with vocabulary)
The same earnings calls, run with the expected company and product names supplied as custom vocabulary.
LibriSpeech Other
Read audiobook speech from the harder half of the LibriSpeech test split, via the End-to-end Speech Benchmark.
Loquacious
Long-form conversational English drawn from a range of recording conditions.
Phonetic vocabulary
Sentences built around names and terms that sound like common words, to test how well custom vocabulary holds up.

Models

Results by model

Every model has its own page with the full dataset breakdown and the runs behind it. For which one to pick day to day, rather than which one scores best, see models in Superwhisper.

Support

Frequently asked questions

Try the models yourself

Superwhisper is free to download, and the on-device models cost nothing to run.

Download free