Benchmarks
We run 16 speech models over 8 datasets and publish what comes out. Our own models and the ones we compete with, on the same audio and the same machines.
Word error rate, speed, and how well each one handles names it has never heard. Including the rows where somebody else wins.
Leaderboard
7 results
| Model | |||||
|---|---|---|---|---|---|
| 83 | 6.8% | 32× | 74.8% | 80.2% | |
| 81 | 7.5% | 14× | 91.0% | 91.4% | |
| 80 | 7.8% | 13× | 90.9% | 93.6% | |
| 76 | 9.7% | — | 68.5% | 77.1% | |
| 75 | 10.1% | 23× | 83.3% | 86.3% | |
| 74 | 8.2% | 7.1× | 87.6% | 90.1% | |
| 73 | 10.6% | 24× | 85.1% | 87.8% |
Updated August 27, 2026 from bench commit 5344cab. Pick a model to see it dataset by dataset.
Word error rate is the share of words a model gets wrong against a human transcript. Response time is how long you wait for the text after you stop talking, measured on dictation-length clips. The blended score puts both into one number out of 100, weighted three to one in favour of accuracy.
Cloud and offline are listed separately. A hosted model and one running on your laptop are answering different questions, and a combined ranking would mostly measure which of the two you happened to pick.
Select the datasets you care about and the table re-averages. Each dataset counts once regardless of how many clips it holds, so a large easy corpus cannot paper over a small hard one.
Scoring
Accuracy is scored against a fixed 30% word error rate. At roughly one word in three wrong, a transcript is faster to retype than to fix, so that is where the accuracy half of the score reaches zero.
The other half is response time: how long you sit waiting for words after you stop talking. Anything under a second scores full marks, because you cannot beat instant, and the score reaches zero at eight seconds.
This is measured, not estimated from throughput. We take the median time to a finished transcript over clips of 8 to 15 seconds, the length of an ordinary dictation, from the same runs the accuracy comes from. A dictation app reports its own end to end latency instead of a wall clock and goes in the same column, because both are the gap between the end of speech and the text arriving.
Accuracy carries three quarters of the weight, because a transcript you have to correct is not a fast transcript. Both anchors are fixed numbers rather than a ranking against the field, so a model's score does not move when another model joins or leaves the benchmark.
Method
Local models are measured by driving the real Superwhisper app, not a reimplementation of it, so the result comes from the same engines and decode settings you get in the product. Cloud models are called over their public APIs.
Competitor apps that have no API get driven as apps. Wispr Flow is measured by running the real product over the same clips and reading back what it types, with its cleanup pass switched off so the transcript can be scored against a word-for-word reference.
Every run writes its clip-level output to the benchmark repository along with the machine it ran on and the commit that produced it. Nothing on this page is a number we typed in by hand.
Datasets
Eight sets, chosen to cover the range between an audiobook read in a quiet room and four people talking over each other in a meeting.
Models
Every model has its own page with the full dataset breakdown and the runs behind it. For which one to pick day to day, rather than which one scores best, see models in Superwhisper.
Runs on your machine
Support
Superwhisper is free to download, and the on-device models cost nothing to run.
Download free