Wispr · Runs in the cloud
Measured on 8 of 8 datasets in the Superwhisper speech benchmark.
In the benchmark for comparison, not shipped in the app.
Results
The same Wispr Flow runs with Auto Cleanup turned off and transforms opted out, read straight from the app's raw transcript. It is the fairer comparison against models that have no formatting layer.
76
Blended score
9.7%
Word error rate
0.54 s
Wait for the text
Macro average over 8 datasets
| Dataset | Score ↑ | WER ↓ | Recall ↑ | F-score ↑ | Response ↓ |
|---|---|---|---|---|---|
| AMI Meetingswispr-raw-esb-ami-20260803 | 71 | 11.8% | — | — | 0.54 s |
| Common Voice, Australian Englishwispr-raw-cv24-au-500-20260803 | 89 | 4.5% | — | — | 0.54 s |
| Common Voice, spontaneous Englishwispr-raw-cv-spontaneous4-20260803 | 67 | 13.1% | — | — | 0.54 s |
| Earnings 22wispr-raw-esb-earnings22-20260803 | 78 | 8.7% | — | — | 0.54 s |
| Earnings 22 (with vocabulary)wispr-raw-contextual-earnings22-20260803 | 53 | 18.8% | 62.8% | 76.6% | 0.54 s |
| LibriSpeech Otherwispr-raw-esb-librispeech-other-20260803 | 91 | 3.8% | — | — | 0.54 s |
| Loquaciouswispr-raw-loquacious-20260803 | 81 | 7.4% | — | — | 0.54 s |
| Phonetic vocabularywispr-raw-phonetic-vocab-20260803 | 76 | 9.5% | 74.3% | 77.6% | 0.54 s |
Cloud results do not depend on the machine that made the request, so they are not split by device.
Updated August 27, 2026 from bench commit 5344cab. The second line of each row is the run that produced it.
Method
Cloud results come from calling the model over its public API with the same audio as everything else in the benchmark.
Word error rate counts insertions, deletions and substitutions against a human transcript, so lower is better. Speed is a multiple of real time, so higher is better. The full method and every other model is on the benchmarks page.
Compare
Same audio, same machines, same method.
Support