Superwhisper · Legacy
Measured on 8 of 8 datasets in the Superwhisper speech benchmark.
Superseded by Superwhisper S1 Voice, which is faster and more accurate.
Results
Ultra was the cloud model behind Superwhisper Pro, built for the audio the on-device models find hardest: accents far from American English, people talking over each other, and names no general model has heard before.
This one is retired. Superwhisper S1 Voice replaced it and beats it on every measure below, so the numbers here are a record of where we were rather than a recommendation.
Audio went to our servers for Ultra, as it does for S1 Voice. If that is not acceptable for your work, the on-device models run without a network connection at all.
73
Blended score
10.6%
Word error rate
0.49 s
Wait for the text
Macro average over 8 datasets
| Dataset | Score ↑ | WER ↓ | Speed ↑ | Recall ↑ | F-score ↑ | Response ↓ |
|---|---|---|---|---|---|---|
| AMI Meetingssw-app-ultra-esb-ami-20260806 | 61 | 15.6% | 27× | — | — | 0.49 s |
| Common Voice, Australian Englishsw-app-ultra-common-voice-24-en-au-20260806 | 89 | 4.4% | 14× | — | — | 0.49 s |
| Common Voice, spontaneous Englishsw-app-ultra-common-voice-spontaneous-4-en-20260806 | 63 | 14.8% | 23× | — | — | 0.49 s |
| Earnings 22sw-app-ultra-esb-earnings22-20260806 | 79 | 8.3% | 33× | — | — | 0.49 s |
| Earnings 22 (with vocabulary)sw-app-ultra-contextual-earnings22-20260806 | 48 | 20.7% | 29× | 84.6% | 89.8% | 0.49 s |
| LibriSpeech Othersw-app-ultra-esb-librispeech-other-20260806 | 92 | 3.4% | 31× | — | — | 0.49 s |
| Loquacioussw-app-ultra-loquacious-test-20260806 | 77 | 9.1% | 20× | — | — | 0.49 s |
| Phonetic vocabularysw-app-ultra-phonetic-vocab-20260806 | 78 | 8.9% | 10.0× | 85.7% | 85.7% | 0.49 s |
Cloud results do not depend on the machine that made the request, so they are not split by device.
Updated August 27, 2026 from bench commit 5344cab. The second line of each row is the run that produced it.
Method
Cloud results come from calling the model over its public API with the same audio as everything else in the benchmark.
Word error rate counts insertions, deletions and substitutions against a human transcript, so lower is better. Speed is a multiple of real time, so higher is better. The full method and every other model is on the benchmarks page.
Compare
Same audio, same machines, same method.
Support
Free to download, and the on-device models cost nothing to run.
Download free