NVIDIA · Runs on your machine
Measured on 8 of 8 datasets in the Superwhisper speech benchmark.
Available as a model choice inside Superwhisper.
Results
NVIDIA's Parakeet V3 runs on the Neural Engine through Argmax ParakeetKit. On Apple silicon it is the fastest thing we ship by a wide margin, and it stays close to the larger Whisper models on clean speech.
Apple M4
73
Blended score
10.7%
Word error rate
0.07 s
Wait for the text
Macro average over 8 datasets on Apple M4
Apple M4 Max
74
Blended score
10.6%
Word error rate
0.06 s
Wait for the text
Macro average over 3 datasets on Apple M4 Max
On Apple M4
| Dataset | Score ↑ | WER ↓ | Speed ↑ | Recall ↑ | F-score ↑ | Response ↓ |
|---|---|---|---|---|---|---|
| AMI Meetingssw-app-coverage-esb-ami-v2-20260802 | 75 | 10.0% | 167× | — | — | 0.07 s |
| Common Voice, Australian Englishsw-app-coverage-cv24-australia-v2-20260802 | 86 | 5.8% | 83× | — | — | 0.07 s |
| Common Voice, spontaneous Englishsw-app-coverage-cv-spontaneous4-v2-20260802 | 67 | 13.4% | 118× | — | — | 0.07 s |
| Earnings 22sw-app-coverage-esb-earnings22-v2-20260802 | 69 | 12.6% | 145× | — | — | 0.07 s |
| Earnings 22 (with vocabulary)sw-app-main-contextual-earnings22-20260802 | 62 | 15.0% | 114× | 54.8% | 70.3% | 0.07 s |
| LibriSpeech Othersw-app-main-esb-librispeech-other-20260802 | 85 | 5.9% | 144× | — | — | 0.07 s |
| Loquacioussw-app-coverage-loquacious-v2-20260802 | 81 | 7.7% | 124× | — | — | 0.07 s |
| Phonetic vocabularysw-app-main-phonetic-vocab-20260802 | 62 | 15.1% | 46× | 71.4% | 76.9% | 0.07 s |
Also measured on Apple M4 Max
| Dataset | Score ↑ | WER ↓ | Speed ↑ | Recall ↑ | F-score ↑ | Response ↓ |
|---|---|---|---|---|---|---|
| Earnings 22 (with vocabulary)20260731-175829 | 61 | 15.4% | 130× | 91.5% | 92.5% | 0.06 s |
| LibriSpeech Other20260721-211510 | 86 | 5.6% | 164× | — | — | 0.06 s |
| Phonetic vocabulary20260731-175829 | 74 | 10.6% | 42× | 65.7% | 73.0% | 0.06 s |
Updated August 27, 2026 from bench commit 5344cab. The second line of each row is the run that produced it.
Method
Local results come from driving the real Superwhisper app over each clip, so they reflect the engine and decode settings that ship in the product rather than a separate implementation.
Word error rate counts insertions, deletions and substitutions against a human transcript, so lower is better. Speed is a multiple of real time, so higher is better. The full method and every other model is on the benchmarks page.
Compare
Same audio, same machines, same method.
Support
Free to download, and the on-device models cost nothing to run.
Download free