Cohere · Runs on your machine

Cohere Transcribe

Measured on 8 of 8 datasets in the Superwhisper speech benchmark.

Available as a model choice inside Superwhisper.

Results

Cohere Transcribe accuracy and speed

Cohere Transcribe runs on your machine through MLX at 4-bit precision. It is the most accurate offline option we ship, and the one to pick when audio cannot leave the laptop.

Quantising to 4-bit costs a little accuracy against the full-precision weights and buys a lot of memory back. These numbers are for the quantised build we actually ship.

Apple M4

79

Blended score

8.4%

Word error rate

0.45 s

Wait for the text

Macro average over 8 datasets on Apple M4

Apple M4 Max

78

Blended score

8.7%

Word error rate

0.18 s

Wait for the text

Macro average over 3 datasets on Apple M4 Max

On Apple M4

DatasetScore WER Speed Recall F-score Response
AMI Meetingssw-app-coverage-esb-ami-v2-20260802817.4%25×0.45 s
Common Voice, Australian Englishsw-app-coverage-cv24-australia-v2-20260802932.8%21×0.45 s
Common Voice, spontaneous Englishsw-app-coverage-cv-spontaneous4-v2-202608025418.5%23×0.45 s
Earnings 22sw-app-coverage-esb-earnings22-v2-20260802808.0%26×0.45 s
Earnings 22 (with vocabulary)sw-app-main-contextual-earnings22-202608026613.8%29×69.1%81.3%0.45 s
LibriSpeech Othersw-app-main-esb-librispeech-other-20260802942.2%26×0.45 s
Loquacioussw-app-coverage-loquacious-v2-20260802875.2%23×0.45 s
Phonetic vocabularysw-app-main-phonetic-vocab-20260802788.9%16×68.6%72.7%0.45 s

Also measured on Apple M4 Max

DatasetScore WER Speed Recall F-score Response
Earnings 22 (with vocabulary)20260731-1758296315.0%79×82.4%89.1%0.18 s
LibriSpeech Other20260721-211510942.2%80×0.18 s
Phonetic vocabulary20260731-175829789.0%33×77.1%79.4%0.18 s

Updated August 27, 2026 from bench commit 5344cab. The second line of each row is the run that produced it.

Method

How this was measured

Local results come from driving the real Superwhisper app over each clip, so they reflect the engine and decode settings that ship in the product rather than a separate implementation.

Word error rate counts insertions, deletions and substitutions against a human transcript, so lower is better. Speed is a multiple of real time, so higher is better. The full method and every other model is on the benchmarks page.

Compare

Other models in the benchmark

Same audio, same machines, same method.

Support

Frequently asked questions

Run Cohere Transcribe yourself

Free to download, and the on-device models cost nothing to run.

Download free