Introducing the S1 family of models

Nico DiPlacido
Nico DiPlacidoAug 19, 2026 · 7 min read

The S1 family of models has arrived.

At Superwhisper we believe that state-of-the-art, voice-to-text tools can be built to be faster, and more accurate without sacrificing your privacy by collecting or training on your data.

Today we’re excited to launch our proprietary Superwhisper models: S1-Voice, S1-Language, and S1-mini.

The S1 models have been trained and fine-tuned in-house at Superwhisper using bespoke synthetic, and publicly-available datasets that capture the hurdles of our customers: noisy office environments, whispered speech, and highly technical medical, financial, and mixed language dictations.

Each model is tuned from our customers feedback: S1-Voice turns speech into polished text in any application, S1-Language providing promptable inference to manipulate or clean up your dictations, and S1-mini offering tone-controlled, more-than-smart-enough clean up and reformatting achieved completely locally on your laptop or phone.

S1-mini: on-device cleanup, open weights

S1-mini is a small 462 MB 0.6B language model that uses local inference to output highly accurate, and tightly contextually aware formatted text with zero network requests.

We’re excited to release S1-mini with open weights now available in the Hugging Face model repository.

By design, S1-mini is highly focused to achieve a handful of common tasks extremely well. This allows us to have an efficient and lightweight model running locally on consumer devices.

Controlling your Tone preferences

After selecting S1-mini as your language model, you’ll have access to the new tone control slider with 5 preferences:

  • casual (all lowercase, minimal punctuation)
  • semi-casual
  • balanced
  • semi-formal
  • formal (contractions expanded, full punctuation)

Automatic formatting, text replacements, and mistake correction

S1-mini automatically detects lists content and transforms them into ordered or unordered lists for better visual clarity and readability. In mail mode, the model automatically contextually formats the transcript as an email with your greeting line, sign-off, and proper spacing:

Raw transcript:

hi Ben thanks for your patience while we looked into this we found three issues with the last shipment the packaging was damaged the wrong size was sent and the invoice amount was incorrect we already proc- we already processed a replacement and a refund for the difference let us know if you have any other questions thanks the support team

Superwhisper S1-mini returns:

Hi Ben, Thanks for your patience while we looked into this. We found three issues with the last shipment: • The packaging was damaged • The wrong size was sent • The invoice amount was incorrect We already processed a replacement and a refund for the difference. Let us know if you have any other questions. Thanks, The Support Team

S1-mini also handles filler words, stutters, and corrections made while dictating by smartly backtracking and capturing what you intended to communicate. For instance, correcting yourself mid-dictation:

Raw transcript:

the meeting is on Tuesday I mean Thursday

Superwhisper S1-mini returns:

the meeting is on Thursday

It also knows what to leave alone: "three or four days" is a timeframe of a trip, and "tea or coffee" is a question rather than a mistake.

The model correctly and consistently renders numbers, dates, currency, percentages, phone numbers with country codes, spoken email addresses and URLs to make transcripts more readable. Say “support at superwhisper dot com" and you get support@superwhisper.com in return.

What it will not do

S1-mini is ruthlessly obedient. It will never add content you did not say, correct facts, soften profanity, flag what you are talking about, or rewrite your dialect. Its sole purpose is to clean the raw ASR transcripts outputted by speech-to-text models.

Thoroughly tested

We evaluated S1-mini on a held-out set of 7,519 cases across 104 transcripts, none of which the model saw during training. Token accuracy reaches 94.8%, with a text-edit error rate of just 11.6%. On email-formatted text, it identifies the greeting line correctly 99.3% of the time and the sign-off 97.9% of the time. It also matches the correct output structure (list versus paragraph) 97.6% of the time, and produces exact email addresses in 92% of cases. Output stability is high as well. Fewer than 1% of generations show any degenerate behavior, such as looping or truncation, and the model correctly withholds output 98.6% of the time when no content should be transcribed.

S1-Voice

S1-Voice is our cloud-hosted speech-to-text model, trained in-house and hosted by us. It transcribes your words up to 46x faster than it took you to speak them, with most dictations under 30s appearing 0.32 seconds after you stop talking.

We evaluated S1-Voice across eight datasets that include meeting audio, earnings calls, and spontaneous speech, not just cleanly read sentences. It averages a 6.8% word error rate across the benchmark and drops to 2.2% on LibriSpeech. Its 6.8% average was the lowest of the 15 models we tested.

S1-Voice also earned the highest blended score in the benchmark at 83 out of 100. For comparison, WisprFlow scored a 76 (finishing in 4th place)

S1-Language

S1-Language is a Superwhisper cloud-hosted instruction-following model with high throughput and precision. It takes dictations, thoughts, meetings recordings, and performs the heavy lifting to clean, format, or summarize into a task specific output at blazing speeds.

S1 language is powerful enough to handle most advanced custom modes, instructions, and rule sets. If you want meeting notes structured a certain way or emails that always sound like you, S1-Language should be your new default choice.

You’ll find S1-Language today in the Superwhisper model picker alongside the latest SOTA models from Anthropic, OpenAI, and Groq.

The new defaults for dictation

As a daily driver, we recommend one of these options depending on your cloud vs offline preference.

  • Offline: Message mode powered by Cohere Transcribe (local), with S1-mini as the language model.
  • Cloud-hosted: Message mode powered by S1-Voice, with S1-Language as the language model.

From the Superwhisper Team:

We have been so impressed with both the S1 model performance, as much as the positive reception from our beta testers over the past few months.

If you’re looking to get involved with early releases of Superwhisper, or provide deeper feedback as we build the default choice for professional speech-to-text, please join our Discord community, or come work with us.

Try Superwhisper free

Free tier that doesn't expire. 30-day refund if Pro isn't for you.

Download free

Available for Mac, Windows and iOS