Ask a dictation app forum which app is most accurate and you’ll get heated opinions. What you won’t get is numbers. We ran five apps through five identical audio recordings — 90 seconds of conference speech, four minutes of longform, a numbers/dates clip, coding dictation, casual speech clean and noisy. Every model in every app. Here’s what actually happens.

Tested May 2026 using methodology v1.2. All numbers are Word Error Rate (WER) — lower is better.

The same app can swing 13× in accuracy depending on the model

Superwhisper is the clearest example, but the pattern shows up everywhere.

AppModelWERType
SuperwhisperS1-Voice (default cloud)22.8%Cloud
SuperwhisperParakeet (default local)6.1%Local
SuperwhisperWhisper Standard1.8%Local
SuperwhisperUltra1.6%Cloud
Wispr FlowStandard3.9%Cloud
Willow VoiceCloud1.5%Cloud
Aqua VoiceDefault2.0%Cloud
TypeLessDefault20.0%*Cloud

WER = Word Error Rate. Lower is better. These are aggregate scores across five recording types. *TypeLess applies mandatory LLM rewriting that cannot be disabled — its WER reflects reformatted output, not raw transcription errors.

The gap between Superwhisper’s worst default (S1-Voice at 22.8%) and its best model (Ultra at 1.6%) is 14 percentage points — within the same app, same subscription. Most people never switch models because the UI doesn’t surface the better ones clearly.

”I tried Superwhisper and it was terrible” — you were probably on S1-Voice or Parakeet

S1-Voice is Superwhisper’s cloud flagship — it’s what you land on if you open the app and start recording without changing anything. In our tests it dropped entire paragraphs, hallucinated content that wasn’t in the source audio, and produced non-deterministic output (37% WER on one run of the conference clip, 16% on the same audio a second time).

Parakeet, the default local model, is better — but it has no ITN (intelligent text normalisation). That means numbers come out as spoken words. “Twelve thousand four hundred dollars and seventy five cents” instead of “$12,400.75.” On our numbers recording, Parakeet scored 86% WER. On casual speech it scored 0.00%.

The model that actually works well — Whisper Standard at 1.8% WER — requires you to go to Settings → Library and search for it manually. It’s not in the default model picker.

Apps with one model remove the problem — and the ceiling

Wispr Flow, Willow Voice, and Aqua Voice each give you a single cloud model with no picker. You get whatever accuracy that model delivers and that’s it.

For most users, this is fine. Willow Voice’s single model hits 1.5% WER — better than Superwhisper’s default by a wide margin, and there’s nothing to configure. Aqua Voice comes in at 2.0%. Wispr Flow at 3.9%.

The trade-off is that you can’t improve. If your use case is coding dictation (where snake_case identifiers and tech terms matter), none of these apps let you switch to a model that handles it better. Superwhisper with Whisper Standard handles coding better than any cloud-only app we tested — but you have to find it first.

The model matters more on specific tasks than on aggregate scores

Aggregate WER is useful for comparison, but it hides a lot. Here’s what we found when we broke results down by recording type:

Numbers and dates — Parakeet (Superwhisper’s default local model) scores 86% WER because it has no ITN. Every other model we tested handles numbers correctly. If you dictate financial data, Parakeet is unusable regardless of how well it scores on casual speech.

Coding dictation — snake_case identifiers trip up almost every cloud model. “last_seen_at” becomes “last scene at” in Wispr Flow and Willow Voice. Whisper Standard handles it better, scoring 7.3% WER on coding versus 4.98%–7.39% for the cloud apps. Not a massive gap, but consistent.

Accented and conference speech — this is where cloud models pull ahead. Willow Voice hit 0.47% WER on our conference clip (accented English speaker, technical vocabulary). Superwhisper’s Ultra cloud model scored 0.45%. Whisper Standard local scored 0.91% — still excellent, but the cloud models have an edge on accent handling.

Long-form recordings — Superwhisper’s Ultra model scored 0.18% WER on our four-minute clip, the best long-form result of anything we tested. Whisper Standard local was 1.09%. Wispr Flow 3.13%. The gap between the best and the rest opens up on longer recordings.

What to actually look for when choosing

If you’re evaluating dictation apps, the question isn’t “which app is most accurate?” It’s “which model is active by default, and does it match my use case?”

If you want the best accuracy with no setup: Willow Voice’s single cloud model at 1.5% WER. Install, record, done.

If you want local processing with no audio upload: Superwhisper with Whisper Standard. Expect 1.8% WER and ~8 seconds of post-stop latency on CPU. You have to find the model in Library settings — it’s not the default.

If you’re on Superwhisper and it’s been bad: check which model is active. If it says S1-Voice or Parakeet, that’s your answer. Switch to Whisper Standard for local or Ultra for cloud.

If you need good numbers and dates: avoid Parakeet. Any cloud model or Whisper Standard handles ITN correctly.

If you dictate in noisy environments: noise resilience is consistent across all models we tested. Café-level background noise had no measurable effect on any of them — this is not a differentiator worth optimising for.

The uncomfortable truth about default settings

Every app defaults to something. The defaults are usually chosen for marketing reasons — fastest response, lowest latency, most impressive demo — not for accuracy across real-world use cases.

Superwhisper defaults to Parakeet locally and S1-Voice in cloud mode. S1-Voice is presented as the flagship AI feature. It scored 22.8% WER in our tests. Whisper Standard, the best local model, is buried.

This isn’t unique to Superwhisper. Aqua Voice’s Pro tier unlocks Avalon and Avalon 1.5 models — but we couldn’t test them because we tested the free tier. Willow Voice advertises a local/offline mode on Pro but it was locked in testing.

The lesson: the accuracy of a dictation app is not fixed. It depends entirely on what’s running under the hood. Check the model, not just the app name.

See all per-model numbers in our full reviews: Superwhisper · Wispr Flow · Willow Voice · Aqua Voice