by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

Reading a speech recognition benchmark without being fooled

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

A word error rate only means something with three details — on which audio, with which text normalisation, measured by whom — and the same model can show 2.7% on read audiobooks and 17.6% on phone calls without either figure being false.

What the word error rate counts

Almost every speech-recognition comparison gives a single number: the word error rate, or WER. The machine's transcript is aligned with a reference transcript written by hand, and three kinds of errors are counted: substitutions (a word replaced by another), deletions (a word spoken but missing) and insertions (a word written but never spoken).

WER = (substitutions + deletions + insertions) ÷ number of words in the reference

Because insertions add up without limit, the WER can exceed 100%: if the person says “yes” and the machine writes “yes yes that's right”, that is three insertions for one word, or 300%.

A worked example, by hand

The person says: “yes I am free on Thursday at ten” — 8 words. The machine writes: “uh yes I free on Tuesday at ten”. Align word by word:

ReferenceTranscriptVerdict
—uhinsertion
yesyescorrect
IIcorrect
am—deletion
freefreecorrect
ononcorrect
ThursdayTuesdaysubstitution
atatcorrect
tentencorrect

One substitution, one deletion, one insertion: 3 errors for 8 words, or 3 ÷ 8 = 37.5%. The three are not equal: “uh” and the missing “am” change nothing in the meaning, “Tuesday” sends the appointment to the wrong day. The WER counts them the same.

Normalisation, the rule of the game nobody reads

If hesitations are removed before counting, “uh” disappears: 2 errors remain, and the WER drops to 2 ÷ 8 = 25%, without a single letter of the transcript changing. Conversely, had the machine written “10” instead of “ten”, a naive count would add one error: 4 ÷ 8 = 50%, for a perfectly correct time. Contractions do the same: “I'm” against a reference reading “I am” counts as two errors.

That is why benchmarks “normalise” both texts before counting. For English, the Open ASR Leaderboard, the most used public leaderboard, removes punctuation and capitals, spells numbers out, standardises spelling and deletes filler words. These choices matter: the Whisper authors write that their normalisation cut WER by up to half on some test sets, often because of a quirk in the references such as contractions split by a space, and warn that a normaliser developed alongside one model risks favouring it. Vexascribe, which compiles comparisons, sums it up: WER comparisons are only valid when the same audio and the same text normalisation are used for every engine.

Read versus conversational speech

The paper introducing Whisper gives for its largest model (table 8) 2.7% errors on LibriSpeech, read audiobooks, but 13.8% on Switchboard and 17.6% on CallHome, two corpora of English telephone conversations. It shows it even more sharply (table 2): another model, wav2vec 2.0, has the same LibriSpeech score, 2.7%, but 34.8% on CallHome. On read books, a tie; on the phone, one makes twice as many errors as the other. The authors conclude that models should be evaluated outside their training conditions, to avoid overstating them.

Public leaderboards inherit this choice. According to the paper that describes it (October 2025), the Open ASR Leaderboard evaluates English on audiobooks, podcasts and videos, meetings, talks, European Parliament speeches and company earnings calls, its other languages on read speech, and has no track devoted to the telephone. Its authors say it: there is no catch-all model, no single dataset is sufficient, and WER alone is not enough.

16 kHz versus 8 kHz

LibriSpeech, the most cited set, is read speech recorded at 16 kilohertz. A phone line only carries frequencies from 300 to 3,400 hertz, sampled at 8 kilohertz and compressed. For Deepgram Nova-3, Vexascribe records 2.6% errors on LibriSpeech and 21.8% on CallHome, and expects 18 to 24% on 8 kilohertz phone conversation: more than eight times as many errors between the read book and the call. A figure measured on wide, clean sound is not a telephony figure (see 8 kHz audio: what the models can't hear and the six causes of errors on the phone).

Who measured: the vendor or a third party

One model can carry three figures without any of them being false. Deepgram announces for Nova-3 a median WER of 5.26% in batch transcription, on a test set it assembled: 2,703 audio files, nine domains, 81.69 hours. Vexascribe records 2.6% on LibriSpeech and 21.8% on CallHome for the same model, and points out that a median across test sets hides the worst domains. Coval, which sells evaluation tools, adds a trap: in its view, streaming WER is typically 1 to 3 points worse than batch WER on the same model, so a “batch” figure does not compare with a “streaming” one. The right attitude is not suspicion but the question.

What an average hides

An “average” WER can be computed in two ways. First sentence: the person says “no thanks”, the machine writes “yes thanks” — one error in two words, 50%. Second sentence: eighteen words, all correct, 0%. The average of the two sentences is 25%; total errors over total words is 1 ÷ 20 = 5%. Both are right, and neither says that the only error turned a refusal into a yes. An average across test sets likewise mixes unrelated conditions: in the Whisper paper, the largest model's 12.8% average (table 2) covers 2.7% on read books and 36.4% on meetings picked up by a single distant microphone.

Six questions before believing a figure

  • Which test set? Read or spontaneous, books or calls, in which language.
  • Which sound? Wideband, or an 8 kilohertz phone line.
  • Which normalisation? Hesitations, numbers, capitals: removed or counted.
  • Which mode? Batch or streaming.
  • Which calculation? Total errors, average per sentence, median per domain.
  • Who measured? The vendor or a third party, and is the method published.

We do not yet have a WER measured on our calls: it needs a hand-transcribed corpus, which is the programme's pilot set. When it exists, the figure will be published in the notebook with these six answers. Until then, we only quote third-party figures, with their name and source. For the questions that go beyond transcription: Evaluating a voice AI vendor: ten questions to ask.

What Nodical can do today

A word error rate does not tell you whether a call achieved anything: an appointment booked on the right day, a customer called back, a file brought up to date. That is what the page what Nodical can do today describes, around the voice agent.

And in practice? what Nodical can do today

Sources

accessed on 1 October 2026.

  1. OpenAI, « Robust Speech Recognition via Large-Scale Weak Supervision » (article Whisper : section 3.2 sur la normalisation, tableaux 2 et 8) — https://cdn.openai.com/papers/whisper.pdf
  2. arXiv 2510.06961, « Open ASR Leaderboard » (jeux de test, normalisation du texte, octobre 2025) — https://arxiv.org/html/2510.06961v1
  3. Vexascribe, « How accurate is Deepgram? Nova-3 WER benchmarks (2026) » (mis à jour le 25 août 2026) — https://vexascribe.com/how-accurate-is-deepgram
  4. Deepgram, « Introducing Nova-3: setting a new standard for AI-driven speech-to-text » (jeu de test de l'éditeur) — https://deepgram.com/learn/introducing-nova-3-speech-to-text-api
  5. Coval, « Best STT providers 2026: independent benchmarks and how to choose » (4 juin 2026) — https://www.coval.ai/blog/best-speech-to-text-providers-in-2026-independent-benchmarks-and-how-to-choose/
  6. OpenSLR, LibriSpeech ASR corpus (« environ 1 000 heures de parole anglaise lue à 16 kHz ») — https://www.openslr.org/12

Corrections

No correction so far.