by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

Why speech recognition fails on the phone

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

An engine that transcribes a read book almost perfectly misses roughly one word in five on the phone, for six reasons that add up — and in a ten-second call, one word in five is often the word that mattered.

The gap between the benchmark and the line

Vendors announce error rates of a few percent. They are true — on read, clean, wideband speech. On the phone, the same engines fail much more: an independent comparison from August 2026 sums up that “even the best engine misses roughly 1 word in 5” on conversational phone audio, and Deepgram itself writes that contact centres see 35 to 50% word errors in production, against 3 to 5% on clean benchmarks. The gap does not have one cause: it has six, and they add up.

1. The narrow band

The line only carries frequencies from 300 to 3,400 hertz, compressed to eight bits. High consonants vanish, “fifteen” sounds like “sixteen”. It is the root cause, the one the others graft onto; it has its own page.

2. Noise

A street, a car, a kitchen, an open-plan office: the person called is never in a studio. Deepgram cites measurements where background voices (“babble”) push the error rate from 5.5% when the voice clearly dominates to 15.2% when the noise is as loud as the voice. Noise changes from one second to the next, and a model trained on quiet sound has not learned to ignore it.

3. Accents

A model learns the accent of its data. A 2025 study evaluated five engines on English spoken by people from six different first languages: all of them degrade between read and spontaneous speech, and none holds across every accent at once. In French the question is the same, and nobody has measured it on the phone: a southern, northern, Belgian or North-African accent, and the “standard” model has only learned part of the country.

4. Overlaps

On the phone, people cut in. The person answers before the end of the question, the agent resumes before the end of the answer, two voices overlap for half a second. Engines transcribe one voice at a time; on an overlap they pick one, mix both, or leave a hole. And in a short call, the overlap lands precisely on the “yes” or “no” that was expected.

5. Proper nouns

A surname, a street, a company, a brand: the model does not know them, so it guesses a common word that sounds alike. For an agent that must record who it spoke to and where to call back, that is the most important word of the call, and the one the engine is least likely to have.

6. Numbers

A phone number, a postcode, a date, a time, an amount. Digits are short, said fast, often in series, and the narrow band makes them all sound a little alike. One wrong digit in a callback number, and the callback never happens — without anyone knowing.

What it costs in a ten-second call

On our September 2026 calls, half of the answered calls last ten seconds or less: about twenty words. One word in five missed is four or five words, and they do not fall at random: the names, the numbers, the “no” said over noise are exactly the words the six causes target. The transcript can look fine and the call be understood wrong.

Where ASDIC stands

We do not yet have an error rate measured on our calls: it needs a hand-transcribed corpus, which is the programme's pilot set. When it exists, each error will be filed under one of the six causes, to know which one to work on first — we do not yet know which weighs most in French on the phone, and we would rather say so. Audio examples per cause will come from that same corpus, anonymised, not before.

Calling back fast is useless if you understood wrong

An agent that calls back within a minute but records the wrong name or the wrong number has lost the contact more surely than a slow one. Nodical explains on its site how to call a lead back within 60 seconds — and what must be understood before calling back.

And in practice? calling back fast is useless if you understood wrong

Sources

accessed on 1 October 2026.

  1. Deepgram, « Noise-robust speech recognition » (centres d'appels : 35 à 50 % d'erreurs en production contre 3 à 5 % sur les bancs d'essai propres) — https://deepgram.com/learn/noise-robust-speech-recognition-techniques
  2. Vexascribe, « How accurate is Deepgram? » (« même le meilleur moteur rate à peu près un mot sur cinq au téléphone », août 2026) — https://vexascribe.com/how-accurate-is-deepgram
  3. arXiv 2503.06924, « Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling » (mars 2025) — https://arxiv.org/abs/2503.06924
  4. arXiv 2510.06961, « Open ASR Leaderboard » — https://arxiv.org/html/2510.06961v1
  5. ASDIC, « Audio 8 kHz : ce que les modèles n'entendent pas » — https://asdic.ai/savoir/audio-8-khz-ce-que-les-modeles-n-entendent-pas

Corrections

No correction so far.