by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

8 kHz audio: what the models can't hear

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

A phone line only carries frequencies from 300 to 3,400 Hz, compressed and sometimes full of holes, while speech-recognition models learned on wide, clean sound; as a result, third-party benchmarks measure five to six times more errors over the phone than on read speech.

What a phone line lets through

The telephone was designed so that people understand each other, not so that they hear each other well. It only carries frequencies between roughly 300 and 3,400 hertz, which translates into sound sampled at 8 kilohertz — half the 16 kilohertz most speech-recognition models learned on, four to six times less than a studio recording. The high consonants, the “s”, the “f”, the “t”, live above that limit: over the phone, “fifteen” and “sixteen” sound alike. Add compression (G.711, A-law in Europe, µ-law in North America), which squeezes each sample to eight bits, and packet loss on mobile and IP networks, which punches twenty-millisecond holes in the signal.

Why a “clean 16 kHz” model collapses

A model learns what it is shown. The big public corpora are made of read books, videos, wideband speech; telephone audio is rare in them, French telephone audio rarer still. Given narrow, compressed, holed sound afterwards, the model does not have bad reflexes: it has no reflexes at all. It resamples, it guesses, and it fails mostly on what matters in a short call: a yes, a no, a number, a name.

What third parties measure

Public figures tell the same story from one model to the next. The paper introducing Whisper, the open reference model, reports for its largest model 2.7% word errors on clean read speech, but 13.8 to 17.6% on English conversational telephone corpora (table 8). For Deepgram Nova-3, an independent comparison updated in August 2026 records 2.6% on clean read speech and 21.8% on CallHome, a corpus of phone calls; its author sums up that “even the best engine misses roughly 1 word in 5 on conversational phone audio”, and expects 18 to 24% errors on 8 kilohertz conversation. Google, for its part, offers a separate transcription model for phone calls, a sign that the case stands apart. And Coval, which publishes independent comparisons, reminds that standard benchmarks reflect “neither your scripts, nor your callers, nor your accents”: the only measurement that counts is the one made on your own calls.

Two things are missing from all this: none of these figures is in French, and none is in Spanish. The most used public leaderboard, the Open ASR Leaderboard, has a multilingual track but no telephone track.

What it changes for an agent that makes calls

One word in five lost is not a slightly uglier transcript: it is a missed intent, an appointment written on the wrong day, a “no” taken for a “yes”. And since our calls last ten seconds (median), there is no second chance to recover from context. That is why ASDIC learns on telephone audio from the start, in French, Spanish and English, rather than adapting a wideband model afterwards.

Where ASDIC stands

We do not yet have a word-error-rate measurement on our calls: it needs a hand-transcribed corpus, which is precisely the programme's pilot set. When it exists, it will be published in the notebook, with the method, in French first, then Spanish and English, on 8 kilohertz sound as it arrives on the line. To our knowledge it will be the first public measurement of its kind in French.

Why call volume changes the game

A telephony model trains on calls, and nobody has as many as a platform that places them every day. Nodical explains on its site why call volume changes the game for agencies and call centres.

And in practice? why call volume changes the game

Sources

accessed on 1 October 2026.

  1. OpenAI, « Robust Speech Recognition via Large-Scale Weak Supervision » (article Whisper, tableau 8) — https://cdn.openai.com/papers/whisper.pdf
  2. Vexascribe, « How accurate is Deepgram? Nova-3 WER benchmarks (2026) » (mis à jour le 25 août 2026) — https://vexascribe.com/how-accurate-is-deepgram
  3. Google Cloud, « Sélectionner un modèle de transcription » (modèle dédié aux appels téléphoniques) — https://docs.cloud.google.com/speech-to-text/docs/v1/transcription-model?hl=fr
  4. Deepgram, « Noise-robust speech recognition » — https://deepgram.com/learn/noise-robust-speech-recognition-techniques
  5. Coval, « Best STT providers 2026: independent benchmarks » (sur la limite des bancs d'essai standard) — https://www.coval.ai/blog/best-speech-to-text-providers-in-2026-independent-benchmarks-and-how-to-choose/

Corrections

No correction so far.