ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.
LearnPhone voice AI in twenty words
Pages about voice agents on the phone speak a jargon made of acronyms; here are the twenty words that come up most, each defined in a few plain sentences and linked to the page of this site that goes further.
How to read this glossary
Voice agents have their own jargon, mostly acronyms. Here are twenty terms, ordered like a call rather than alphabetically, each linked to the page of this site that goes further. Third-party figures carry the name of their author; their source is at the bottom of the page. To see these words at work in a real call, read Ten seconds: what happens when a call is picked up.
The sound on the line
Telephone band
The frequencies a classic phone line carries: 300 to 3,400 hertz, according to the recommendations of the International Telecommunication Union (ITU). The rest is cut off, and high consonants such as “s” or “f” suffer most. See 8 kHz audio.
8 kHz
The telephone sampling rate: 8,000 measurements of the sound per second, a value set by ITU Recommendation G.711. Speech-recognition models generally learn on richer, so-called “wideband” sound, and rarely hear this kind. See 8 kHz audio.
G.711 (A-law, µ-law)
The historic coding of telephone voice, standardised by the ITU: 8,000 samples per second, eight bits each, which makes 64,000 bits per second. It comes in two variants, A-law and µ-law (“mu-law”), which compress sound differently; between two countries that do not use the same one, A-law is what travels. See 8 kHz audio.
Who speaks, and when
AMD (answering machine detection)
The “human or voicemail” decision at the start of an outbound call. Twilio states that its detector returns a result in about four seconds with default settings; on Nodical's calls, in September 2026, a voicemail is recognised after 8 seconds (median). See Detecting voicemail and The voicemail beep.
VAD (voice activity detection)
A detector that says, moment by moment, whether there is speech or not. It understands nothing: it separates voice from silence and, more or less well, from noise. See Silence on the phone.
Endpointing
The decision that a speech segment is over, usually because silence has exceeded a threshold. Nodical's published response delay (0.14 seconds median) is measured after that decision, so it does not count it. See End of turn.
End of turn
The decision that the person has really finished and that it is the agent's turn: “I'd like to…” followed by silence is not one. On its English validation data, LiveKit's eot-bench shows that a detector listening to silence alone wrongly cuts people off 55.6% of the time if it answers after 300 milliseconds. See End of turn.
Barge-in
Speaking while the other party is speaking. An agent must be interruptible and fall silent at once; when the agent is the one cutting in, it is a fault. The cut-off rate will be one of the measurements ASDIC publishes; it is not measured yet. See Measuring a voice agent.
Diarisation
Knowing who speaks when, by separating the agent's voice from the person's. It is simple when each side of the call is recorded on its own track, harder when voices overlap. See Why speech recognition fails on the phone.
How the agent is built
Full-duplex
A system that listens and speaks at the same time, instead of waiting its turn. Luqia Technologies, which published a French full-duplex benchmark in September 2026, describes such models as able to listen, speak, pause and respond during the conversation. See Pipeline or speech-to-speech.
Speech-to-speech
A single model that takes speech in and gives speech out, without going through text. It is contrasted with the pipeline, which chains speech recognition, a language model and speech synthesis, and can be checked step by step. See Pipeline or speech-to-speech.
Understanding what is said
WER (word error rate)
Words substituted, deleted and inserted, divided by the number of words spoken; ten words said, one replaced and one dropped, makes 20%. For Deepgram Nova-3, Vexascribe's comparison (August 2026) records 2.6% on clean read speech and 21.8% on CallHome, a corpus of phone calls. See Reading a benchmark.
SLU (spoken language understanding)
Going from what is said to what it means: the intent and the useful details. In French, the reference is the MEDIA corpus, distributed by ELRA since 2005 and now free for academic research. See Understanding what the caller wants.
Intent
What the person wants, in one label: “book an appointment”, “call me back”, “not interested”. MEDIA was annotated only in concepts; a version annotated with intents was published in 2024. See Understanding what the caller wants.
Measuring
Latency
The time the agent takes to react. The word is vague: depending on who measures, it starts at the person's last word or at the end-of-speech decision, and stops at the agent's first sound or first byte sent. Before comparing, ask where the measurement starts and ends. See Measuring a voice agent and Evaluating a voice AI vendor.
TTFAB (or TTFA)
Time to first audio byte: the time between the moment the person goes quiet and the agent's first sound. OpenBenchmarks measures it on real phone calls, both sides recorded on one clock; its readings (last call on 1 August 2026) give, for five platforms, medians from 1.3 to 1.7 seconds. See Measuring a voice agent.
Median, p90 and p95
The median splits calls into two halves: half do better, half do worse. The p90 is the value under which 9 calls out of 10 fall, the p95 the value under which 19 out of 20 fall: they describe the people who waited the most, which an average hides. See Evaluating a voice AI vendor.
The data the agent learns from
Annotated corpus
Recordings to which people have added by hand what needed to be known: the transcript, the moment a voicemail can be recognised, the ends of turns, the intents. It is what makes it possible to train a model and measure its errors; there is no public one for real French phone calls. See A French phone-call corpus.
Pseudonymisation
Processing data so that it can no longer be attributed to a person without additional information, kept separately and protected (GDPR, Article 4); for a call, it means removing names, numbers and addresses. Pseudonymised data remains personal data (Recital 26); only data rendered anonymous falls outside the GDPR. See A French phone-call corpus.
Voiceprint
A representation computed from a voice to recognise one specific person. The GDPR calls “biometric data” the data “resulting from specific technical processing” that allow or confirm the unique identification of a person; a model that recognises a voicemail or an intent does not need one, and ASDIC will not make any. See Sovereign AI.
These words in a real project
These twenty words are there to ask the right questions before letting a voice agent make or answer calls: how fast it knows who it is talking to, whether it cuts people off, what it understands, what happens to the recordings. To move on to practice, Nodical publishes Nodical's practical guides.
And in practice? Nodical's practical guides
Sources
accessed on 1 October 2026.
- UIT-T, Recommandation P.310, « Transmission characteristics for telephone band (300-3400 Hz) digital telephones » — https://www.itu.int/rec/T-REC-P.310
- UIT-T, Recommandation G.711, « Pulse code modulation (PCM) of voice frequencies » (8 000 échantillons par seconde, huit bits par échantillon, loi A et loi µ) — https://www.itu.int/rec/T-REC-G.711
- Twilio, « AMD FAQ & Best Practices » (résultat en environ 4 secondes avec les réglages par défaut) — https://www.twilio.com/docs/voice/answering-machine-detection-faq-best-practices
- LiveKit, eot-bench (classement et jeu de données) — https://github.com/livekit/eot-bench
- arXiv 2609.10765, Luqia Technologies, « French Full-Duplex Benchmark for Spoken Dialogue Models » (septembre 2026) — https://arxiv.org/abs/2609.10765
- Vexascribe, « How accurate is Deepgram? Nova-3 WER benchmarks (2026) » (mis à jour le 25 août 2026) — https://vexascribe.com/how-accurate-is-deepgram
- arXiv 2403.19727, MEDIA : nouvelle annotation en intentions (ELRA, distribué depuis 2005) — https://arxiv.org/abs/2403.19727
- OpenBenchmarks, « Voice agent latency benchmark (2026), TTFAB from real phone calls » (dernier appel : 1er août 2026) — https://openbenchmarks.com/voice-agent-latency
- Règlement (UE) 2016/679 (RGPD), article 4, définitions « pseudonymisation » et « données biométriques » — https://gdpr-info.eu/art-4-gdpr/
- RGPD, considérant 26 (données pseudonymisées et données anonymes) — https://gdpr-info.eu/recitals/no-26/
- ASDIC, carnet n° 1 : « Septembre 2026 : première mesure » — https://asdic.ai/carnet/2026-09-premiere-mesure
Corrections
No correction so far.