ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.
LearnMeasuring a voice agent: the five numbers that matter
Voice-agent comparisons almost all measure one thing, the time before the first sound, while an outbound call rides on five decisions; we publish those five measurements, as a median and for 9 calls out of 10, from our own calls.
What the comparisons measure, and why it is not enough
When voice agents are compared, what is measured is almost always latency: the time between the moment the person goes quiet and the moment the agent starts speaking. The most serious benchmark we know, OpenBenchmarks, calls it TTFAB (time to first audio byte) and measures it properly, on real phone calls, recording both sides on one clock. Its August 2026 readings give, for five platforms on the market, medians between 1.3 and 1.7 seconds and p95s between 1.8 and 2.3 seconds. That is useful, and credit is due. But it is one number, and it says nothing about whether the agent recognised a voicemail, cut someone off, handed over at the right moment, or understood anything at all. An agent can answer fast and do nonsense.
The five numbers we publish
1. The time to know who you are talking to
How long after pick-up the agent knows it is talking to a human or a voicemail. It is the first decision of the call, and it conditions all the others. Today: 8 seconds (median), 14 seconds for 9 calls out of 10. Target: under 2 seconds.
2. The end-of-turn delay
How long between the person's last word and the agent's first — the real wait, the one people feel. It is the number closest to the comparisons' “latency”, with one difference: it must include the decision that the sentence is over, not only the making of the answer. Today we only measure the second part (0.14 seconds median, 1.32 seconds for 9 answers out of 10); the full measurement needs an annotated corpus. Target: under 0.5 seconds.
3. The cut-off rate
The share of turns where the agent spoke while the person had not finished. This number goes hand in hand with the previous one: you can always answer faster by cutting off more often. LiveKit's eot-bench shows the size of the trade-off for a detector that listens to silence only: 55.6% false cut-offs when answering after 300 milliseconds, 21.7% after 600, on its English data. We have not measured it on our calls yet. Target: under 5%.
4. The wrong-handover rate
When the agent passes the call to a human, was it the right moment? Too early wastes an adviser; too late, the person has hung up. Today: 1.19% of answered calls are handed over, at 135 seconds (median); we do not yet know what share of those handovers was relevant. Target: more than 9 handovers out of 10 judged relevant.
5. Call duration
The simplest number, and the most forgotten. It sets the scale for all the others: on our September 2026 calls, half of the answered calls last 10 seconds or less, and 8 out of 10 under 20 seconds. A decision that comes at 8 seconds comes, for many calls, after the conversation has ended.
How to read a median and a p90
We never publish an average: a few very long calls distort it. The median is the value that splits calls into two halves: half do better, half do worse. The p90 is the value under which 9 calls out of 10 fall; it describes the experience of the people who waited the most. An agent can have a good median and a bad p90: fast in general and sometimes very slow, which, on the phone, gets noticed. Serious benchmarks publish both; commercial pages usually publish only the better one.
What we do not measure yet
Of these five numbers, two are measured today (the human-or-voicemail delay, the duration), one is half measured (the end-of-turn delay, excluding the decision), and two await a hand-annotated corpus (cut-offs, handover relevance). The notebook says at each entry which one has moved. We would rather publish “not measured yet” than a number borrowed from elsewhere.
What Nodical does with these measurements
These five numbers are the ones a customer looks at, whether they know it or not: a call that runs long costs money, a cut-off makes people hang up, a missed handover loses a contact. Nodical explains on its site what it does with these measurements in its campaigns.
And in practice? what Nodical does with these measurements
Sources
accessed on 1 October 2026.
- OpenBenchmarks, « Voice agent latency benchmark (2026), TTFAB from real phone calls » (dernier appel : 1er août 2026) — https://openbenchmarks.com/voice-agent-latency
- LiveKit, eot-bench (classement et jeu de données) — https://github.com/livekit/eot-bench
- Hamming.ai, « Voice AI latency: what's fast, what's slow » — https://hamming.ai/resources/voice-ai-latency-whats-fast-whats-slow-how-to-fix-it
- Telnyx, « Voice AI agents compared on latency » — https://telnyx.com/resources/voice-ai-agents-compared-latency
- ASDIC, carnet n° 1 : « Septembre 2026 : première mesure » — https://asdic.ai/carnet/2026-09-premiere-mesure
Corrections
No correction so far.