ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.
LearnEnd of turn: waiting for silence is not enough
A voice agent must decide at every moment whether the person has finished their sentence or is catching their breath; waiting for a fixed silence means interrupting if you wait little and making people wait if you wait long, and no published measurement says how this plays out in French, over the phone.
Two questions that get confused
“Is he still talking?” and “Is he done?” are not the same question. The first is acoustic: there is sound, or there is not. That is the job of a voice activity detector (VAD), which marks the end of a speech segment when silence exceeds a threshold. The second is a matter of meaning and music: “I'd like to book an appointment…” followed by silence is not the end of a sentence, it is a hesitation; “No thanks.” followed by the same silence is. Humans decide without thinking, from intonation, rhythm and words. A machine that listens only to silence gets it wrong both ways.
The trade-off silence imposes
With a silence threshold, everything comes down to one number: how many milliseconds to wait. Wait little, and the agent interrupts at every breath. Wait long, and every answer arrives after a gap the person reads as slowness, or as a dropped line. LiveKit's open benchmark, eot-bench, quantifies this trade-off for a detector that listens to silence only (on its English validation data): answering after 300 milliseconds of silence, you wrongly cut people off 55.6% of the time; after 600 milliseconds, still 21.7%; to cut off only one turn in twenty, you have to wait 1.6 seconds. A model that also listens to words and intonation does much better: for its own, LiveKit reports in its June 2026 post 9.9% false cut-offs with a 300-millisecond budget and 4.5% with 600. Those figures are theirs, on their data; above all they say one thing: you do not escape this trade-off by tuning the threshold, you escape it by listening to something other than silence.
What the people working on it publish
LiveKit released an end-of-turn model and an open benchmark, eot-bench, whose dataset covers fourteen languages, French and Spanish included. The post gives per-language results, but says nothing about telephony: the conversations are wideband recordings between a human and an agent, never over a line. Daily (Pipecat) offers Smart Turn v3, an open model (BSD licence) of eight megabytes that runs in about twelve milliseconds on an ordinary CPU and covers twenty-three languages; its corpus was collected from the community, not from real calls, and again no per-language or telephone measurement is published. Picovoice wrote the most pedagogical guide on latency, turn-taking and barge-in, for on-device assistants, in English. Kyutai, a French lab, publishes streaming speech recognition with a “semantic VAD” that decides in about half a second whether the sentence is over — open, French-English, but with no telephone measurement. On the research side, two 2026 papers show the direction: one learns to predict the length of the coming silence rather than wait it out (Next-Turn), the other asks which aspects of speech really carry end-of-turn information. Finally, Luqia published in September 2026 the only French dialogue benchmark we know of, which says enough about how empty the field is.
Why the telephone changes everything
Over the phone, sound is narrow (300–3,400 Hz) and compressed, packets get lost, and intonation — precisely what says “I'm done” — is harder to hear. The models cited learned on wideband speech; nothing says how they behave on a line, in French, with someone talking fast because they are in a hurry to hang up. It is one of the reasons 8 kHz telephone audio has its own page on this site.
Where ASDIC stands
In September 2026, on Nodical's calls, the agent answers 0.14 seconds after deciding the person had finished (1.32 seconds for 9 answers out of 10). That figure excludes the decision itself: it measures the generation of the answer, not the wait. So it is the end-of-turn decision that makes people wait, and it is the one we cannot yet measure without a hand-annotated corpus — one that records, for each turn, the moment the person had really finished. Our target is written as a target: answer in under 0.5 seconds, interrupting less than 5% of the time. It is not reached; it will be published with the measurement that proves it. The method and limits of the September measurement are in the notebook.
What it changes during a call
An agent that interrupts makes people hang up; an agent that waits too long makes them doubt anyone is there. Between the two lies the difference between a conversation and a voice form. Nodical describes on its site what its agents do today during the call.
And in practice? what Nodical does today during the call
Sources
accessed on 1 October 2026.
- LiveKit, « Solving end-of-turn detection: Turn Detector v1.0 » (17 juin 2026) — https://livekit.com/blog/solving-end-of-turn-detection
- LiveKit, eot-bench (classement et jeu de données) — https://github.com/livekit/eot-bench
- Daily, « Announcing Smart Turn v3 » — https://www.daily.co/blog/announcing-smart-turn-v3-with-cpu-inference-in-just-12ms/
- Picovoice, « Voice agent latency, turn-taking, and barge-in » — https://picovoice.ai/guide/voice-agents/voice-ux-latency-turn-taking/
- Kyutai, reconnaissance vocale en flux avec « semantic VAD » — https://kyutai.org/stt/
- arXiv 2606.18094, « Next-Turn: duration-aware streaming endpoint detection » — https://arxiv.org/pdf/2606.18094
- arXiv 2609.11066, « Less can be More: what aspects of speech drive end-of-turn detection » — https://arxiv.org/pdf/2609.11066
- arXiv 2609.10765, Luqia, « French Full-Duplex Benchmark » (septembre 2026) — https://arxiv.org/abs/2609.10765
Corrections
No correction so far.