ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.
LearnSilence on the phone: noise, breath, hesitation
A voice activity detector decides, every few tens of milliseconds, whether it hears speech or not; on the phone, noise, breathing and hesitation fool it, and the most quoted figures, Picovoice's on its own benchmark, concern read English, not phone calls.
What a voice activity detector does
Before understanding what someone says, a voice agent has to know whether they are saying anything. That is the job of the voice activity detector (VAD): it cuts the sound into slices a few tens of milliseconds long and returns, for each one, a probability — speech or not speech. Above a threshold, the agent considers the person is talking; below it, that they are silent. Everything else depends on it: when to transcribe, when to stop if the person cuts in, when to start counting the silence.
The WebRTC project's detector, open-sourced by Google, relies on classic signal processing; more recent ones are neural networks, such as Silero (open) or Cobra (sold by Picovoice).
What is not silence
On the phone, the “silence” between two sentences is almost never silent. The detector has to tell apart:
- noise: a street, a car, a television. Other people's voices are the worst: it is speech, but not from the person called;
- breathing: an intake of air before answering: not speech, but often a sign that speech is coming;
- hesitation: “um”, “er”, “well…”: speech without useful words, which still says something important — the person has not finished;
- the line itself: crackle, the agent's own voice echoing back through the person's phone.
Taking noise for speech makes the agent fall silent for nothing. Taking speech for noise means missing the start of a “no”, or cutting off someone resuming after a hesitation.
How a detector is measured
A detector is tuned with a threshold, and every threshold gives two numbers. The detection rate (or true positive rate): the share of speech slices recognised as speech. The false alarm rate (or false positive rate): the share of non-speech slices wrongly taken for speech. Lowering the threshold raises both. Two detectors can therefore only be compared at the same false alarm rate, or on the curve linking the two numbers across all thresholds (the ROC curve). An “accuracy” percentage published alone, without the matching false alarm rate, says nothing.
What Picovoice publishes
Picovoice, a Canadian publisher of on-device voice software, maintains the most quoted benchmark on the subject. It should be read for what it is: Picovoice's own benchmark, built to compare its own product, Cobra, with two open detectors, WebRTC's and Silero (version 5.1). Its code and data are public, so it can be rerun. The method: read English speech from LibriSpeech, an audiobook corpus recorded at 16 kHz, mixed with real-world noise from the DEMAND corpus (18 environments, including kitchen, office and traffic), at a signal-to-noise ratio of 0 dB — noise as loud as the voice.
The results, detailed by Picovoice in a comparison published in November 2025 and updated in January 2026, are read at a fixed false alarm rate. They are detection rates per slice of sound, not success rates per sentence or per call:
| Detector | Detection rate at 5% false alarms | Detection rate at 1% false alarms |
|---|---|---|
| Cobra (Picovoice) | 98.9% | 95% |
| Silero 5.1 | 87.7% | 80.4% |
| WebRTC | 50% | not published (rate deemed too low by Picovoice) |
Picovoice benchmark: LibriSpeech test-clean (read English, 16 kHz) mixed with DEMAND noise at 0 dB. Figures published by Picovoice, not rerun by ASDIC.
Picovoice itself notes that the ranking depends on the threshold: at 25% false alarms, WebRTC does better than Silero. Its conclusion is the right one, and it applies to its own figures too: evaluate a detector at the false alarm rate your application can tolerate.
Why this is not yet the phone
This benchmark measures something real, robustness to noise, in conditions that are not those of a call. The speech is read, not conversational: no hesitation, no “um”, no restarts. It is in English. It is wideband, at 16 kHz, whereas a phone line cuts everything above 3,400 Hz (see the page on 8 kHz audio). The noise is added afterwards, not picked up with the voice and then compressed by the line. Silero, for its part, accepts audio at 8 kHz as well as 16 kHz, but its own published evaluations are run at 16 kHz. To our knowledge, nobody has published a comparison of detectors on real phone calls, in French or in Spanish.
Detecting speech does not tell you when to answer
A perfect detector would say exactly when the person is talking and when they are silent. It still would not say whether they have finished. “I'd like to… um…” followed by silence is not the end of a sentence; “No thanks.” followed by the same silence is. LiveKit's eot-bench quantifies the cost of answering on silence alone, on its English validation data: after 300 milliseconds of silence, you wrongly cut people off 55.6% of the time; after 600 milliseconds, still 21.7%; and you have to wait 1.6 seconds to cut off only one turn in twenty. VAD is the starting point of the end-of-turn decision, not the decision itself; that is the subject of the page on end of turn.
Where ASDIC stands
We have no voice activity detection measurement on our calls: it would take calls in which every passage is marked by hand — speech, noise, breathing, hesitation. What we measure today comes afterwards: in September 2026, the agent answers 0.14 seconds (median; 1.32 seconds for 9 answers out of 10) after deciding the person had finished, excluding end-of-speech detection. Our target is written as a target: answer in under 0.5 seconds, interrupting less than 5% of the time. It is not reached; it will be published with the measurement that proves it, in the notebook.
A line that picks up at any hour
A line that answers day and night hears everything: calls from a car or a building site, people searching for their words. The difference between silence, noise and hesitation plays out in every sentence. Nodical explains on its site how its agents pick up day and night.
And in practice? how its agents pick up day and night
Sources
accessed on 1 October 2026.
- Picovoice, « Voice Activity Detection Benchmark » (page officielle du banc d'essai : Cobra, WebRTC VAD, Silero VAD ; courbe ROC) — https://picovoice.ai/docs/benchmark/vad/
- Picovoice, dépôt « voice-activity-benchmark » (LibriSpeech test-clean mélangé au bruit DEMAND, courbes tracées à 0 dB de rapport signal sur bruit ; Silero version 5.1) — https://github.com/Picovoice/voice-activity-benchmark
- Picovoice, « Choosing the Best Voice Activity Detection in 2026: Cobra vs Silero vs WebRTC VAD » (12 novembre 2025, mis à jour le 23 janvier 2026 ; taux de détection à 5 % et 1 % de fausses alertes) — https://picovoice.ai/blog/best-voice-activity-detection-vad/
- Silero VAD, dépôt officiel (licence MIT, 8 000 et 16 000 Hz, moins d'1 ms par tranche de 30 ms et plus sur un seul fil d'exécution du processeur) — https://github.com/snakers4/silero-vad
- OpenSLR, LibriSpeech (« environ 1 000 heures d'anglais lu à 16 kHz ») — https://www.openslr.org/12/
- LiveKit, eot-bench (classement et jeu de données) — https://github.com/livekit/eot-bench
- ASDIC, « Fin de tour de parole : attendre le silence ne suffit pas » — https://asdic.ai/savoir/fin-de-tour-de-parole-attendre-le-silence-ne-suffit-pas
Corrections
No correction so far.