by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

Evaluating a voice AI vendor: ten questions to ask

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

A sales sheet gives a number; ten simple questions tell you what it is worth — on which audio, in which language, over which line, as a median or at worst, measured by whom — and what becomes of the voices of the people called, where, for how long, and when everything stops.

Why questions rather than a ranking

No ranking tells you which voice agent will suit your calls. Published figures are often accurate, but measured elsewhere: other audio, another language, another line. Coval, a company that sells evaluation tools for voice agents, put it bluntly in June 2026: vendor benchmarks are “marketing copy with measurements attached”; the numbers are real, the conditions are picked to flatter. It files its objections under three questions: under what conditions, measured how, compared to what. We break them down here into ten. None requires an engineer; all call for a written answer, with names. This page recommends no product: it compares answers, not brands.

Five questions on measurement

1. What audio did you measure on?

Read books, videos, meetings, or real calls? Coval points out that standard test sets do not reflect “your scripts, your callers, your accents”. A good answer names the test set, gives its size and says whether it resembles your calls: outbound or inbound, short or long. “In real-world conditions”, with nothing more, is not an answer.

2. In which language?

A figure measured on one English accent does not cover them all. According to the paper that describes it (October 2025), the Open ASR Leaderboard, the most used public speech-recognition leaderboard, evaluates English on audiobooks, podcasts and videos, meetings, talks, European Parliament speeches and company earnings calls, and its other languages on read speech only. Ask for the figure in your callers' language and, if their accents vary — Scottish, Irish, Indian, Southern US —, on those accents.

3. Over the phone (8 kHz) or wideband?

A phone line only carries frequencies from 300 to 3,400 hertz, sampled at 8 kilohertz. The gap is no detail: in the paper introducing Whisper, the same model makes 2.7% word errors on read audiobooks and 17.6% on CallHome, a corpus of English phone calls. Ask for a figure measured on line audio (see 8 kHz audio: what the models can't hear).

4. The median, and the p95?

The median tells you what an ordinary call experiences; the p95, the value under which 95 measurements out of 100 fall, tells you what the people who wait the most experience. Coval calls a median latency measured under ideal conditions “a marketing number”. OpenBenchmarks, which times the response delay of five platforms on real phone calls, publishes both: in August 2026, medians of 1.3 to 1.7 seconds and p95s of 1.8 to 2.3 seconds. Also ask what is being timed, from when to when (see Measuring a voice agent: the five numbers that matter).

5. Who measured: you, or a third party?

A measurement made by the vendor on its own test set is not false; it is unverifiable. Has a third party repeated it? Is the method public, like those of OpenBenchmarks and the Open ASR Leaderboard, which publish their code? Is it repeated when the model changes? Coval reports vendor-pushed updates that do not change the version string. The safest course is still to measure on a sample of your own calls (see Reading a speech recognition benchmark without being fooled).

Four questions on data

6. Where does the audio go?

Telephone carrier, speech recognition, language model, speech synthesis, storage: the voice of a person called passes through several services. Ask, for each one, for the company and the country. “Hosted in Europe” may cover only the database (see Sovereign AI: what it means for a phone conversation).

7. How long is it kept?

Recordings, transcripts, summaries: each has its own retention period. For call recording, the CNIL, the French data-protection authority, states that retention varies with the purpose pursued; for proof of a contract, it notes that the ordinary limitation period in France is five years. A serious answer gives a period per use, says what happens at the end (deletion or anonymisation) and states whether your calls are used to train the vendor's models.

8. Who are your subprocessors?

The GDPR (Article 28) forbids a processor from engaging another without the controller's written authorisation, and requires it to give notice of any change so that the controller can object. Ask for the list by name, with each one's country and role. A published list is worth more than one “available on request”.

9. How does a person called object?

It is not up to the person called to work out how to say no. The CNIL asks for an oral notice at the start of the conversation announcing the recording and its purpose and informing people that they can object. Since 2 August 2026, the EU AI Act (Article 50) also requires that people know they are talking to an AI. Ask what the agent says in the first seconds, what happens if the person refuses, and how their number is then excluded.

One question on outages

10. What happens when something goes down?

If a link in the chain fails, what becomes of the call in progress? Does the agent go silent, hang up, hand over to a human or to another provider? What happens to scheduled calls, and who is alerted? Ask for the story of a past outage. A vendor that has never had one does not exist; a vendor that can tell you about its own does.

Answers that should raise an eyebrow

  • An accuracy rate with no test set, no language and no line quality.
  • A single “latency”, without saying what is timed and with no p95.
  • “Hosted in Europe” with no list of subprocessors.
  • “GDPR compliant” as the answer to a precise question.
  • Nothing in writing about what happens during an outage.

None of these answers proves a product is bad. They only prove that nobody knows yet whether it is good.

Our own answers, on measurement

We ask ourselves the same questions. Which audio? Our production calls: in September 2026, 60,287 calls answered by a human, 14 campaigns, in French, measured without reading the content of conversations. Median and high value? Voicemail is recognised at 8 seconds (median), and at 14 seconds for 9 calls out of 10. Who measured? Only us: no third party has repeated the measurement. How much of the chain is ours? No comprehension model of our own today — 0% —; the voice and the speech recognition rely on an American provider, named with the other subprocessors in Nodical's privacy policy. The detail is in notebook entry no. 1.

What Nodical answers to these ten questions

The last questions concern the service more than the research: how a call is billed, when the agent hands over to a human, what record is kept of each call, how you stop. Nodical publishes its answer by comparing a call centre and a voice agent point by point: what Nodical answers to these ten questions.

And in practice? what Nodical answers to these ten questions

Sources

accessed on 1 October 2026.

  1. Coval, « Best STT providers 2026: independent benchmarks and how to choose » (4 juin 2026) — https://www.coval.ai/blog/best-speech-to-text-providers-in-2026-independent-benchmarks-and-how-to-choose/
  2. OpenBenchmarks, « Voice agent latency benchmark (2026), TTFAB from real phone calls » (dernier appel : 1er août 2026) — https://openbenchmarks.com/voice-agent-latency
  3. arXiv 2510.06961, « Open ASR Leaderboard » (jeux de test, octobre 2025) — https://arxiv.org/html/2510.06961v1
  4. OpenAI, « Robust Speech Recognition via Large-Scale Weak Supervision » (article Whisper, tableau 8) — https://cdn.openai.com/papers/whisper.pdf
  5. CNIL, « L'enregistrement des conversations téléphoniques afin d'établir la preuve de la formation d'un contrat » — https://www.cnil.fr/fr/lenregistrement-des-conversations-telephoniques-afin-detablir-la-preuve-de-la-formation-dun-contrat
  6. Règlement (UE) 2016/679 (RGPD), article 28 (sous-traitant) — https://gdpr-info.eu/art-28-gdpr/
  7. AI Act, article 50 (transparence : informer la personne qu'elle parle à une IA), applicable au 2 août 2026 — https://artificialintelligenceact.eu/article/50/
  8. Nodical, politique de confidentialité (liste des sous-traitants) — https://nodical.ai/confidentialite

Corrections

No correction so far.