by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

Pipeline or speech-to-speech: two ways to put an AI on the phone

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

A voice agent can chain three models — one writes down what it hears, another drafts the answer, a third speaks it — or hand everything to a single model that listens and talks at the same time; the first is easier to control, the second reacts faster, and we know of no published measurement comparing the two on real phone calls in French.

Two ways to make a machine talk

To hold a conversation, a voice agent has to understand what is said, decide what to answer, and say it. There are two ways to organise this. The first chains three models: speech recognition that turns the voice into text, a language model that drafts an answer, and speech synthesis that speaks it. This is called a pipeline, or cascade. The second hands everything to a single model that takes in sound and puts out sound, without going through writing: speech-to-speech. Some of these listen and speak at the same time, as people do on the phone: they are called full-duplex. The choice decides how fast the agent answers, what happens when someone cuts in, what can be checked, and which languages it speaks.

The pipeline: hear, write, think, speak

In a pipeline, each step has its own model, and one can be swapped without touching the others. According to OpenAI's documentation, the chained path is for when you want to “inspect or transform intermediate text and replace each component independently”. Kyutai, a French lab, made the same choice for Unmute: its speech recognition (stt-1b-en_fr, which understands English and French) and its voice plug into any language model, and the lab writes that “the key strength of Unmute is its modularity”.

The price is time, and what gets lost along the way. Each step waits for the previous one: Kyutai explains that once the end of speech is detected, its system still has to wait an additional 500 milliseconds, the delay of its speech recognition, before it has the full text. And text does not keep tone, hesitation, the “yes” that means “no”. On the phone, where speech recognition fails far more often than in a studio, a transcription error becomes an understanding error.

The single model: from speech to speech

A speech-to-speech model takes in sound and answers in sound, with no written step. Moshi, released by Kyutai in 2024 with open weights, is presented by its authors as the first real-time full-duplex spoken large language model: it follows its own speech and that of the other party in two parallel streams, with a theoretical latency of 160 milliseconds, 200 in practice on an L4 graphics card. OpenAI offers a “Realtime” API that uses “one model to interpret audio, decide what to do, and respond in speech”, and also describes a third path: a model that listens and speaks at the same time and delegates reasoning and tool use to a separate backend.

The price is the reverse. Kyutai, which built both, says so: Moshi “doesn't yet match” text models on function calling, reasoning and in-context learning. And a single model only speaks the languages it learned, whereas a pipeline can switch speech recognition for each language.

What each one gains and loses

CriterionPipelineSingle model (speech-to-speech)
Response delayWaits add up from one step to the nextShort by design, no intermediate step
InterruptionsHandled by a separate detectorThe model listens while it speaks, if it is full-duplex
ControlReadable text, answer checkable before it is spokenNo intermediate text to review
Tools, business rulesThose of the chosen language modelBehind, according to Kyutai for Moshi
LanguagesOne speech recogniser per language, as you chooseThe model's own, and only those
Tone, hesitationLost in writingKept in the sound

What benchmarks measure, and what they don't

Artificial Analysis ranks speech-to-speech models, along with several providers' “default” cascaded systems, on reasoning (Big Bench Audio: 1,000 questions in English, read by synthetic voices), on pauses, turn-taking and interruptions (Full Duplex Bench), and on customer-service tasks (τ-Voice). Full-Duplex-Bench, published in March 2025, proposes a way to measure these behaviours; in 2026 Luqia built French versions, Canadian and European, and notes that optimising a benchmark's numbers can hurt conversational naturalness (see French corpora).

None of these measurements comes off a phone line. Moshi's 200 milliseconds are measured on a graphics card, not at the end of a call, and its audio codec, Mimi, processes sound sampled at 24 kHz, three times the 8 kHz of a phone line. On real calls, OpenBenchmarks records, for five platforms on the market, a time to first audio of 1.3 to 1.7 seconds (median). These numbers do not measure the same thing and cannot be compared; they remind us that lab latency is not call latency (see the five numbers that matter).

Where ASDIC fits

ASDIC is neither a pipeline nor a speech-to-speech model: it is an understanding model. It does not draft answers and does not speak. It is meant to listen to the call and make the decisions that keep a conversation together: human or voicemail, has the person finished their sentence, what do they want, should the call go to an adviser. These questions arise whatever architecture does the talking: in a pipeline, such a model can decide on the sound without waiting for the text; next to a single model, it provides separate decisions that can be measured.

This is a choice of method, not a result. Nodical's agents run today as a pipeline, with speech recognition and a voice supplied by an American provider, and Nodical has, to date, no understanding model of its own. On the September 2026 calls, the agent answers in 0.14 seconds (median) once it has decided the person has finished, and in 1.32 seconds for 9 answers out of 10; that delay excludes end-of-speech detection, which is precisely what ASDIC has to learn. Our target is written as a target: answer in under 0.5 seconds, interrupting less than 5% of the time. It is not reached; it will be published with the measurement, in the notebook.

What it changes during the call

For a company, the architecture matters less than what it makes possible during the call: booking an appointment, updating a record, alerting a salesperson. Those actions go through text and software, which is where, by Kyutai's own account, the single model still lags. Nodical's site details the workflows Nodical triggers during the call.

And in practice? the workflows Nodical triggers during the call

Sources

accessed on 1 October 2026.

  1. Kyutai, « Unmute » (reconnaissance vocale et voix branchées sur n'importe quel modèle de langue ; comparaison avec Moshi) — https://kyutai.org/unmute
  2. Kyutai, « Kyutai STT » (délai de 500 ms de stt-1b-en_fr, VAD sémantique, « flush trick » d'Unmute) — https://kyutai.org/stt/
  3. Hugging Face, kyutai/stt-1b-en_fr (anglais et français, délai de 0,5 s) — https://huggingface.co/kyutai/stt-1b-en_fr
  4. arXiv 2410.00037, Kyutai, « Moshi: a speech-text foundation model for real-time dialogue » (2024) — https://arxiv.org/abs/2410.00037
  5. Kyutai, dépôt Moshi (latence de 160 ms théorique, 200 ms sur carte L4 ; codec Mimi à 24 kHz) — https://github.com/kyutai-labs/moshi
  6. OpenAI, « Voice agents » (documentation : GPT-Live, Realtime API, chaîne) — https://developers.openai.com/api/docs/guides/voice-agents
  7. OpenAI, « Realtime API » (documentation) — https://developers.openai.com/api/docs/guides/realtime
  8. Artificial Analysis, classement speech-to-speech (Big Bench Audio, Full Duplex Bench, τ-Voice, « Default Cascaded System ») — https://artificialanalysis.ai/speech-to-speech
  9. Artificial Analysis, jeu de données Big Bench Audio (1 000 questions, anglais, voix de synthèse) — https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio
  10. arXiv 2503.04721, « Full-Duplex-Bench: a benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities » (mars 2025) — https://arxiv.org/abs/2503.04721
  11. arXiv 2609.10765, Luqia, « French Full-Duplex Benchmark » (EMNLP 2026 Industry) — https://arxiv.org/abs/2609.10765
  12. OpenBenchmarks, « Voice agent latency benchmark (2026), TTFAB from real phone calls » (dernier appel : 1er août 2026) — https://openbenchmarks.com/voice-agent-latency
  13. ASDIC, carnet n° 1 : « Septembre 2026 : première mesure » — https://asdic.ai/carnet/2026-09-premiere-mesure

Corrections

No correction so far.