by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

Understanding what the caller wants, without going through text

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

Understanding someone on the phone is not writing down what they say but knowing what they want; research can do it straight from the sound on hotel bookings or commands to a home assistant, but nobody has yet published a measurement on a ten-second outbound call, in French, over a phone line.

Writing is not understanding

When someone answers “Oh, no, not right now”, a perfect transcript gives five correct words but does not say what to do: is the person declining, or asking to be called back? Spoken language understanding — SLU, as researchers call it — answers that question. It looks for the intent, meaning what the person wants (to decline, to accept, to be called back, to ask a question), and the information that goes with it (a day, a time, a name). For an agent that makes calls, that decision is what matters, not the text: a perfect transcript misread does as much damage as a wrong one.

Two paths: through text, or straight from sound to meaning

The classic method is a cascade: speech recognition writes down what it hears, then a second model reads that text and extracts the intent. It inherits every error of the first step: on the phone, where transcription fails far more often than in a studio, a misheard word becomes a misunderstood intent. And text erases tone, hesitation, a voice that drops.

The other method goes straight from sound to meaning, with no intermediate transcript; it is called “end-to-end”. The authors of SLURP, an English corpus published in 2020, write that it infers meaning “directly from audio data” and promises to reduce error propagation. Those of LeBenchmark 2.0, a French project, note that such systems are now preferred to cascades for that reason, and because they can draw on acoustic cues. It is the path ASDIC has chosen to explore, with one caution: preferred does not mean better everywhere. The same choice arises for the whole agent, between a chain of models and a single model.

MEDIA, the French reference

In French, almost all work is compared on MEDIA. Distributed by ELRA since 2005 and free for academic research, the corpus is described by the team that updated it in 2022 as the most challenging of those available to the research community. According to the paper that added eleven intents to it in 2024, it holds 1,258 recorded phone dialogues, 250 speakers and about 70 hours of conversation, on a single topic: hotel booking. The dialogues were collected with the “Wizard of Oz” method: the person believes they are talking to a machine, but a hidden human answers.

Its limits fit in three words: one topic, one era, one staging. The data dates from the 2000s, the dialogues are simulated, and a booking takes time: 70 hours for 1,258 dialogues makes more than three minutes per dialogue on average (our calculation). The full inventory is on our page on French corpora.

What the results say

The 2024 paper compares the two paths on MEDIA. For precise information — dates, number of rooms, prices — the error is higher when going straight from sound to meaning: the best end-to-end system gets 18.30% of that information wrong on the full version of the task. For intent, this is not the case: the direct path gets the best scores on almost every version of the corpus, and about 9 intents out of 10 are recognised correctly whichever path is used. The authors consider the gap small, and the cascade remains better overall when both tasks are added up. Sound is enough for intent; text is still useful for detail.

These numbers do not come from real calls. SLURP, the English reference, gathers about 72,000 recordings, or 58 hours, of single-turn requests to a home assistant, read from a tablet by more than a hundred participants. LeBenchmark 2.0 trained its models on up to 14,000 hours of French, of which 38 hours are telephone dialogues, and acted ones.

An intent in ten seconds is not an intent in three minutes

On our September 2026 calls, half of the calls answered by a human last ten seconds or less, and 8 out of 10 under twenty seconds. In ten seconds, a person says one or two sentences: “Hello?”, “Who's calling?”, “No thanks”, “Call me back later”. The intent has to be recognised from those few words, often at the first turn. In a three-minute dialogue, the person repeats, clarifies, corrects; in a ten-second call, there is no second chance.

The intents are not the same either. In MEDIA, the person called in order to book. In an outbound call, they asked for nothing, and what they want at the first turn may simply be to know who is calling. As for the intent that matters most to a company, “I'm interested, put me through to someone”, it comes late when it comes at all: on our calls, a handover to a human happens on 1.19% of answered calls, at 135 seconds (median). A model that learned on hotel bookings has seen none of these situations.

Where ASDIC stands

Nodical has, today, no understanding model of its own: the agent relies on a provider's speech recognition, hence on the cascade described above. We do not yet know how often it gets the intent wrong: that would take a corpus in which every turn is annotated by hand. That is what the programme's pilot set provides for: real pseudonymised calls, annotated with the intent at each turn, the answers to the questions asked, and the moment a handover would have been relevant. Our target is written as a target: more than 9 handovers out of 10 judged relevant. It is not reached; it will be published with the measurement, in the notebook.

What it changes for a handover

In an outbound campaign, the most expensive decision is passing the call to a salesperson: too early, they waste time on someone who was not ready; too late, the interested person has hung up. That decision rests entirely on intent. Nodical describes on its site the handover decision in an outbound campaign.

And in practice? the handover decision in an outbound campaign

Sources

accessed on 1 October 2026.

  1. arXiv 2403.19727, Alavoine et al., « New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark » (LREC-COLING 2024) — https://arxiv.org/abs/2403.19727
  2. Laperrière et al., « The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning » (LREC 2022) — https://aclanthology.org/2022.lrec-1.171/
  3. arXiv 2309.05472, « LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech » — https://arxiv.org/abs/2309.05472
  4. arXiv 2011.13205, Bastianelli et al., « SLURP: A Spoken Language Understanding Resource Package » (EMNLP 2020) — https://arxiv.org/abs/2011.13205
  5. ASDIC, « Un corpus d'appels téléphoniques en français : pourquoi il n'en existe pas » — https://asdic.ai/savoir/corpus-d-appels-telephoniques-en-francais-pourquoi-il-n-en-existe-pas
  6. ASDIC, carnet n° 1 : « Septembre 2026 : première mesure » — https://asdic.ai/carnet/2026-09-premiere-mesure

Corrections

No correction so far.