ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.
LearnA French phone-call corpus: why none exists
Everything that exists in French is either old, or Canadian, or not telephone audio, or not annotated for the decisions of a call; ASDIC is building its own from real pseudonymised calls, and will publish an evaluation set.
What an agent that makes calls needs
To learn to decide during a call, a model needs calls: telephone-line sound (8 kilohertz, compressed), real and short conversations, in French, with two voices — and above all annotations that say, for each call, what had to be decided and when: human or voicemail, at what moment it could be known, where the turns end, what the person wanted. That fourfold criterion — telephone, real, French, annotated for decisions — is what no public corpus meets today.
The inventory, without indulgence
CALLFRIEND French (1996)
The only genuine French phone-call corpus distributed by the Linguistic Data Consortium: 60 unscripted conversations of 5 to 30 minutes, at 8 kilohertz, between native Canadian speakers, placed from North America. It is thirty years old, Canadian, and was designed for language identification, not for understanding what is said.
ESLO (Orléans)
The Orléans sociolinguistic surveys are a large corpus of spoken French, first collected in the late 1960s and again since 2008; it contains a few phone calls. The telephone subset republished on Hugging Face in 2026 holds 36 excerpts and under two hours, under a non-commercial research licence. Precious for linguistics; too small, and off-limits for commercial use, to train anything.
Claire (2023)
The Claire French Dialogue Dataset, published by LINAGORA, gathers about 160 million words of French dialogue from 24 corpora: transcripts and stage plays. It is text, not sound; it teaches a language model how people talk to each other, not how they hear each other over the phone. Non-commercial licence.
MEDIA (2005)
Distributed by ELRA since 2005 and long the French reference for spoken language understanding, MEDIA consists of tourist-booking dialogues recorded in a laboratory in the early 2000s, annotated in concepts (and, since 2024, in intents). Free for research since 2020. Twenty years old, a single domain, simulated dialogues: useful to compare models, not to learn a 2026 outbound call.
Luqia's benchmark (2026)
In September 2026, Luqia Technologies published the first French “full-duplex” dialogue benchmark (pauses, turn-taking, interruptions, backchannels), in two variants: Canadian, built on CALLFRIEND, and European, built on MEDIA. Serious work, worth citing — and which, by construction, rests on the two corpora above, with their limits.
And the public leaderboards
The most used speech-recognition leaderboard, the Open ASR Leaderboard, evaluates French on CoVoST-2 and FLEURS: read speech, in wideband. It has no telephone track, in any language.
Why nobody has done it
Because a corpus of real calls is hard to build legally and expensive to annotate. You need the right to record, you must tell people, remove what identifies them, and have thousands of calls listened to by annotators. Laboratories do not have the calls; the companies that have them do not publish. ASDIC sits at the meeting point: Nodical places the calls, and the research programme's mission is to turn them into publishable data.
What we are building
A corpus of real outbound calls, in French first, then Spanish and English, at 8 kilohertz as the sound arrives on the line, pseudonymised before any annotation: names, numbers and addresses removed from audio and transcripts. The annotations target the decisions of the call: human or voicemail with the moment it becomes recognisable, intent at each turn, answers to the questions asked, turn boundaries and endings, the moment a handover would have been relevant. A pilot set of about ten hours in 2027, then 300 to 500 hours over the programme.
What will be published, and within which framework
The training corpus will not be published: it contains real conversations, and pseudonymisation is not enough to make it public data. What we will publish: an anonymised, reduced evaluation set, so that anyone can replay their models on French telephone audio and compare; the results of open models on that set; and the method, in the notebook. The framework is that of the CNIL's recommendations for the development of AI systems: reuse based on legitimate interest, people informed at the start of the call, right to object, impact assessment reviewed by legal counsel before the first training. Until that framework is validated, nothing is published.
Who is building ASDIC
ASDIC is the research programme of Nodical, which has voice agents make calls every day on behalf of companies. Nodical introduces itself on its site: who is building ASDIC.
And in practice? who is building ASDIC
Sources
accessed on 1 October 2026.
- LDC, CALLFRIEND Canadian French (LDC96S48, 1996) — https://catalog.ldc.upenn.edu/LDC96S48
- Hugging Face, BrunoHays/eslo-telephony-fr (sous-ensemble téléphonique d'ESLO 1) — https://huggingface.co/datasets/BrunoHays/eslo-telephony-fr
- arXiv 2311.16840, « The Claire French Dialogue Dataset » (LINAGORA Labs, 2023) — https://arxiv.org/abs/2311.16840
- arXiv 2403.19727, MEDIA : nouvelle annotation en intentions (ELRA, distribué depuis 2005) — https://arxiv.org/abs/2403.19727
- arXiv 2609.10765, Luqia, « French Full-Duplex Benchmark » (EMNLP 2026 Industry) — https://arxiv.org/abs/2609.10765
- arXiv 2510.06961, « Open ASR Leaderboard » (pistes et jeux de données) — https://arxiv.org/html/2510.06961v1
- CNIL, recommandations pour le développement des systèmes d'IA (corpus de fiches, 2025) — https://www.cnil.fr/fr/intelligence-artificielle/recommandations-developpement-systemes-ia
Corrections
No correction so far.