by Nodical

ASDIC is a 2027-2028 research project. Published measurements are those of the notebook; targets are targets.

Learn

Spanish on the phone: accents, speed and “¿diga?”

Published By Douglas DemartProject status as of 1 October 2026

In one sentence

A voice agent calling in Spanish has to understand “¿Diga?”, “¿Bueno?” and “¿Aló?”, an accent from Madrid, the Canary Islands or Buenos Aires, and fast speech, while public leaderboards only evaluate read Spanish, never over the phone; and since Nodical's first Spanish calls have not taken place yet, ASDIC publishes no Spanish figures so far.

The first second already changes from one country to the next

In English you pick up with “Hello?”; in France, with “Allô ?”. In Spain it is “¿Diga?”, “¿Dígame?” (“say”, “tell me”) or a plain “¿Sí?”. The dictionary of the Royal Spanish Academy (RAE) lists “diga” and “dígame” as expressions used “when answering the phone”, marks “¿bueno?” as Mexican and “aló” as Latin American, with a nice detail: the word comes from the French allô. Different words with the same shape: one short word, a question, then silence. That is what tells a person from a voicemail, which talks at length without waiting for an answer.

For a machine, the problem is what the word means in each country. “Bueno” also means “all right”: a “¿Bueno?” at pick-up in Mexico is not a yes. And in Mexico, according to the Instituto Cervantes' Catálogo de voces hispánicas, “¿mande?” means “pardon?”: a request to repeat, not an answer.

One language, many accents

About 520 million people speak Spanish as their native language, according to the Instituto Cervantes (2025), and they do not pronounce it the same way. The Catálogo de voces hispánicas describes the differences city by city; some fall exactly where a machine notices them:

  • Madrid separates s from z (“casa”, house, and “caza”, hunt, sound different) and tends to aspirate syllable-final s.
  • Las Palmas de Gran Canaria pronounces them alike (seseo), aspirates or drops final s (“mesas”, tables, sounds almost like “mesah”), and its intonation has points in common with the Caribbean one.
  • Buenos Aires also has seseo, tends to aspirate final s, and uses “vos” for “you” (“vos tenés”).
  • Mexico City has seseo but keeps final s tense and long, and weakens unstressed vowels (“antes” sounds almost like “ants”).

Final s is so telling that the Linguistic Data Consortium used it to sort “Caribbean” and “non-Caribbean” speakers in its Spanish telephone corpora. And it often carries the plural: “las mesas” versus “la mesa”. Shrunk to a breath and sent through a narrow line (see the page on 8 kHz audio), we expect it to be even harder to hear; that is a hypothesis we have not measured yet.

Speed: more syllables, not more information

A study by the CNRS and Université Lumière Lyon 2, published in Science Advances in 2019, compared 17 languages, Spanish, French and English among them, using 15 short texts read by 10 native speakers per language. Spanish is cited as an example of a language spoken fast, with little information per syllable; in the end, all the languages studied convey about 39 bits of information per second. Speaking fast is not saying more.

For a machine, though, something does change: more syllables to split in the same time, in an already narrow sound, and an end-of-turn decision to make at a faster pace. But the study measures read texts: we know of no published measurement of Spanish speech rate over the phone.

What public leaderboards evaluate in Spanish

The most used public speech-recognition leaderboard, the Open ASR Leaderboard, evaluates Spanish in its multilingual track. The paper introducing it (October 2025) lists CoVoST-2 (read sentences), FLEURS (Wikipedia sentences read aloud, in 102 languages) and MLS (LibriVox audiobooks); its current dataset, updated in July 2026, holds FLEURS, Common Voice and MLS for Spanish. Read speech every time, in wideband.

On that ground, Spanish does well: in the paper's multilingual table, five models (Canary, Whisper, Phi-4, Parakeet and Voxtral) get between 3.2 and 3.8% of Spanish words wrong on average, and for each of them Spanish is, of the five languages evaluated, the one with the fewest errors. Good news for read Spanish. It says nothing about Spanish on the phone.

What they do not evaluate: the telephone

The leaderboard has no telephone track, in any language, and the private sets it added in May 2026 are English only. Yet, unlike French, Spanish does have research telephone audio, distributed by the LDC: CALLHOME Spanish (about 38 hours, 120 conversations at 8 kHz, mostly with family or friends), Fisher Spanish (about 163 hours, 819 conversations of 10 to 12 minutes among 136 Caribbean and non-Caribbean speakers) and two CALLFRIEND sets of 60 conversations each.

These are serious corpora, but none resembles a ten-second sales call: they are long conversations, recorded from the Americas (United States, Canada, Puerto Rico, Dominican Republic), released between 1996 and 2010, and no public leaderboard uses them. LiveKit's eot-bench counts Spanish among its 14 languages of real human-to-agent conversations, but says nothing about telephony. In French the situation is worse still, as we explain on the page about a French phone-call corpus.

Where ASDIC stands in Spanish

In Spanish, ASDIC has no figures at all: Nodical's first Spanish calls have not taken place yet. The measurements published in the notebook come from calls in French, and we do not carry them over. When the first Spanish campaigns start, we will measure the same things as in French (voicemail, end of turn, cut-offs, handovers, duration), by country and by accent, on 8 kHz sound as it arrives on the line. The programme's targets apply to French, Spanish and English; they are not reached in any language and will be published with the measurement that proves them.

What it changes for a campaign in Spain

An agent calling Spain should not sound as if it were calling Mexico, take a “¿Bueno?” for a yes, or lose the plural of someone who aspirates their s. Nodical explains on its site how an outbound prospecting campaign works, from qualification to warm transfer.

And in practice? how an outbound prospecting campaign works

Sources

accessed on 1 October 2026.

  1. RAE-ASALE, Diccionario de la lengua española, « aló » (« interj. Am. U. para responder al teléfono », du français allô) — https://dle.rae.es/al%C3%B3
  2. RAE-ASALE, Diccionario de la lengua española, « bueno » (« interj. Méx. U. para contestar al teléfono ») — https://dle.rae.es/bueno
  3. RAE-ASALE, Diccionario de la lengua española, « decir » (« diga, o dígame : exprs. U. cuando se responde al teléfono ») — https://dle.rae.es/decir
  4. Instituto Cervantes, Catálogo de voces hispánicas, Madrid — https://cvc.cervantes.es/lengua/voces_hispanicas/espana/madrid.htm
  5. Instituto Cervantes, Catálogo de voces hispánicas, Las Palmas de Gran Canaria — https://cvc.cervantes.es/lengua/voces_hispanicas/espana/las_palmas.htm
  6. Instituto Cervantes, Catálogo de voces hispánicas, Buenos Aires — https://cvc.cervantes.es/lengua/voces_hispanicas/argentina/buenosaires.htm
  7. Instituto Cervantes, Catálogo de voces hispánicas, Ciudad de México — https://cvc.cervantes.es/lengua/voces_hispanicas/mexico/mexicodf.htm
  8. Instituto Cervantes, « El español: lengua para el mundo 2025 », 20 claves (520 millions de locuteurs natifs) — https://cvc.cervantes.es/lengua/anuario/anuario_25/elm/p01.htm
  9. CNRS, « Similar information rates across languages, despite divergent speech rates » (communiqué, septembre 2019) — https://www.cnrs.fr/en/press/similar-information-rates-across-languages-despite-divergent-speech-rates
  10. Coupé, Oh, Dediu, Pellegrino, « Different languages, similar encoding efficiency », Science Advances, 4 septembre 2019 — https://doi.org/10.1126/sciadv.aaw2594
  11. arXiv 2510.06961, « Open ASR Leaderboard » (tableau 1 : jeux de données ; tableau 4 : erreurs par langue) — https://arxiv.org/html/2510.06961v1
  12. Open ASR Leaderboard, dépôt GitHub (jeux multilingues actuels : FLEURS, Common Voice, MLS) — https://github.com/huggingface/open_asr_leaderboard
  13. Hugging Face, « Open ASR Leaderboard: private data » (6 mai 2026, jeux privés en anglais seulement) — https://huggingface.co/blog/open-asr-leaderboard-private-data
  14. arXiv 2205.12446, « FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech » — https://arxiv.org/abs/2205.12446
  15. arXiv 2012.03411, « MLS: A Large-Scale Multilingual Dataset for Speech Research » — https://arxiv.org/abs/2012.03411
  16. LDC, CALLHOME Spanish Speech (LDC96S35) — https://catalog.ldc.upenn.edu/LDC96S35
  17. LDC, Fisher Spanish Speech (LDC2010S01) — https://catalog.ldc.upenn.edu/LDC2010S01
  18. LDC, CALLFRIEND Spanish-Caribbean Dialect (LDC96S57) — https://catalog.ldc.upenn.edu/LDC96S57
  19. LDC, CALLFRIEND Spanish-Non-Caribbean Dialect (LDC96S58) — https://catalog.ldc.upenn.edu/LDC96S58
  20. LiveKit, eot-bench (14 langues, dont l'espagnol) — https://github.com/livekit/eot-bench
  21. ASDIC, carnet n° 1 : « Septembre 2026 : première mesure » — https://asdic.ai/carnet/2026-09-premiere-mesure

Corrections

No correction so far.