Guides / Voice AI

Multilingual and code-switching speech data

Average accuracy hides where multilingual voice models actually break. This guide explains the language-equity gap and code-switching failures our research found, and the data that addresses them.

By Datoric · Published July 21, 2026 · Last reviewed July 21, 2026

The language-equity gap is real and measurable

GlobalVoice-Bench reports results per language-resource tier rather than only as an overall average. It defines the language-equity gap as the performance differential between high- and low-resource languages under identical evaluation across twenty languages in three resource tiers.

The report finds a large, monotonic gap, with several evaluated systems unable to produce usable transcripts on the low-resource tier. Per-language results reveal deployment gaps that an overall average can conceal.

Code-switching breaks quietly

Bilingual speakers mix languages mid-sentence, and standard WER summaries do not capture every resulting failure mode. On the Mandarin-English code-switching axis of GlobalVoice-Bench, some evaluated systems returned empty transcripts on a large fraction of samples, while switch-point language identification remained low across the evaluated models.

The lesson for data is direct: code-switched speech has to be collected and represented explicitly, because it is a distinct phenomenon that clean monolingual corpora do not teach and averaged metrics do not surface.

Building coverage that holds up

Datoric's conversational voice program spans 30+ languages in 48kHz WAV with accent and locale labels, transcripts, and diarization. That structure lets a buyer target specific languages and balance accents. Code-switching is scoped as a dedicated custom collection with target language pairs, switch patterns, and transcription requirements defined explicitly.

What to build in

  • Explicit per-language coverage, including lower-resource languages, rather than an averaged mix.
  • Accent and locale labels so within-language variation is represented and controllable.
  • Code-switched speech collected as its own category, with transcripts that capture the switches.
  • Diarization and speaker attributes for multi-speaker, multilingual scenarios.

How to evaluate multilingual coverage

Move past the headline number

  • Ask for per-language coverage and quality, not a single averaged figure.
  • Check whether code-switched and overlapping speech is represented at all.
  • Confirm accent and locale labels exist for the languages you care about.
  • Treat silent empty transcripts as a first-class failure mode to test for.

Evidence and procurement

Review the records behind the program

Related datasets

Keep reading

Frequently asked

Questions buyers ask

What is the language-equity gap?
It is the performance differential between high- and low-resource languages under identical evaluation. GlobalVoice-Bench found it to be large and monotonic across frontier speech systems, with several unable to serve the low-resource tier reliably.
Why is code-switching so hard for speech models?
Speakers switch languages mid-utterance, which monolingual training does not teach. GlobalVoice-Bench documented systems silently returning empty transcripts on code-switched speech and low switch-point language-ID accuracy across models.
How does Datoric support multilingual and code-switching needs?
The conversational program covers 30+ languages with accent and locale labels, transcripts, and diarization. Code-switching needs are scoped as a dedicated custom collection with explicit language-pair and transcription requirements.

Build multilingual coverage that holds up

Tell us your target languages, accents, and code-switching needs, and we will scope a multilingual speech program with per-language coverage.

Browse datasets