Multilingual speech and synchronized audio-video data for text-to-speech, speech recognition, and voice agents. Capture transcripts, diarization, speaker attributes, recording quality, timing, and interaction labels.
Conversational voice capacity
Available · capacity up to 20,000 hours
Conversational voice languages
30+
Audio quality
48kHz WAV
Audio-video capacity
Available · capacity up to 4,000 hours
01
Speech data decides voice-model behavior
Voice models reflect their training distribution. Language, accent, overlap, interruption, and expressive range are data choices before they are modeling choices.
Explore the details+
Datoric's published research makes the stakes concrete. GlobalVoice-Bench documents a large, monotonic language-equity gap across the evaluated speech systems and a silent-drop failure mode on code-switched speech. VoicePro-Bench reports that audio understanding, rather than text reasoning, is the dominant source of downstream error in its matched evaluation. Both reports are linked below.
02
What we collect
Choose conversational voice, synchronized audio-video, or spoken-preference programs, with languages, formats, and annotations scoped to the model workflow.
Available and build-to-spec programs+
Conversational voice: 48kHz WAV across 30+ languages with transcripts, diarization, speaker attributes, accent and locale labels, and per-clip SNR, LUFS, clipping, and silence metrics plus TTS-suitability scores.
Audio-video conversation: 48kHz audio with 1080p+ video, transcripts, diarization, timestamps, face and mouth visibility, gesture and emotion labels, and laughter, interruption, and overlap events.
Voice-preference data: build-to-spec spoken-preference pairs with spoken rationales for model alignment.
03
Multilingual and code-switched coverage
Multilingual quality depends on per-language coverage, not one average. Datoric scopes languages, accents, locales, and code-switching patterns to the target deployment.
Explore the details+
GlobalVoice-Bench found that two evaluated systems, ElevenLabs Scribe and GPT-4o-mini Audio, achieved 100% effective coverage across all twenty languages on its cultural-QA axis, while several systems returned empty transcripts on Mandarin-English code-switched speech. Datoric handles code-switching as a dedicated custom collection with target pairs and switch patterns defined during review.
04
Turn-taking, overlap, and expressive speech
Voice agents need turn-taking, overlap, interruption, laughter, and expressive variation. Programs align those events with diarization, timestamps, and emotion labels.
05
Provenance, recording conditions, and speaker distribution
Each delivery pairs audio-quality metrics and speaker-distribution labels with consent and sourcing records. Audio-video programs add synchronization, visibility, transcript, and diarization QA.
Explore the details+
Conversational voice deliveries can include per-clip SNR, LUFS, clipping, and silence metrics plus speaker attributes and accent or locale labels. Audio-video deliveries add synchronization, visibility, transcript, diarization, and paralinguistic QA. Buyers can review these records against the target environment and participant population.
06
How to evaluate a voice dataset
Evaluate coverage by language, accent, speaker, and interaction type, then verify recording quality, sourcing, consent, and review evidence.
Questions worth asking any provider+
Which languages and accents are covered, and is coverage per-language rather than an averaged figure?
Is code-switched and overlapping speech represented, or only clean single-speaker audio?
What recording-quality metrics are reported per clip (SNR, LUFS, clipping, silence)?
What speaker attributes and locale labels support distribution control?
How is the audio sourced and consented, and how is it QA-reviewed?
Program standards
Review how the data is sourced, checked, and licensed.
What languages does the conversational voice program cover?+
It spans 30+ languages in 48kHz WAV, with accent and locale labels, transcripts, diarization, and speaker attributes so multilingual behavior can be trained and measured per language.
Do you provide data for expressive and conversational TTS?+
Yes. The audio-video conversational specification includes laughter, interruption, and overlap events plus emotion labels, while the conversational voice specification includes TTS-suitability scores.
How do you support code-switching and multilingual robustness?+
The conversational program supports multilingual distribution control with 30+ language coverage and accent and locale labels. Datoric scopes code-switching as a dedicated custom collection informed by the failure modes documented in GlobalVoice-Bench.
What quality metrics come with the audio?+
Conversational voice can include per-clip SNR, LUFS, clipping, silence, quality, and TTS-suitability scores. Audio-video review materials include synchronization, visibility, audio clarity, video quality, transcript, and diarization QA fields.
Build a voice dataset to spec
Tell us the languages, speakers, and conditions your voice model needs, and we will scope a conversational or audio-video program around them.