ASR Voice Dataset
Available · transcript-aligned speech
Train and evaluate multilingual speech-recognition systems on conversational audio aligned with transcripts, speaker turns, language and accent metadata, and audio-quality signals.
Dataset specification
- 48kHz WAV audio
- Coverage: 30+ languages
- Transcripts
- Diarization metadata
- Language, accent, and locale labels
- Speaker attributes
- SNR, LUFS, clipping, and silence metrics
- CSV/JSON metadata
Intended uses
- Automatic speech recognition and speech-to-text
- Multilingual and accent-aware transcription
- Speaker diarization and conversational transcription
- Voice-agent speech understanding
Scope notes
- Language, accent, speaker, and recording-condition distributions are confirmed for the selected delivery.
- Transcript, diarization, normalization, and timestamp fields vary by subset and are identified in the review package.
- Repository metadata illustrates the schema; real audio is supplied in the buyer review package.
Collection scope
Speech coverage
- 30+ global, regional, and underrepresented languages
- Natural conversational dialogue
- Accent, locale, and speaker-profile metadata where available
ASR annotations
- Transcripts aligned to utterances
- Speaker turns and diarization where available
- Language, locale, accent, and audio-quality fields
Available program
- 48kHz conversational WAV or PCM audio where available
- Transcript and metadata files for selected subsets
Buyer review package
- Annotation schema, data dictionary, and sample metadata in the repository
- Real audio, transcripts, metadata, and QA summaries for review
- Target languages, accents, speaker mix, and hours confirmed for the selected delivery
Annotation & metadata fields
- Language, locale, and accent
- Transcript and utterance timestamps
- Speaker turns and diarization where available
- Speaker count and attributes where available
- Topic category and noise level
- SNR, LUFS, clipping, and silence metrics
- Transcript and audio-quality review fields
Capture methodology
- Natural conversational speech recorded across supported languages, accents, and speaking styles
- 48kHz audio in WAV or PCM delivery formats where available
- Utterance metadata aligned with transcripts and diarization files where available
- Recording conditions and audio-quality signals documented for the selected subset
Provenance & rights chain
- Contributor consent and chain-of-custody documentation included in licensing review
- Commercial license issued directly by Datoric
- Language, locale, and speaker-profile records travel with the selected delivery where available
- Collection and annotation activity handled under Datoric's published privacy notice
Quality, duplicates & PII
How submissions are reviewed and cleaned before they are accepted into the dataset.
- Transcript alignment and language-label review
- Diarization and speaker-turn checks where available
- SNR, loudness, clipping, silence, and background-noise checks
- Duplicate, malformed, and privacy-sensitive records handled under the agreed acceptance and redaction criteria
Formats & delivery
- 48kHz WAV audio
- CSV metadata
- JSON transcripts and diarization files where available
- Buyer review materials and production transfer arranged directly with Datoric
Rights & license scope
Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.
Version & verification
- Availability
- Available
- Datasheet version
- July 21, 2026
- Release date
- July 21, 2026
- Last verified
- July 21, 2026
- Owner
- Datoric
Frequently asked
- What ASR fields are available?
- Available fields include transcripts, utterance timestamps, language and locale labels, accent metadata, speaker turns, diarization, and audio-quality signals, with exact coverage confirmed for the selected subset.
- Which languages and accents are available?
- The source program covers 30+ languages across major commercial, regional, and long-tail groups. The exact language, accent, and speaker mix is confirmed for the selected delivery.
- Can we review audio and transcripts before selecting a subset?
- Yes. The buyer review package includes real audio, transcripts, metadata, QA summaries, and licensing documentation for the selected languages and recording conditions.
- Can Datoric collect a different ASR specification?
- Yes. Custom programs can target languages, accents, domains, speaker profiles, recording conditions, transcript conventions, and acceptance criteria.