Datasets / Voice

ASR Voice Dataset

Available · transcript-aligned speech

Train and evaluate multilingual speech-recognition systems on conversational audio aligned with transcripts, speaker turns, language and accent metadata, and audio-quality signals.

Dataset specification

  • 48kHz WAV audio
  • Coverage: 30+ languages
  • Transcripts
  • Diarization metadata
  • Language, accent, and locale labels
  • Speaker attributes
  • SNR, LUFS, clipping, and silence metrics
  • CSV/JSON metadata

Intended uses

  • Automatic speech recognition and speech-to-text
  • Multilingual and accent-aware transcription
  • Speaker diarization and conversational transcription
  • Voice-agent speech understanding

Scope notes

  • Language, accent, speaker, and recording-condition distributions are confirmed for the selected delivery.
  • Transcript, diarization, normalization, and timestamp fields vary by subset and are identified in the review package.
  • Repository metadata illustrates the schema; real audio is supplied in the buyer review package.

Collection scope

Speech coverage

  • 30+ global, regional, and underrepresented languages
  • Natural conversational dialogue
  • Accent, locale, and speaker-profile metadata where available

ASR annotations

  • Transcripts aligned to utterances
  • Speaker turns and diarization where available
  • Language, locale, accent, and audio-quality fields

Available program

  • 48kHz conversational WAV or PCM audio where available
  • Transcript and metadata files for selected subsets

Buyer review package

  • Annotation schema, data dictionary, and sample metadata in the repository
  • Real audio, transcripts, metadata, and QA summaries for review
  • Target languages, accents, speaker mix, and hours confirmed for the selected delivery

Annotation & metadata fields

  • Language, locale, and accent
  • Transcript and utterance timestamps
  • Speaker turns and diarization where available
  • Speaker count and attributes where available
  • Topic category and noise level
  • SNR, LUFS, clipping, and silence metrics
  • Transcript and audio-quality review fields

Capture methodology

  • Natural conversational speech recorded across supported languages, accents, and speaking styles
  • 48kHz audio in WAV or PCM delivery formats where available
  • Utterance metadata aligned with transcripts and diarization files where available
  • Recording conditions and audio-quality signals documented for the selected subset

Provenance & rights chain

  • Contributor consent and chain-of-custody documentation included in licensing review
  • Commercial license issued directly by Datoric
  • Language, locale, and speaker-profile records travel with the selected delivery where available
  • Collection and annotation activity handled under Datoric's published privacy notice

Quality, duplicates & PII

How submissions are reviewed and cleaned before they are accepted into the dataset.

  • Transcript alignment and language-label review
  • Diarization and speaker-turn checks where available
  • SNR, loudness, clipping, silence, and background-noise checks
  • Duplicate, malformed, and privacy-sensitive records handled under the agreed acceptance and redaction criteria

Formats & delivery

  • 48kHz WAV audio
  • CSV metadata
  • JSON transcripts and diarization files where available
  • Buyer review materials and production transfer arranged directly with Datoric

Rights & license scope

Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.

Version & verification

Availability
Available
Datasheet version
July 21, 2026
Release date
July 21, 2026
Last verified
July 21, 2026
Owner
Datoric

Frequently asked

What ASR fields are available?
Available fields include transcripts, utterance timestamps, language and locale labels, accent metadata, speaker turns, diarization, and audio-quality signals, with exact coverage confirmed for the selected subset.
Which languages and accents are available?
The source program covers 30+ languages across major commercial, regional, and long-tail groups. The exact language, accent, and speaker mix is confirmed for the selected delivery.
Can we review audio and transcripts before selecting a subset?
Yes. The buyer review package includes real audio, transcripts, metadata, QA summaries, and licensing documentation for the selected languages and recording conditions.
Can Datoric collect a different ASR specification?
Yes. Custom programs can target languages, accents, domains, speaker profiles, recording conditions, transcript conventions, and acceptance criteria.

Related dataset specifications