Datasets / Voice

TTS Voice Dataset

Available · 30+ languages

Build expressive multilingual TTS and voice systems from natural 48 kHz conversation with transcripts, speaker metadata, pronunciation coverage, and audio quality signals.

Request access on Hugging Face

Dataset specification

  • 48 kHz WAV audio
  • Coverage: 30+ languages
  • Transcripts
  • Diarization metadata
  • Speaker attributes
  • Accent and locale labels
  • SNR, LUFS, clipping, and silence metrics
  • Quality scores
  • TTS suitability scores
  • CSV/JSON metadata

Intended uses

  • Text-to-speech and speech synthesis
  • Voice cloning and speaker modeling
  • Multilingual conversational AI and voice agents
  • Pronunciation, accent, and expressive-speech modeling

Scope notes

  • Language, accent, speaker, and recording distributions are confirmed for the selected delivery.
  • Transcript, diarization, normalization, and pronunciation fields vary by subset and are identified in the review package.
  • Repository metadata illustrates the schema; real audio is supplied in the buyer review package.

Collection scope

Speech coverage

  • 30+ global, regional, and underrepresented languages
  • Natural conversational dialogue
  • Accent, locale, and speaker-profile metadata where available

Language mix

  • English and major accents
  • Major commercial languages
  • High-growth regional languages
  • Low-resource and long-tail languages

Available scale

  • Available program
  • 48 kHz WAV or PCM where available

Buyer review package

  • Annotation schema, data dictionary, and sample metadata in the repository
  • Real audio, transcripts, metadata, and QA summaries for review
  • Target languages, accents, speaker mix, and hours confirmed for the selected delivery

Annotation & metadata fields

  • Language, locale, and accent
  • Speaker count and speaker attributes where available
  • Transcript and diarization availability
  • Topic category and noise level
  • Clean and expressive speech duration
  • Text-normalization and pronunciation metadata where available
  • SNR, LUFS, clipping, and silence metrics
  • Quality and TTS-suitability scores

Capture methodology

  • Natural conversational speech recorded for expressive pacing, tone, turn-taking, and pronunciation variation
  • 48 kHz audio in WAV or PCM delivery formats where available
  • 16-bit or 24-bit audio depending on the selected subset
  • Utterance metadata aligned with transcripts and diarization files where available

Provenance & rights chain

  • Contributor consent and chain-of-custody documentation included in licensing review
  • Commercial license issued directly by Datoric
  • Language, locale, and speaker-profile records travel with the selected delivery where available
  • Collection and annotation activity handled under Datoric's published privacy notice

Quality, duplicates & PII

How submissions are reviewed and cleaned before they are accepted into the dataset.

  • SNR, loudness, clipping, silence, and background-noise checks
  • Transcript and diarization quality evaluated by subset where available
  • Speaker consistency, pronunciation coverage, phoneme coverage, and expressive range review
  • Duplicate, malformed, and privacy-sensitive records handled under the agreed acceptance and redaction criteria

Formats & delivery

  • 48 kHz WAV audio
  • CSV metadata
  • JSON transcripts and diarization files where available
  • Repository review materials on Hugging Face; production transfer arranged directly

Rights & license scope

Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.

Version & verification

Availability
Available
Datasheet version
July 21, 2026
Release date
July 21, 2026
Last verified
July 21, 2026
Owner
Datoric

Frequently asked

Can we review audio before selecting a subset?
Yes. Request a buyer review package with real 48 kHz audio in your target languages, transcripts, metadata, QA summaries, and licensing documentation.
Which languages and accents are available?
The program covers 30+ languages across major commercial, high-growth regional, and long-tail groups. The exact language, accent, and speaker mix is confirmed for the selected delivery.
What quality signals are included?
Available metadata includes SNR, LUFS, clipping, silence, noise level, transcript quality, and TTS-suitability fields, with subset-specific coverage documented during review.
Can Datoric collect a different voice specification?
Yes. Custom programs can target languages, accents, speaker profiles, recording conditions, scripts, and acceptance criteria.

Related dataset specifications