Datasets / Multimodal

Audio-Video Conversational Dataset

Available · 20+ languages

Train multimodal conversational systems on synchronized speech and video that capture words, speaker turns, mouth motion, gestures, emotion, overlap, and interruption behavior.

Request access on Hugging Face

Dataset specification

  • 48 kHz audio
  • 1080p+ video
  • Multilingual coverage
  • Transcripts
  • Diarization
  • Timestamps
  • Face and mouth visibility labels
  • Gesture labels
  • Emotion labels
  • Laughter and interruption events
  • Overlap labels
  • Sync quality metrics
  • QA review

Intended uses

  • Multimodal voice agents and conversational AI
  • Avatar and lip-sync model training
  • Turn-taking, overlap, and interruption modeling
  • Emotion, gesture, and paralinguistic understanding

Scope notes

  • Language, speaker, environment, and interaction-style distributions are confirmed for the selected delivery.
  • 1080p video, 48 kHz audio, separated audio tracks, and annotation fields vary by subset and are identified during review.
  • Repository metadata illustrates the schema; real synchronized media is supplied in the buyer review package.

Collection scope

Conversation coverage

  • Natural multi-speaker conversation
  • 20+ languages
  • Turn-taking, overlap, laughter, interruption, gesture, gaze, and emotion

Language mix

  • English and major accents
  • Spanish and Portuguese
  • East Asian and South Asian languages
  • European, Southeast Asian, MENA, and long-tail languages

Available scale

  • Available program
  • 1080p+ video and 48 kHz audio where available

Buyer review package

  • Annotation schema, data dictionary, and sample metadata in the repository
  • Real synchronized audio-video, transcripts, annotations, and QA summaries
  • Target languages, interaction styles, speaker mix, and hours confirmed for the selected delivery

Annotation & metadata fields

  • Transcript segments and speaker turns
  • Timestamps and diarization
  • Face and mouth visibility
  • Gesture, gaze, emotion, and tone labels
  • Laughter and interruption events
  • Overlapping-speech labels
  • Speaker role
  • Audio-video sync quality and QA fields

Capture methodology

  • Natural conversations captured as synchronized audio and video
  • Face movement, mouth motion, gesture, gaze, expression, and body language preserved with speech
  • Speaker turns, pauses, laughter, interruption, and overlap aligned to timestamps
  • 1080p+ video and 48 kHz audio used where available for the selected subset

Provenance & rights chain

  • On-camera contributor consent and chain-of-custody documentation included in licensing review
  • Commercial license issued directly by Datoric
  • Speaker, language, and recording metadata travel with the selected delivery where available
  • Collection and annotation activity handled under Datoric's published privacy notice

Quality, duplicates & PII

How submissions are reviewed and cleaned before they are accepted into the dataset.

  • Audio-video synchronization and drift validation
  • Face and mouth visibility checks
  • Audio clarity, video exposure, and blur review
  • Transcript, diarization, paralinguistic annotation, duplicate, malformed-record, and privacy review

Formats & delivery

  • MP4 video
  • WAV audio tracks where separated
  • CSV metadata and transcript files
  • JSON annotations

Rights & license scope

Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.

Version & verification

Availability
Available
Datasheet version
July 21, 2026
Release date
July 21, 2026
Last verified
July 21, 2026
Owner
Datoric

Frequently asked

Can we inspect synchronization and visibility quality?
Yes. Review materials include audio-video sync, face and mouth visibility, audio clarity, video quality, annotation, and transcript QA fields for the selected subset.
Which conversational events are labeled?
Available fields cover speaker turns, pauses, laughter, interruptions, overlapping speech, gestures, gaze, emotion or tone, and speaker roles.
Can we select languages or interaction styles?
Yes. The pilot and delivery can be scoped by language, speaker mix, conversational setting, interaction pattern, annotation coverage, and hours.
Can Datoric collect a custom audio-video specification?
Yes. Custom programs can target languages, environments, camera and audio requirements, conversational scenarios, labels, and acceptance criteria.

Related dataset specifications