Audio-Video Conversational Dataset
Available · 20+ languages
Train multimodal conversational systems on synchronized speech and video that capture words, speaker turns, mouth motion, gestures, emotion, overlap, and interruption behavior.
Dataset specification
- 48 kHz audio
- 1080p+ video
- Multilingual coverage
- Transcripts
- Diarization
- Timestamps
- Face and mouth visibility labels
- Gesture labels
- Emotion labels
- Laughter and interruption events
- Overlap labels
- Sync quality metrics
- QA review
Intended uses
- Multimodal voice agents and conversational AI
- Avatar and lip-sync model training
- Turn-taking, overlap, and interruption modeling
- Emotion, gesture, and paralinguistic understanding
Scope notes
- Language, speaker, environment, and interaction-style distributions are confirmed for the selected delivery.
- 1080p video, 48 kHz audio, separated audio tracks, and annotation fields vary by subset and are identified during review.
- Repository metadata illustrates the schema; real synchronized media is supplied in the buyer review package.
Collection scope
Conversation coverage
- Natural multi-speaker conversation
- 20+ languages
- Turn-taking, overlap, laughter, interruption, gesture, gaze, and emotion
Language mix
- English and major accents
- Spanish and Portuguese
- East Asian and South Asian languages
- European, Southeast Asian, MENA, and long-tail languages
Available scale
- Available program
- 1080p+ video and 48 kHz audio where available
Buyer review package
- Annotation schema, data dictionary, and sample metadata in the repository
- Real synchronized audio-video, transcripts, annotations, and QA summaries
- Target languages, interaction styles, speaker mix, and hours confirmed for the selected delivery
Annotation & metadata fields
- Transcript segments and speaker turns
- Timestamps and diarization
- Face and mouth visibility
- Gesture, gaze, emotion, and tone labels
- Laughter and interruption events
- Overlapping-speech labels
- Speaker role
- Audio-video sync quality and QA fields
Capture methodology
- Natural conversations captured as synchronized audio and video
- Face movement, mouth motion, gesture, gaze, expression, and body language preserved with speech
- Speaker turns, pauses, laughter, interruption, and overlap aligned to timestamps
- 1080p+ video and 48 kHz audio used where available for the selected subset
Provenance & rights chain
- On-camera contributor consent and chain-of-custody documentation included in licensing review
- Commercial license issued directly by Datoric
- Speaker, language, and recording metadata travel with the selected delivery where available
- Collection and annotation activity handled under Datoric's published privacy notice
Quality, duplicates & PII
How submissions are reviewed and cleaned before they are accepted into the dataset.
- Audio-video synchronization and drift validation
- Face and mouth visibility checks
- Audio clarity, video exposure, and blur review
- Transcript, diarization, paralinguistic annotation, duplicate, malformed-record, and privacy review
Formats & delivery
- MP4 video
- WAV audio tracks where separated
- CSV metadata and transcript files
- JSON annotations
Rights & license scope
Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.
Version & verification
- Availability
- Available
- Datasheet version
- July 21, 2026
- Release date
- July 21, 2026
- Last verified
- July 21, 2026
- Owner
- Datoric
Frequently asked
- Can we inspect synchronization and visibility quality?
- Yes. Review materials include audio-video sync, face and mouth visibility, audio clarity, video quality, annotation, and transcript QA fields for the selected subset.
- Which conversational events are labeled?
- Available fields cover speaker turns, pauses, laughter, interruptions, overlapping speech, gestures, gaze, emotion or tone, and speaker roles.
- Can we select languages or interaction styles?
- Yes. The pilot and delivery can be scoped by language, speaker mix, conversational setting, interaction pattern, annotation coverage, and hours.
- Can Datoric collect a custom audio-video specification?
- Yes. Custom programs can target languages, environments, camera and audio requirements, conversational scenarios, labels, and acceptance criteria.