Training datasets.
Robotics datasets
Physical AI
UMI Dual-Camera Annotated Manipulation
500 hours · bimanual demonstrations
Bimanual UMI demonstrations with left and right gripper views, end-effector poses, gripper width, and synchronized episode annotations.
View detailsPhysical AI
Robot Teleoperation
Custom collection · specified robot
Robot demonstrations with synchronized video, depth, states, actions, and task outcomes. Includes the robot and sensor configuration.
View detailsPhysical AI
Tactile Glove and Stereo Manipulation
500 hours · visual and tactile signals
Bimanual demonstrations pairing stereo manipulation video with instrumented-glove signals, per-zone contacts, task phases, and frame states.
View detailsPhysical AI
Paired Egocentric RGB and IMU
20,000 hours · synchronized video and motion
Egocentric RGB video synchronized with dual IMU streams across real-world activities.
View detailsPhysical AI
Egocentric Residential Video
Available · 80 task categories
Available 1080p+ egocentric residential video across 80 non-cooking task categories, with hand-object interaction and state-change labels.
View detailsPhysical AI
Multimodal Human Manipulation
Available · multimodal streams
Human manipulation demonstrations with synchronized video, depth, motion capture, and IMU signals.
View detailsPhysical AI
Industrial First-Person Video
Available · 1080p+ video
Available 1080p+ industrial egocentric video with workflow, tool, outcome, and safety labels.
View details
More datasets
Agentic
Computer-Use Traces
Available · 250,000 traces
Available computer-use data pairing screen states with UI actions, outcomes, and optional DOM/accessibility metadata.
View detailsVoice
TTS Voice
Available · 30+ languages
Available 48 kHz conversational voice data across 30+ languages, with transcripts and diarization.
View detailsMultimodal
Audio-Video Conversational
Available · 20+ languages
Available synchronized conversational audio-video data across 20+ languages, including transcripts, gestures, and overlap labels.
View detailsVoice
ASR Voice
Available · transcript-aligned speech
Available multilingual conversational speech aligned with transcripts, diarization, language, accent, and audio-quality metadata.
View detailsWorld Models
Action-Conditioned Gameplay
Available · action-conditioned trajectories
Available action-conditioned gameplay data with synchronized video, inputs, state, event, and outcome logs.
View detailsAlignment
RLHF-Voice
Build-to-spec · preference pairs
Build-to-spec voice-preference data with human rankings and spoken rationales, scoped to the target model and evaluation criteria.
View details