Egocentric video, human manipulation, and robot teleoperation for systems that learn to act in the physical world. Available programs pair structured labels with calibration, synchronization, and QA.
Residential egocentric capacity
Available · capacity up to 100,000 hours
Industrial first-person capacity
Available · capacity up to 50,000 hours
Manipulation capacity
Available · capacity up to 6,000 hours
Teleoperation capacity
Available · capacity up to 2,000 episodes
01
Physical AI runs on demonstrations, not scrapes
Policies, controllers, and world models learn from demonstrations that pair actions with observations and outcomes. Synchronized sensors and annotations make each trajectory reviewable.
Explore the details+
Datoric collects egocentric activity, human manipulation, and robot teleoperation with task specifications and labels for actions, objects, tools, interactions, state changes, and outcomes.
02
What we capture
Egocentric video captures real-world tasks. Manipulation data adds depth, pose, and IMU signals. Teleoperation pairs robot observations with controls for imitation learning and VLA.
Available programs+
Residential egocentric video: 1080p+ recordings across 80 non-cooking household tasks, labeled by task, action, object, room, and before/after state.
Industrial first-person video: 1080p+ workplace recordings with workflow, pick/place, scan, tool, location, and safety labels.
Human manipulation: synchronized egocentric and external video with RGB-D, motion capture, IMU, hand/body pose, and contact labels.
Robot teleoperation: multi-camera and wrist video with joint states, end-effector pose, gripper state, commands, and trajectory logs.
03
Signals that feed VLA and world models
VLA policies need synchronized perception and action. World models need coherent sequences with labeled state changes. Datoric preserves both pairings through training delivery.
Structured signal types+
Action and trajectory logs: action commands, end-effector pose, gripper state, and joint states timestamped against video.
Perception labels: object, tool, and location identities plus hand-object interaction and contact events.
State transitions: before/after state changes and task phases for temporal and world-model training.
Outcome labels: success, failure, and completion states for reward modeling and policy evaluation.
04
Capture rigs, calibration, and QA
Programs synchronize egocentric and external cameras, RGB-D, motion capture, and IMU. Calibration aligns pose, depth, and video; teleoperation includes trajectory QA.
Explore the details+
QA covers hand visibility, blur, exposure, human review, and privacy redaction. Residential records also include consent documentation, face blurring, PII redaction, and identifier removal.
05
How to evaluate a physical-AI dataset
Look beyond hour counts. Evaluate stream alignment, label density, task coverage, and outcome coverage to judge what a physical-AI dataset can actually teach a policy.
Questions worth asking any provider+
Are action and perception streams synchronized, and how is that verified (calibration and sensor-alignment checks)?
Which labels are present per frame or per episode, including objects, contacts, phases, and outcomes, versus inferred later?
What task taxonomy is covered, and what is explicitly out of scope (for example, the residential set covers 80 non-cooking task categories)?
How are privacy and consent handled for footage of real people and homes?
Which quantities describe available inventory and which describe collection capacity for the selected delivery?
06
Choose the program that fits
Choose by task, sensors, labels, QA, provenance, and delivery. The catalog marks availability and capacity, including 80 residential task categories and episode-based teleoperation, before a pilot is scoped.
Program standards
Review how the data is sourced, checked, and licensed.
It is demonstration data that teaches models to perceive and act in the physical world: egocentric video of real tasks, multimodal human-manipulation recordings, and robot-teleoperation episodes, annotated with the objects, actions, state changes, and outcomes a policy needs to learn.
Do you support vision-language-action (VLA) and imitation-learning workflows?+
Yes. Teleoperation episodes pair multi-camera and wrist-camera video with action commands, end-effector pose, gripper state, and trajectory logs. This provides aligned perception and action records for VLA and imitation-learning workflows.
How is multimodal data kept in sync?+
Manipulation programs run synchronized egocentric and external cameras with RGB-D, motion capture, and IMU, and include calibration checks and sensor-alignment QA so video, depth, and pose share a common reference frame.
How do you handle privacy for footage of real people and homes?+
Video specifications include human QA review and privacy redaction. The residential record also includes contributor consent documentation, PII redaction, face blurring where needed, and identifier removal.
Build a physical-AI dataset to spec
Tell us the tasks, sensors, and outcomes your policy needs to learn, and we will scope an egocentric, manipulation, or teleoperation program around them.