Capture: many streams, one clock
Physical-AI data is inherently multimodal. A single manipulation recording can combine synchronized egocentric and external video, RGB-D/depth, motion capture, and IMU signal; a teleoperation episode combines multi-camera and wrist-camera video with the robot's own joint states, end-effector pose, and action commands. The engineering challenge is keeping all of it on one timeline.
The available manipulation specification includes calibration checks and sensor-alignment QA so pose, depth, and video share a common reference frame. The teleoperation record includes QA reports alongside trajectories. These records let buyers review how contact events, depth frames, action commands, and video are aligned.