Guides / Physical AI

Physical-AI data capture and annotation

Multimodal physical data is only trainable if the streams line up and the labels match the task. This guide walks through the capture and annotation choices behind Datoric's physical-AI programs.

By Datoric · Published July 21, 2026 · Last reviewed July 21, 2026

Capture: many streams, one clock

Physical-AI data is inherently multimodal. A single manipulation recording can combine synchronized egocentric and external video, RGB-D/depth, motion capture, and IMU signal; a teleoperation episode combines multi-camera and wrist-camera video with the robot's own joint states, end-effector pose, and action commands. The engineering challenge is keeping all of it on one timeline.

The available manipulation specification includes calibration checks and sensor-alignment QA so pose, depth, and video share a common reference frame. The teleoperation record includes QA reports alongside trajectories. These records let buyers review how contact events, depth frames, action commands, and video are aligned.

Annotation: labeling for the task, not the pixels

Good physical-AI annotation is organized around what a policy needs to predict, not around what is easy to click. Datoric structures labels into perception, action, and outcome layers so a training pipeline can select exactly the supervision it needs.

Label layers

  • Perception: object, tool, location, and room identities; hand-object interaction and contact events.
  • Temporal structure: action segments, task phases, and before/after state changes.
  • Action: joint states, end-effector pose, gripper state, action commands, and trajectory logs for teleoperation.
  • Outcome: success, failure, and completion states for reward modeling and evaluation.

Quality metrics and review

Capture quality is measured, not assumed. Egocentric residential footage carries hand-visibility metrics; industrial footage carries blur and exposure metrics; manipulation carries calibration and sensor-alignment QA. Human QA review sits on top of the automated metrics across programs.

Consent and privacy by design

Because physical data records real people, homes, and workplaces, privacy handling is part of capture rather than an afterthought. Programs include privacy redaction, with contributor consent documentation, PII redaction, face blurring where needed, and identifier removal on the residential set. Documenting how footage was sourced and cleaned is what lets buyers deploy it with confidence.

A capture-and-annotation checklist

Before committing to a program

  • Confirm which sensors are synchronized and how alignment is verified.
  • Map required labels to the perception, temporal, action, and outcome layers above.
  • Agree the task taxonomy and what is explicitly out of scope.
  • Confirm capture-quality metrics and the QA review process.
  • Confirm consent sourcing and redaction for any footage of people or private spaces.

Evidence and procurement

Review the records behind the program

Related datasets

Keep reading

Frequently asked

Questions buyers ask

How is multimodal sensor data synchronized?
Manipulation programs run synchronized egocentric and external cameras with RGB-D, motion capture, and IMU, and include calibration checks and sensor-alignment QA so pose, depth, and video share a common reference frame.
What annotation layers do you provide?
Perception labels (objects, tools, locations, interactions, contacts), temporal structure (action segments, task phases, state changes), action signal (joint states, pose, gripper, action commands, trajectories), and outcome labels (success/failure/completion).
How is capture quality measured?
With program-specific metrics including hand visibility on residential egocentric footage, blur and exposure on industrial footage, calibration and sensor-alignment QA on manipulation, plus human QA review.

Design a capture-and-annotation spec

Share your sensors, labels, and task list, and we will design a capture-and-annotation program with the sync and QA your models need.

Browse datasets