Datasets / Physical AI

Multimodal Human Manipulation Dataset

Available · multimodal streams

Train physical AI systems on synchronized human manipulation demonstrations that combine first-person and external video with depth, motion, pose, and interaction signals.

Request access on Hugging Face

Dataset specification

  • Synchronized egocentric and external video
  • RGB-D and depth
  • Motion capture
  • IMU signals
  • Hand pose
  • Body pose
  • Object and tool labels
  • Action segments
  • Hand-object contact labels
  • Task phases
  • Completion states
  • Calibration checks
  • Sensor alignment QA

Intended uses

  • Human-demonstration learning for Physical AI
  • Manipulation, contact, and action recognition
  • Motion, pose, and object-interaction modeling
  • Embodied world-model training

Scope notes

  • Sensor availability, camera coverage, task distribution, and environment mix are confirmed for the selected delivery.
  • RGB-D, depth, pose, motion-capture, and IMU streams are subset-specific and identified during review.
  • Repository metadata illustrates the schema; real synchronized sensor data is supplied in the buyer review package.

Collection scope

Modalities

  • First-person and external video where available
  • RGB-D and depth where available
  • Motion capture, IMU, hand pose, and body pose

Task coverage

  • Pick and place, sorting, folding, wiping, opening, closing, and tool use
  • Package handling, stocking, light assembly, and object transfer
  • Contact-rich human demonstrations

Available scale

  • Available program
  • Synchronized visual, motion, sensor, and annotation streams

Buyer review package

  • Session schema, data dictionary, and sample metadata in the repository
  • Real synchronized multi-sensor samples, logs, annotations, and QA summaries
  • Target tasks, sensor mix, environments, and hours confirmed for the selected delivery

Annotation & metadata fields

  • Task phase and action segment
  • Object and tool labels
  • Hand-object contact events
  • Completion state and before or after state
  • Sensor-stream descriptors and timestamps
  • Hand and body pose where available
  • Calibration, synchronization, and sensor-dropout QA fields
  • Multi-view alignment metadata

Capture methodology

  • Human demonstrations captured with synchronized first-person and external views where available
  • RGB-D, depth, motion-capture, IMU, hand-pose, and body-pose streams aligned by session
  • Task phases, action segments, contact events, objects, tools, and outcomes aligned to sensor time
  • Calibration and stream descriptors recorded with the applicable sensor package

Provenance & rights chain

  • Contributor consent and chain-of-custody documentation included in licensing review
  • Commercial license issued directly by Datoric
  • Sensor, task, and annotation lineage recorded for the selected delivery
  • Collection and annotation activity handled under Datoric's published privacy notice

Quality, duplicates & PII

How submissions are reviewed and cleaned before they are accepted into the dataset.

  • Calibration and timestamp-alignment checks
  • Sensor-dropout flags and multi-view alignment review
  • Video usability, object visibility, hand-pose coverage, and annotation quality checks
  • Duplicate, malformed, incomplete, and privacy-sensitive sessions handled under the agreed acceptance criteria

Formats & delivery

  • Synchronized video and sensor logs
  • CSV metadata
  • JSON annotations
  • Pose and depth files where available

Rights & license scope

Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.

Version & verification

Availability
Available
Datasheet version
July 21, 2026
Release date
July 21, 2026
Last verified
July 21, 2026
Owner
Datoric

Frequently asked

How is this different from a robot teleoperation program?
This dataset captures human demonstrations with visual, motion, and sensor context. A build-to-spec robot teleoperation program would capture robot executions with robot states and action commands.
Which sensor streams are available?
Available modalities include first-person and external video, RGB-D or depth, motion capture, IMU, hand pose, and body pose. The exact stream mix is confirmed for the selected subset.
Can we inspect synchronization quality?
Yes. Review materials include calibration, timestamp alignment, sensor-dropout, visibility, pose-coverage, and multi-view alignment fields where applicable.
Can the task and sensor specification be customized?
Yes. Custom collection can target task categories, environments, camera placement, sensor stack, annotations, and acceptance criteria.

Related dataset specifications