Multimodal Human Manipulation Dataset
Available · multimodal streams
Train physical AI systems on synchronized human manipulation demonstrations that combine first-person and external video with depth, motion, pose, and interaction signals.
Dataset specification
- Synchronized egocentric and external video
- RGB-D and depth
- Motion capture
- IMU signals
- Hand pose
- Body pose
- Object and tool labels
- Action segments
- Hand-object contact labels
- Task phases
- Completion states
- Calibration checks
- Sensor alignment QA
Intended uses
- Human-demonstration learning for Physical AI
- Manipulation, contact, and action recognition
- Motion, pose, and object-interaction modeling
- Embodied world-model training
Scope notes
- Sensor availability, camera coverage, task distribution, and environment mix are confirmed for the selected delivery.
- RGB-D, depth, pose, motion-capture, and IMU streams are subset-specific and identified during review.
- Repository metadata illustrates the schema; real synchronized sensor data is supplied in the buyer review package.
Collection scope
Modalities
- First-person and external video where available
- RGB-D and depth where available
- Motion capture, IMU, hand pose, and body pose
Task coverage
- Pick and place, sorting, folding, wiping, opening, closing, and tool use
- Package handling, stocking, light assembly, and object transfer
- Contact-rich human demonstrations
Available scale
- Available program
- Synchronized visual, motion, sensor, and annotation streams
Buyer review package
- Session schema, data dictionary, and sample metadata in the repository
- Real synchronized multi-sensor samples, logs, annotations, and QA summaries
- Target tasks, sensor mix, environments, and hours confirmed for the selected delivery
Annotation & metadata fields
- Task phase and action segment
- Object and tool labels
- Hand-object contact events
- Completion state and before or after state
- Sensor-stream descriptors and timestamps
- Hand and body pose where available
- Calibration, synchronization, and sensor-dropout QA fields
- Multi-view alignment metadata
Capture methodology
- Human demonstrations captured with synchronized first-person and external views where available
- RGB-D, depth, motion-capture, IMU, hand-pose, and body-pose streams aligned by session
- Task phases, action segments, contact events, objects, tools, and outcomes aligned to sensor time
- Calibration and stream descriptors recorded with the applicable sensor package
Provenance & rights chain
- Contributor consent and chain-of-custody documentation included in licensing review
- Commercial license issued directly by Datoric
- Sensor, task, and annotation lineage recorded for the selected delivery
- Collection and annotation activity handled under Datoric's published privacy notice
Quality, duplicates & PII
How submissions are reviewed and cleaned before they are accepted into the dataset.
- Calibration and timestamp-alignment checks
- Sensor-dropout flags and multi-view alignment review
- Video usability, object visibility, hand-pose coverage, and annotation quality checks
- Duplicate, malformed, incomplete, and privacy-sensitive sessions handled under the agreed acceptance criteria
Formats & delivery
- Synchronized video and sensor logs
- CSV metadata
- JSON annotations
- Pose and depth files where available
Rights & license scope
Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.
Version & verification
- Availability
- Available
- Datasheet version
- July 21, 2026
- Release date
- July 21, 2026
- Last verified
- July 21, 2026
- Owner
- Datoric
Frequently asked
- How is this different from a robot teleoperation program?
- This dataset captures human demonstrations with visual, motion, and sensor context. A build-to-spec robot teleoperation program would capture robot executions with robot states and action commands.
- Which sensor streams are available?
- Available modalities include first-person and external video, RGB-D or depth, motion capture, IMU, hand pose, and body pose. The exact stream mix is confirmed for the selected subset.
- Can we inspect synchronization quality?
- Yes. Review materials include calibration, timestamp alignment, sensor-dropout, visibility, pose-coverage, and multi-view alignment fields where applicable.
- Can the task and sensor specification be customized?
- Yes. Custom collection can target task categories, environments, camera placement, sensor stack, annotations, and acceptance criteria.