Guides / Physical AI

Egocentric video for training data

First-person video is the closest widely-capturable proxy for what an embodied agent sees. This guide explains what egocentric training data is, what labels make it useful, and how to evaluate a corpus for physical-AI work.

By Datoric · Published July 21, 2026 · Last reviewed July 21, 2026

What egocentric video is

Egocentric, or first-person, video is recorded from the camera position of the person doing a task, roughly matching the viewpoint an embodied agent would have. It captures hands, tools, and manipulated objects from the actor's perspective, providing a useful frame of reference for manipulation and navigation policies.

That viewpoint is why egocentric video anchors so much physical-AI work: it supplies broad coverage of real activity without instrumenting a robot for every scene, and it can be labeled with the actions and state changes a downstream policy needs to imitate.

The labels that make it trainable

Raw first-person footage is a starting point, not a dataset. What turns it into training data is the annotation layer: what task is happening, which objects and tools are involved, how the hands interact with them, and how the scene changes as a result.

Labels Datoric attaches to egocentric footage

  • Task, action, object, and room labels for the residential corpus.
  • Hand-object interaction labels and before/after state changes.
  • Workflow phases, pick/place and scan events, and tool and location labels on the industrial corpus.
  • Completion outcomes and safety-relevant labels where the task warrants them.

Coverage: residential and industrial

Datoric offers two available egocentric video programs with different coverage. The residential set is 1080p+ first-person video across 80 non-cooking household task categories, with hand-visibility metrics, PII redaction, face blurring where needed, and identifier removal. The industrial set is 1080p+ first-person workplace video for warehouse and industrial tasks, with blur and exposure metrics and privacy redaction.

Choosing between them is a coverage decision: residential activity for home and general manipulation tasks, industrial activity for logistics, warehouse, and manufacturing workflows.

Quality and consent

Purpose-built egocentric collection makes quality and privacy controls part of the specification. Datoric's records include human QA review, capture-quality metrics, privacy redaction, and, for the residential set, contributor consent documentation, PII redaction, face blurring where needed, and identifier removal.

How to evaluate an egocentric corpus

What to check before you buy

  • Resolution and viewpoint stability: is the footage genuinely first-person and high enough resolution (1080p+) for the target task?
  • Label density: are hands, objects, interactions, and state changes annotated, or only coarse activity tags?
  • Task taxonomy and explicit scope: for example, the residential set covers 80 non-cooking categories.
  • Consent and redaction: how is the footage sourced, and how is PII removed?

Evidence and procurement

Review the records behind the program

Related datasets

Keep reading

Frequently asked

Questions buyers ask

Why use egocentric video instead of third-person video?
First-person footage matches the viewpoint an embodied agent has at inference time and captures hands, tools, and objects the way a manipulation policy must reason about them, which third-person angles often occlude.
What tasks does the residential egocentric set cover?
It spans 80 non-cooking household task categories in 1080p+ first-person video, labeled with task, action, object, and room identities plus hand-object interaction and before/after state changes.
What privacy controls are documented?
The specifications include human QA review and privacy redaction. The residential record also includes contributor consent documentation, PII redaction, face blurring where needed, and identifier removal.

Scope an egocentric video program

Bring your task taxonomy and label requirements; we will scope a residential or industrial first-person video program to match.

Browse datasets