What egocentric video is
Egocentric, or first-person, video is recorded from the camera position of the person doing a task, roughly matching the viewpoint an embodied agent would have. It captures hands, tools, and manipulated objects from the actor's perspective, providing a useful frame of reference for manipulation and navigation policies.
That viewpoint is why egocentric video anchors so much physical-AI work: it supplies broad coverage of real activity without instrumenting a robot for every scene, and it can be labeled with the actions and state changes a downstream policy needs to imitate.
The labels that make it trainable
Raw first-person footage is a starting point, not a dataset. What turns it into training data is the annotation layer: what task is happening, which objects and tools are involved, how the hands interact with them, and how the scene changes as a result.
Labels Datoric attaches to egocentric footage
- Task, action, object, and room labels for the residential corpus.
- Hand-object interaction labels and before/after state changes.
- Workflow phases, pick/place and scan events, and tool and location labels on the industrial corpus.
- Completion outcomes and safety-relevant labels where the task warrants them.
Coverage: residential and industrial
Datoric offers two available egocentric video programs with different coverage. The residential set is 1080p+ first-person video across 80 non-cooking household task categories, with hand-visibility metrics, PII redaction, face blurring where needed, and identifier removal. The industrial set is 1080p+ first-person workplace video for warehouse and industrial tasks, with blur and exposure metrics and privacy redaction.
Choosing between them is a coverage decision: residential activity for home and general manipulation tasks, industrial activity for logistics, warehouse, and manufacturing workflows.
Quality and consent
Purpose-built egocentric collection makes quality and privacy controls part of the specification. Datoric's records include human QA review, capture-quality metrics, privacy redaction, and, for the residential set, contributor consent documentation, PII redaction, face blurring where needed, and identifier removal.
How to evaluate an egocentric corpus
What to check before you buy
- Resolution and viewpoint stability: is the footage genuinely first-person and high enough resolution (1080p+) for the target task?
- Label density: are hands, objects, interactions, and state changes annotated, or only coarse activity tags?
- Task taxonomy and explicit scope: for example, the residential set covers 80 non-cooking categories.
- Consent and redaction: how is the footage sourced, and how is PII removed?