How it works
From a spec to data
you can defend.
A single, repeatable path takes your requirement from specification to delivered, licensed data — with consent established up front and quality reviewed before anything ships.
The workflow
From requirement to delivered data.
Every engagement follows the same disciplined path, whether you start from a catalog dataset or a blank specification.
- 01
Define the specification
We work with your team to write down exactly what the data must contain — modalities, tasks, labels, environments, and the acceptance criteria that define success.
- 02
Source consented contributors
We recruit contributors who fit the specification and collect explicit consent and licensing at the point of contribution, so rights are established before any data is captured.
- 03
Collect at the origin
Contributors record voice, video, computer-use, or physical-AI data to the spec. Structured metadata — timestamps, actions, and labels — is captured alongside the raw signal.
- 04
Review and QA
Submissions are reviewed against modality-specific signal, metadata, and labeling criteria, with redaction applied where the specification requires it.
- 05
Package and pilot
We assemble a review package so your team can assess fit, then run a pilot against your evaluation before scaling the program.
- 06
Deliver with rights attached
Accepted data is delivered with the provenance and license records defined for the engagement, through the agreed access channel.
See it for yourself
Each step maps to something you can inspect.
The workflow is not a diagram — it is reflected in the datasheets, methodology, and notices published on this site.
| What we state | Primary evidence |
|---|---|
| Collection targets and available fields are published per dataset. | Dataset catalogChecked July 21, 2026 |
| The egocentric-video specification documents contributor consent and redaction requirements. | Egocentric video repositoryChecked July 21, 2026 |
| The computer-use repository documents human review, PII review, and duplicate and malformed-trace filtering. | Computer-Use Traces repositoryChecked July 21, 2026 |
| Official catalog records publish the current access-review and buyer-package turnaround. | Official Datoric dataset repositoriesChecked July 21, 2026 |
Questions
Working with us
- Can we start from an existing dataset?
- Yes. Every catalog entry publishes a full buyer datasheet and links to its repository review materials. If an existing dataset fits, evaluation can begin from those records.
- What if the data we need does not exist yet?
- That is the common case. We run a build-to-spec collection program against your requirements, with QA and acceptance criteria agreed before collection begins.
- How do we evaluate quality before buying?
- Start with the published datasheet. Each repository provides its schema and sample metadata, with representative media, QA summaries, and licensing documentation available in a buyer review package.
- How is delivered data accessed?
- Catalog pages link to Datoric's Hugging Face access endpoints. Production delivery format and transfer method are confirmed for the engagement.
Continue
More of the record
Data rights & provenance
How consent, licensing, and origin are defined and reviewed.
Quality methodology
Modality-specific checks for datasets and benchmarks.
Security & privacy
Request-based access, privacy handling, and program safeguards.
Licensing
Commercial license scope and the terms buyers review.
Custom data collection
Build-to-spec programs for data that does not exist yet.
About Datoric
Company identity, research group, and contacts.