Datasets / Agentic

Computer-Use Traces Dataset

Available · 250,000 traces

Train and evaluate computer-use agents on human demonstrations that pair screen states with the UI actions behind them, including recordings, timestamped events, outcomes, and DOM/accessibility context where available.

Request access on Hugging Face

Dataset specification

  • Production scale: 250,000 traces
  • Approximately 15,000 hours of screen activity
  • Approximately 10M timestamped UI actions
  • Approximately 8M screen states
  • 100+ task categories
  • Screen recordings
  • Screenshots
  • JSON action logs
  • DOM and accessibility metadata where available
  • Success labels
  • QA review
  • PII redaction

Intended uses

  • Training GUI and computer-use agents
  • Behavior cloning and imitation learning from human demonstrations
  • Grounding models in UI state via DOM and accessibility context
  • Evaluating multi-step task completion

Scope notes

  • The repository sample metadata illustrates structure and field distributions; production records are supplied in buyer review packages.
  • DOM and accessibility files are available for applicable subsets rather than every trace.
  • Collection-window, geography, contributor, device, and environment mix are delivery-specific and reviewed during dataset selection.

Collection scope

Environments

  • Web applications in the browser
  • Desktop operating systems and native apps
  • Productivity suites and email
  • Data entry, CRM, and admin consoles
  • Multiple locales and languages

Task coverage

  • 100+ task categories
  • Multi-step goal-directed workflows
  • Success and failure trajectories

Production scale

  • 250,000 traces
  • Approximately 15,000 hours of screen activity
  • Approximately 10M timestamped UI actions
  • Approximately 8M screen states

Buyer review package

  • Illustrative schema and sample metadata in the repository
  • Representative media, metadata, QA summaries, and data dictionary in the buyer review package
  • Delivery-specific collection window, geography, contributor, device, and environment mix reviewed during dataset selection

Annotation & metadata fields

  • Screen recordings
  • Screenshots per step
  • JSON action logs (clicks, keystrokes, scrolls, navigation)
  • Timestamps aligned to screen state
  • DOM snapshots where available
  • Accessibility-tree metadata where available
  • Task and goal labels
  • Step-level success labels

Capture methodology

  • Screen recordings paired with per-step screenshot states
  • Timestamped mouse, keyboard, scroll, click, and text-entry events
  • Task instruction, action sequence, and completion outcome captured as one trace
  • Optional browser, DOM, accessibility-tree, visible-text, and UI-element context
  • Reviewer notes and recovery steps included where available

Provenance & rights chain

  • Chain-of-custody and contributor consent documentation included in licensing review
  • Commercial license issued directly by Datoric
  • Annotation activity and quality records handled under Datoric's published privacy notice
  • Contributor compensation is administered through the collection platform and tied to recorded annotation activity

Quality, duplicates & PII

How submissions are reviewed and cleaned before they are accepted into the dataset.

  • Human-reviewed completion and success/failure labels
  • PII redaction review
  • Duplicate trajectory filtering
  • Malformed or incomplete action logs rejected

Formats & delivery

  • Video (screen recordings) and image (screenshots)
  • JSON action logs and step metadata
  • Optional DOM and accessibility files
  • Repository review materials on Hugging Face; production transfer arranged directly

Rights & license scope

Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.

Version & verification

Availability
Available
Datasheet version
July 19, 2026
Release date
July 19, 2026
Last verified
July 21, 2026
Owner
Datoric

Frequently asked

Can we inspect a schema and samples?
Yes. The repository provides the annotation schema and illustrative sample metadata after its access conditions are accepted. Request the buyer review package for real trace samples, metadata, QA summaries, and licensing documentation.
What exactly is captured per step?
Each step pairs the screen state with the UI action and timestamp. Applicable subsets also include DOM or accessibility context, visible text, and reviewer notes.
How is personal data handled?
The repository documents PII redaction review as part of QA. Our security overview and privacy notice describe the broader personal-data handling.
Can you collect additional task categories?
Yes. Custom-collection options can target additional software environments, task categories, fields, and evaluation criteria.

Related dataset specifications