Datasets / Alignment

RLHF-Voice Dataset

Build-to-spec · preference pairs

Align and evaluate voice models with human preference data that connects prompts and candidate responses to rankings, spoken rationales, reviewer agreement, and program-specific criteria.

Dataset specification

  • Paired or multi-candidate voice responses
  • Human preference rankings
  • Spoken rationale audio and transcripts
  • Prompt and response metadata
  • Language, locale, and accent metadata
  • Audio quality and recording-condition fields
  • Program-specific evaluation criteria
  • Reviewer agreement and QA fields
  • WAV audio with CSV and JSON metadata

Intended uses

  • Voice-model preference optimization
  • Reward-model and preference-model training
  • Voice-agent alignment and evaluation
  • Comparative evaluation of response quality, style, and safety

Scope notes

  • Target models, languages, response types, reviewer mix, and evaluation criteria are defined for each program.
  • Human preferences are rubric- and population-dependent; the pilot establishes agreement and known disagreement before scaling.
  • This is a build-to-spec program, so final schema, volume, and delivery fields are confirmed during scoping.

Collection scope

Comparison units

  • Prompts paired with two or more candidate voice responses
  • Human rankings or pairwise preferences
  • Spoken rationales and transcripts when included in the program

Evaluation criteria

  • Naturalness, helpfulness, and instruction adherence
  • Style, tone, and conversational fit
  • Safety and program-specific dimensions

Program configuration

  • Target models and response candidates
  • Languages, locales, accents, and interaction types
  • Ranking rubric, rationale format, and acceptance criteria

Buyer review package

  • Program specification, annotation schema, and data dictionary
  • Pilot preference pairs, rationales, metadata, and QA summaries
  • Target models, languages, reviewer mix, and production scope confirmed before scaling

Annotation & metadata fields

  • Prompt, task, and comparison identifiers
  • Candidate response identifiers and ordering
  • Pairwise preference or ranked-choice label
  • Evaluation criteria and per-dimension scores
  • Spoken rationale audio and transcript when included
  • Language, locale, accent, and interaction metadata
  • Reviewer agreement and adjudication fields
  • Audio quality, completion, and QA status

Capture methodology

  • Candidate voice responses grouped under a shared prompt and evaluation rubric
  • Candidate ordering, reviewer instructions, and comparison protocol defined before collection
  • Human reviewers submit preferences and spoken rationales when required by the program
  • Rankings, rationale audio, transcripts, and metadata aligned to each comparison unit

Provenance & rights chain

  • Reviewer consent and chain-of-custody documentation included in licensing review
  • Commercial license issued directly by Datoric
  • Prompt, model-output, reviewer, and annotation lineage recorded for the selected program
  • Collection and annotation activity handled under Datoric's published privacy notice

Quality, duplicates & PII

How submissions are reviewed and cleaned before they are accepted into the dataset.

  • Reviewer qualification and calibration against the agreed rubric
  • Agreement, consistency, and adjudication checks
  • Audio quality and transcript review for spoken rationales
  • Duplicate, malformed, incomplete, and low-information responses handled under the acceptance criteria

Formats & delivery

  • WAV or PCM rationale audio when included
  • CSV preference and evaluation records
  • JSON prompt, response, ranking, and rationale records
  • Data dictionary, rubric, and QA summary

Rights & license scope

Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.

Version & verification

Availability
Build-to-spec
Datasheet version
July 23, 2026
Release date
July 23, 2026
Last verified
July 23, 2026
Owner
Datoric

Frequently asked

What does each preference record contain?
A record can pair a prompt with two or more candidate voice responses, a human preference or ranking, per-dimension evaluations, and a spoken rationale with transcript when that field is included in the program.
Can the evaluation rubric match our model and product?
Yes. The pilot defines the target models, response types, languages, ranking method, evaluation dimensions, rationale format, and acceptance criteria before production collection begins.
How is reviewer consistency measured?
The QA plan can include qualification, calibration items, agreement checks, repeat items, adjudication, and low-information response review, with thresholds agreed during scoping.
Can we review a pilot before scaling?
Yes. The buyer review package includes pilot preference pairs, rationales, metadata, reviewer-agreement results, QA summaries, and licensing documentation for the proposed program.

Related dataset specifications