RLHF-Voice Dataset
Build-to-spec · preference pairs
Align and evaluate voice models with human preference data that connects prompts and candidate responses to rankings, spoken rationales, reviewer agreement, and program-specific criteria.
Dataset specification
- Paired or multi-candidate voice responses
- Human preference rankings
- Spoken rationale audio and transcripts
- Prompt and response metadata
- Language, locale, and accent metadata
- Audio quality and recording-condition fields
- Program-specific evaluation criteria
- Reviewer agreement and QA fields
- WAV audio with CSV and JSON metadata
Intended uses
- Voice-model preference optimization
- Reward-model and preference-model training
- Voice-agent alignment and evaluation
- Comparative evaluation of response quality, style, and safety
Scope notes
- Target models, languages, response types, reviewer mix, and evaluation criteria are defined for each program.
- Human preferences are rubric- and population-dependent; the pilot establishes agreement and known disagreement before scaling.
- This is a build-to-spec program, so final schema, volume, and delivery fields are confirmed during scoping.
Collection scope
Comparison units
- Prompts paired with two or more candidate voice responses
- Human rankings or pairwise preferences
- Spoken rationales and transcripts when included in the program
Evaluation criteria
- Naturalness, helpfulness, and instruction adherence
- Style, tone, and conversational fit
- Safety and program-specific dimensions
Program configuration
- Target models and response candidates
- Languages, locales, accents, and interaction types
- Ranking rubric, rationale format, and acceptance criteria
Buyer review package
- Program specification, annotation schema, and data dictionary
- Pilot preference pairs, rationales, metadata, and QA summaries
- Target models, languages, reviewer mix, and production scope confirmed before scaling
Annotation & metadata fields
- Prompt, task, and comparison identifiers
- Candidate response identifiers and ordering
- Pairwise preference or ranked-choice label
- Evaluation criteria and per-dimension scores
- Spoken rationale audio and transcript when included
- Language, locale, accent, and interaction metadata
- Reviewer agreement and adjudication fields
- Audio quality, completion, and QA status
Capture methodology
- Candidate voice responses grouped under a shared prompt and evaluation rubric
- Candidate ordering, reviewer instructions, and comparison protocol defined before collection
- Human reviewers submit preferences and spoken rationales when required by the program
- Rankings, rationale audio, transcripts, and metadata aligned to each comparison unit
Provenance & rights chain
- Reviewer consent and chain-of-custody documentation included in licensing review
- Commercial license issued directly by Datoric
- Prompt, model-output, reviewer, and annotation lineage recorded for the selected program
- Collection and annotation activity handled under Datoric's published privacy notice
Quality, duplicates & PII
How submissions are reviewed and cleaned before they are accepted into the dataset.
- Reviewer qualification and calibration against the agreed rubric
- Agreement, consistency, and adjudication checks
- Audio quality and transcript review for spoken rationales
- Duplicate, malformed, incomplete, and low-information responses handled under the acceptance criteria
Formats & delivery
- WAV or PCM rationale audio when included
- CSV preference and evaluation records
- JSON prompt, response, ranking, and rationale records
- Data dictionary, rubric, and QA summary
Rights & license scope
Licensed directly by Datoric for commercial AI training, with final scope controlled by the signed agreement for the selected delivery.
Version & verification
- Availability
- Build-to-spec
- Datasheet version
- July 23, 2026
- Release date
- July 23, 2026
- Last verified
- July 23, 2026
- Owner
- Datoric
Frequently asked
- What does each preference record contain?
- A record can pair a prompt with two or more candidate voice responses, a human preference or ranking, per-dimension evaluations, and a spoken rationale with transcript when that field is included in the program.
- Can the evaluation rubric match our model and product?
- Yes. The pilot defines the target models, response types, languages, ranking method, evaluation dimensions, rationale format, and acceptance criteria before production collection begins.
- How is reviewer consistency measured?
- The QA plan can include qualification, calibration items, agreement checks, repeat items, adjudication, and low-information response review, with thresholds agreed during scoping.
- Can we review a pilot before scaling?
- Yes. The buyer review package includes pilot preference pairs, rationales, metadata, reviewer-agreement results, QA summaries, and licensing documentation for the proposed program.