# Build a representative dataset

Canonical URL: https://agentlearn.dev/learn/evals/dataset-design
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 10 minutes

What you choose to test determines what you are able to discover.

## Map the input space

List the workflows, languages, difficulty levels, and failure costs in your application. Sample realistic inputs from each relevant group. Include ordinary traffic and deliberately difficult cases, but label those groups separately: an adversarial challenge set is useful without being representative of daily traffic.

## Separate development from measurement

Use a development set to iterate on prompts and a held-out test set to estimate performance after those choices are fixed. If you repeatedly inspect test failures and tune against them, that set becomes development data. Keep related examples, such as messages from one conversation, in the same split to avoid leakage.

## Make labels auditable

Record the expected answer or rubric, where it came from, and the dataset version. Have a second annotator review ambiguous examples. Document disagreements instead of silently treating one opinion as ground truth. Remove personal information you do not need for the task.

## Worked example

Worked example: A dataset with 900 English questions and 100 Spanish questions gives an overall score dominated by English. Report both language slices. If you oversample Spanish to diagnose failures, use production traffic weights only when estimating a production-wide score.

## Design the sampling frame

Start with the requests the product should handle, not examples that are easy to write. Include common questions in realistic proportions, and create explicit challenge slices for consequential exceptions. A representative sample estimates ordinary traffic behavior; an adversarial set probes failure modes. Do not combine them into one unlabeled average and call it production accuracy.

For a fictional 100-case suite, use 60 ordinary returns questions, 20 exception cases, and 20 requests with missing evidence. These counts are an instructional choice, not a universal recipe. Record how cases were collected, which users and languages they cover, and what was excluded.

## Split at the right unit

If one support conversation produces five paraphrased questions, splitting individual questions can put near-duplicates into both development and test sets. Group related cases by conversation, customer, document family, or another leakage-relevant unit before splitting. Keep a separate final test set that is not repeatedly used to tune prompts.

A synthetic generator can expand coverage but can also reproduce its own style and blind spots. Human-review generated cases, verify reference answers, and compare their distribution with real allowed-use traffic.

**Your artifact:** a dataset card listing intended use, sampling, labels, privacy handling, split unit, known gaps, and version.


## Key takeaway

Represent your users, preserve a holdout, and keep important slices visible.

## Knowledge check

You have tuned your prompt after reading every test-set error. What next?

1. Report the same test score as an unbiased estimate
2. Delete the hardest examples
3. Evaluate the fixed prompt on a fresh holdout

Answer: Evaluate the fixed prompt on a fresh holdout

The original test set has influenced development. A fresh, representative holdout provides a less biased estimate of generalization.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
