AgentLearn

Sample size & confidence intervals

82% on 50 examples and 82% on 5,000 examples are very different amounts of evidence.

A sample rate is an estimate

For binary independent outcomes with a common success probability, the number of passes follows a binomial model. The observed rate p̂ = passes / n estimates that probability. Its approximate standard error is √(p̂(1−p̂)/n), so four times as many independent examples roughly halves the standard error.

Use a sensible interval

The lab uses a 95% Wilson score interval, which behaves better than a simple normal interval near zero or one and for small samples. Under repeated sampling, a 95% confidence procedure aims to produce intervals containing the fixed population proportion about 95% of the time. It is not a 95% posterior probability statement about this particular interval.

More data does not fix bias

Duplicated questions, correlated conversations, and repeated outputs from the same prompt violate a naive independence assumption. A million unrepresentative examples can produce a very precise estimate of the wrong population. Sample at the right unit and inspect coverage before increasing n.

Worked example

Worked example: 82 passes out of 100 gives a Wilson 95% interval of approximately 73.3%–88.3%. Zero failures out of a small sample does not imply zero failure risk: even a perfect observed score has an interval with a lower bound below 100%.

Calculate a confidence interval

A pass rate is an estimate from a sample. For independent binary cases, the Wilson interval is a useful alternative to a naive normal interval, especially near 0% or 100%. Let p̂ = successes/n and z approximately 1.96 for a two-sided 95% interval:

center = (p̂ + z²/(2n)) / (1 + z²/n)

halfWidth = z × sqrt(p̂(1−p̂)/n + z²/(4n²)) / (1 + z²/n)

The interval is center ± halfWidth. The lab calculates this directly. Even a perfect observed score has uncertainty: zero observed failures is not proof of zero failure risk.

State the assumptions

The usual interpretation concerns repeated sampling: the procedure covers the fixed population rate in about 95% of repeated samples under its assumptions. It is not a 95% probability that this particular model is safe, nor does it account for biased sampling, missing task categories, or an incorrect grader.

Correlated cases reduce effective information. Ten paraphrases from the same conversation are not equivalent to ten independent conversations. Repeated model generations for one case also form a cluster. Resample or model at the appropriate unit rather than pretending every output is independent.

Your experiment: keep the success rate fixed and compare n = 25, 100, and 400. Roughly quadrupling n halves uncertainty near the same rate; it does not fix a biased dataset.

Key takeaway

Report the score, the sample size, the interval method, and the sampling assumptions.

Knowledge check

Roughly how much independent data halves standard error?

  1. Twice as much
  2. Four times as much
  3. Ten times as much
Answer and explanation

Four times as much

Standard error scales approximately as 1/√n, holding the underlying success rate fixed. Increasing n by four gives a factor of one-half.

Sources

Read this lesson as Markdown

Continue learning