# Confidence is not correctness

Canonical URL: https://agentlearn.dev/learn/evals/calibration
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 10 minutes

A model that knows when it might be wrong can be more useful than one that is always certain.

## Read a reliability diagram

Group predictions into confidence bins. Plot the mean confidence of each bin against its observed accuracy. If predictions assigned roughly 80% confidence are correct about 80% of the time, that bin is calibrated. Points below the diagonal indicate overconfidence; points above it indicate underconfidence.

## Separate calibration and accuracy

A system can be well calibrated but uninformative: always predicting the base rate is an example. Expected calibration error averages bin gaps, weighted by bin size, but depends on bin choices. Brier score averages squared probability error and captures both calibration and discrimination.

## Design an abstention policy

Choose when to defer to a person or ask for clarification. Measure accuracy on answered cases together with coverage, the fraction of cases answered. Model-written confidence is not automatically a reliable probability. Validate whichever confidence signal you actually use on representative held-out data.

## Worked example

Worked example: Among 50 fictional predictions labeled 90% confident, only 35 are correct. Observed accuracy is 70%, giving a 20 percentage-point gap in this bin. That is evidence of overconfidence on this sample, not a guarantee about every individual prediction.

## Confidence needs an observable target

A model saying “90% confident” is a prediction about correctness only if you define correctness and test that prediction across many cases. A calibrated collection of 0.9-confidence answers should be correct about 90% of the time under the measured distribution. This is a group property, not a guarantee about any one answer.

A reliability diagram groups predictions into bins and compares average predicted confidence with observed success frequency. Bin definitions and sample sizes matter: a bin with three examples provides little evidence. Distribution shift can break previously observed calibration.

## Calculate a proper scoring rule

For binary correctness, the Brier score is the mean of `(p − y)²`, where p is predicted probability and y is 0 or 1. Lower is better. For probabilities [0.9, 0.8, 0.6, 0.2] and outcomes [1, 0, 1, 0], the squared errors are [0.01, 0.64, 0.16, 0.04], averaging 0.2125.

The confidently wrong 0.8 prediction contributes most of the error. This is why simply increasing verbal confidence is not improvement.

**Your experiment:** choose an abstention threshold on development data, then report coverage and error rate on held-out data. A lower answered-case error rate can be purchased by answering fewer questions; show both quantities.


## Key takeaway

Validate confidence empirically, and report what happens when the system abstains.

## Knowledge check

A bin has 90% average confidence and 60% accuracy. The model is…

1. Underconfident in this bin
2. Perfectly calibrated
3. Overconfident in this bin

Answer: Overconfident in this bin

Its stated confidence exceeds its empirical success rate by 30 percentage points. More data is needed to estimate that gap precisely.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
