Accuracy, precision, recall & F1
The same predictions can look excellent or terrible depending on the question your metric asks.
Count outcomes first
For a binary detector, true positives are positive cases flagged correctly. False positives are negative cases flagged incorrectly. False negatives are missed positive cases. True negatives are correctly rejected negative cases. A confusion matrix exposes all four counts before compressing them into a score.
Match the metric to the cost
Precision is TP / (TP + FP): among flagged items, how many are positive? Recall is TP / (TP + FN): among positive items, how many were found? Accuracy is (TP + TN) / N. F1 is 2PR / (P + R), which balances precision and recall but ignores true negatives and does not encode business costs.
Thresholds create tradeoffs
Lowering a classifier threshold usually increases recall and false positives. Choose the threshold on development data, based on the cost of misses and false alarms. When a denominator is zero, declare how the metric is handled. Report prevalence and counts so readers can interpret the percentages.
Worked example
Worked example: In 100 messages, 10 are unsafe. A detector flags 12 messages: 8 correctly, 4 incorrectly. It misses 2. Precision = 8/12 = 66.7%; recall = 8/10 = 80%; F1 = 72.7%. A detector that flags nothing has 90% accuracy but zero recall.
Calculate the confusion matrix
Consider a synthetic detector that flags policy-violating answers. Among 100 answers, 25 truly violate policy. At one threshold, the detector finds 18 violations, misses 7, and incorrectly flags 12 acceptable answers. The remaining 63 answers are correctly accepted.
| Outcome | Count |
|---|---|
| True positive: violation correctly flagged | 18 |
| False positive: acceptable answer flagged | 12 |
| False negative: violation missed | 7 |
| True negative: acceptable answer accepted | 63 |
Precision = TP / (TP + FP) = 18 / 30 = 60%. Recall = TP / (TP + FN) = 18 / 25 = 72%. F1 = 2TP / (2TP + FP + FN) = 36 / 55, approximately 65.5%. Accuracy = (TP + TN) / N = 81%.
Choose the operating point
Lowering a threshold usually increases recall and also increases false positives. The right operating point depends on the consequence of a miss and the cost of review. Do not tune the threshold on the final test set and then report that set as untouched evidence.
When no answers are flagged, precision has a zero denominator; it is undefined, not perfect. The lab displays unavailable metrics explicitly. With multiple classes, macro averages weight classes equally, while micro aggregation weights individual decisions.
Your experiment: find a threshold with high recall, then count the additional human-review workload.
Key takeaway
A metric is a choice about which errors matter.
Knowledge check
You need to minimize missed unsafe messages. Which metric directly tracks finding them?
- Recall
- Precision
- Overall accuracy
Answer and explanation
Recall
Recall measures the share of actual positives detected. Monitor false positives too; maximizing recall alone can mean flagging everything.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- Accuracy, precision, recall & F1 — The same predictions can look excellent or terrible depending on the question your metric asks.
- Scoring open-ended generation — There can be many good answers to a question, and a fluent answer can still be wrong.
- Confidence is not correctness — A model that knows when it might be wrong can be more useful than one that is always certain.