AgentLearn

Accuracy, precision, recall & F1

The same predictions can look excellent or terrible depending on the question your metric asks.

Count outcomes first

For a binary detector, true positives are positive cases flagged correctly. False positives are negative cases flagged incorrectly. False negatives are missed positive cases. True negatives are correctly rejected negative cases. A confusion matrix exposes all four counts before compressing them into a score.

Match the metric to the cost

Precision is TP / (TP + FP): among flagged items, how many are positive? Recall is TP / (TP + FN): among positive items, how many were found? Accuracy is (TP + TN) / N. F1 is 2PR / (P + R), which balances precision and recall but ignores true negatives and does not encode business costs.

Thresholds create tradeoffs

Lowering a classifier threshold usually increases recall and false positives. Choose the threshold on development data, based on the cost of misses and false alarms. When a denominator is zero, declare how the metric is handled. Report prevalence and counts so readers can interpret the percentages.

Worked example

Worked example: In 100 messages, 10 are unsafe. A detector flags 12 messages: 8 correctly, 4 incorrectly. It misses 2. Precision = 8/12 = 66.7%; recall = 8/10 = 80%; F1 = 72.7%. A detector that flags nothing has 90% accuracy but zero recall.

Calculate the confusion matrix

Consider a synthetic detector that flags policy-violating answers. Among 100 answers, 25 truly violate policy. At one threshold, the detector finds 18 violations, misses 7, and incorrectly flags 12 acceptable answers. The remaining 63 answers are correctly accepted.

OutcomeCount
True positive: violation correctly flagged18
False positive: acceptable answer flagged12
False negative: violation missed7
True negative: acceptable answer accepted63

Precision = TP / (TP + FP) = 18 / 30 = 60%. Recall = TP / (TP + FN) = 18 / 25 = 72%. F1 = 2TP / (2TP + FP + FN) = 36 / 55, approximately 65.5%. Accuracy = (TP + TN) / N = 81%.

Choose the operating point

Lowering a threshold usually increases recall and also increases false positives. The right operating point depends on the consequence of a miss and the cost of review. Do not tune the threshold on the final test set and then report that set as untouched evidence.

When no answers are flagged, precision has a zero denominator; it is undefined, not perfect. The lab displays unavailable metrics explicitly. With multiple classes, macro averages weight classes equally, while micro aggregation weights individual decisions.

Your experiment: find a threshold with high recall, then count the additional human-review workload.

Key takeaway

A metric is a choice about which errors matter.

Knowledge check

You need to minimize missed unsafe messages. Which metric directly tracks finding them?

  1. Recall
  2. Precision
  3. Overall accuracy
Answer and explanation

Recall

Recall measures the share of actual positives detected. Monitor false positives too; maximizing recall alone can mean flagging everything.

Sources

Read this lesson as Markdown

Continue learning