# Separate retrieval from answer quality

Canonical URL: https://agentlearn.dev/learn/evals/rag
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 13 minutes

When a grounded assistant fails, locate the failure before changing the prompt.

## Evaluate the retriever

Label which documents or passages support each question. Recall@k measures the share of relevant items retrieved in the top k; the exact definition of relevance and unit of measurement matters. Ranking metrics can reward putting helpful evidence earlier. Report latency and whether the corpus actually contains the answer.

## Evaluate the generator

Score answer correctness, supported claims, citation accuracy, and appropriate abstention separately. A citation that exists is not necessarily a citation that supports the claim. For questions without evidence in the corpus, a good assistant should follow the product’s explicit policy for uncertainty or escalation.

## Run an oracle-context experiment

Give the generator human-selected supporting passages. If quality improves substantially, retrieval is a likely bottleneck. If it still fails, investigate instruction following, reasoning, or the rubric. This intervention narrows the diagnosis, but the production system still needs end-to-end evaluation.

## Worked example

Worked example: On 60 fictional questions, the answer is retrieved for 45. The assistant answers 36 of those correctly. Retrieval coverage is 75%; conditional generation accuracy is 80%; total correct answers are 36/60 = 60% if none of the other cases is answered correctly.

## Decompose the pipeline

Measure corpus coverage first: does an authoritative answer exist in the available documents? Then evaluate retrieval: did the relevant passages reach the context? Finally evaluate generation: did the answer use those passages faithfully? An end-to-end failure can originate at any of these stages.

Retrieval precision@k is relevant retrieved items divided by retrieved items; recall@k is relevant retrieved items divided by all relevant items under your labeling scheme. Stable passage IDs and carefully defined relevance labels are essential. Different chunking can change the denominator, so compare configurations with an explicit evaluation unit.

## Score citations and support separately

A citation can name a real document without supporting the attached claim. Check both identifier validity and entailment or support. An answer can also be factually correct by outside knowledge but unsupported by the supplied corpus. Decide whether that counts as a failure for your product.

Use an oracle-context experiment to isolate generation: supply the known relevant passages directly and see whether the answer improves. Conversely, score retrieved passages without generating an answer to inspect retrieval alone.

**Your experiment:** open the retrieval-ranking lab at /labs/5. Compare top-1 and top-2 for the personalized-product question. Explain why the exception passage changes eligibility even when the general returns passage looks relevant.


## Key takeaway

Measure retrieval, generation given evidence, and the complete user outcome.

## Knowledge check

Why test with hand-selected supporting passages?

1. To isolate whether retrieval is limiting answer quality
2. To prove the production retriever is correct
3. To remove the need for end-to-end testing

Answer: To isolate whether retrieval is limiting answer quality

Oracle context is a diagnostic intervention. It helps separate retrieval failures from failures that remain when good evidence is supplied.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
