Retrieval-augmented generation
RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully.
Two chances to fail
A retrieval system can miss the relevant policy, retrieve an outdated policy, or include irrelevant passages. A generator can then ignore correct evidence or make claims the evidence does not support. Evaluate retrieval and answer quality separately so a single end-to-end score does not hide the cause.
How it works
Build a versioned corpus with stable document IDs and access controls. Choose chunk boundaries that preserve meaning, retrieve candidate passages, optionally rerank, and pass a small evidence set to the model. For a labeled query set, measure retrieval recall and precision at a chosen k. For answers, check correctness, support, citation accuracy, and appropriate abstention.
A concrete example
If two policy passages are relevant and top-3 retrieves one relevant and two irrelevant passages, recall@3 is 1/2 and precision@3 is 1/3. Increasing k may improve recall while adding distraction. The retrieval lab lets you see that tradeoff using a small lexical corpus, not live embeddings.
Apply it to your assistant
Include policy-2 in retrieved and recompute both metrics. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
Measure evidence selection and evidence use as separate stages.
JavaScript exercise: Retrieval-augmented generation · code experiment
Include policy-2 in retrieved and recompute both metrics.
const relevant = new Set(['policy-1', 'policy-2']);
const retrieved = ['policy-1', 'shipping-1', 'warranty-1'];
const hits = retrieved.filter(id => relevant.has(id)).length;
console.log({ precision: hits / retrieved.length, recall: hits / relevant.size });
Knowledge check
Two passages are relevant; the top three results contain one of them. What is recall@3?
- 1/3
- 1/2
- 3/2
Answer and explanation
1/2
Recall is retrieved relevant items divided by all relevant items: 1/2. Precision would be 1/3.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., 2020. A foundational approach to combining retrieval with text generation.
Continue learning
- Retrieval-augmented generation — RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully.
- Testing & debugging agents — A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.
- Agent security & prompt injection — An agent can encounter instructions inside data: a retrieved page, email, file, or tool result. Those instructions must not acquire authority.
- Evaluating the complete agent — A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.