Capstone: a support assistant
Connect the architecture to an evaluation plan. Your finished project should explain not only how it works, but why it is ready—or not ready—to ship.
One narrow product, complete boundaries
Build a read-only support assistant that answers return-policy questions using a small, versioned corpus. It must cite evidence, ask for missing information, and abstain when no policy supports an answer. Keep refunds out of scope for the first release. This gives you a tractable system with meaningful failure cases.
How it works
Implement retrieval, bounded orchestration, output validation, and redacted traces. Assemble a development set and a held-out set with ordinary questions, exceptions, outdated policies, unanswerable requests, and prompt-injection attempts. Version the prompt and configuration. Record correctness, support, abstention, latency, and cost for each run.
A concrete example
Deliver four artifacts: a runnable assistant, a dataset with scoring guidance, a reproducible evaluation report, and a written release decision. The evals track teaches the statistics and grader design needed to defend that decision. The starter below shows a toy release report, not a production-ready assistant or a statistically adequate sample.
Apply it to your assistant
Add a failed critical case and confirm it blocks release. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
A complete agent project includes its evidence, failure analysis, and release criteria.
JavaScript exercise: Capstone: a support assistant · code experiment
Add a failed critical case and confirm it blocks release.
const cases = [
{ id: 'ordinary', pass: true, critical: false },
{ id: 'unsupported-policy', pass: true, critical: true },
{ id: 'injection', pass: true, critical: true },
];
const criticalFailures = cases.filter(c => c.critical && !c.pass);
console.log({ cases: cases.length, criticalFailures: criticalFailures.length, gate: criticalFailures.length ? 'BLOCK' : 'Needs full quality review' });
Knowledge check
Which artifact is essential beyond a working happy-path demo?
- A reproducible evaluation report and explicit release decision
- A claim that it is autonomous
- A larger logo
Answer and explanation
A reproducible evaluation report and explicit release decision
A demo shows possibility. A versioned evaluation and release rationale show what behavior was tested and what uncertainty remains.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- On-device agents — Local inference can change privacy, connectivity, and latency tradeoffs, but it brings device limits and model-distribution costs.
- Choosing an agent framework — A framework should make your state, permissions, and failures easier to understand. Start from requirements, not a popularity list.
- Capstone: a support assistant — Connect the architecture to an evaluation plan. Your finished project should explain not only how it works, but why it is ready—or not ready—to ship.
- Reading agent case studies critically — A case study is evidence about a particular system under particular conditions. Learn to separate transferable ideas from headline claims.