# Quality, latency & cost tradeoffs

Canonical URL: https://agentlearn.dev/learn/evals/cost-frontier
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 11 minutes

The best system is often the one that satisfies the task at an acceptable operating cost.

## Measure complete requests

Count input and output tokens, retries, retrieval, tool execution, and verifier calls. Report latency percentiles, such as median and p95, instead of only an average. State concurrency, warmup, caching, and region because these conditions affect operational measurements.

## Find the Pareto frontier

A system is dominated when another measured configuration is at least as good on every chosen objective and strictly better on at least one. Nondominated systems form a Pareto frontier. A frontier is descriptive: it does not choose the right tradeoff for you or account for unmeasured qualities.

## Turn tradeoffs into a policy

Set minimum quality and maximum latency requirements before minimizing cost. Consider routing easy requests to a cheaper system and escalating uncertain cases, then evaluate the complete routing policy. A per-token price alone cannot tell you the cost per successfully completed user task.

## Worked example

Worked example: System A costs $0.02 per fictional task and passes 80%; B costs $0.05 and passes 88%; C costs $0.07 and passes 85%. B dominates C on cost and quality. Neither A nor B dominates the other. The lab lets you set a minimum quality requirement.

## Separate dominance from feasibility

A configuration is dominated when another is at least as good on all considered objectives and strictly better on one. On a two-dimensional cost–quality plot, a cheaper configuration with equal or higher quality dominates the more expensive one. Add latency as a third objective and the frontier may change.

Feasibility comes first. If the product requires a minimum quality level, a cheap configuration below that requirement is not a valid recommendation. Critical-failure constraints may exclude a high-average-quality configuration as well.

## Account for uncertainty and scope

The lab uses invented configurations with exact illustrative numbers. Real quality estimates have uncertainty, costs vary with request length and retries, and latency depends on load and geography. A tiny score difference may not justify declaring one configuration superior.

Suppose A costs $0.01 per request and resolves 80%, while B costs $0.012 and resolves 96%. Under this simplified accounting both cost $0.0125 per resolution. Human escalations or tool costs can change that result. State whether the measurement includes those costs.

**Your experiment:** set the quality floor, identify feasible options, and explain why the cheapest feasible option is not necessarily the best under a latency or risk constraint. Never treat the lab's synthetic dollars as current vendor pricing.


## Key takeaway

Compare the full workflow and choose among feasible tradeoffs.

## Knowledge check

A is cheaper and more accurate than B, with other objectives equal. B is…

1. Dominated by A
2. Necessarily the better deployment
3. On the cost-quality frontier

Answer: Dominated by A

A improves cost and quality, so B is dominated for these measured objectives. Additional objectives could change the decision.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
