Context engineering
Context engineering is deciding what evidence and instructions the model gets, in what order, and within what budget.
Useful context beats maximal context
The support assistant might have system instructions, conversation history, retrieved policy passages, tool schemas, and order details. These compete for a finite context window. More text can add distraction, conflicting versions, privacy exposure, and cost. Start by asking which facts are necessary for this decision.
How it works
Reserve output capacity before adding inputs. Preserve controlling instructions and current user intent, select relevant evidence, and trim low-value history. A character-based token estimate is only a planning heuristic; use the model's actual tokenizer or provider usage accounting for enforcement. Long-context behavior depends on task and evidence position, so test it instead of assuming all included facts will be used.
A concrete example
With a toy 4,000-token total budget and 800 tokens reserved for output, inputs must fit within 3,200. If instructions use 500, the current request 200, and evidence 1,500, only 1,000 remain for history. Blindly appending 2,000 history tokens breaks the budget.
Apply it to your assistant
Increase history to 2,000 and inspect which budget is exceeded. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
Allocate context deliberately and evaluate whether the model uses the evidence it receives.
JavaScript exercise: Context engineering · code experiment
Increase history to 2,000 and inspect which budget is exceeded.
const capacity = 4000;
const budget = { instructions: 500, request: 200, evidence: 1500, history: 1000, output: 800 };
const total = Object.values(budget).reduce((a, b) => a + b, 0);
console.log({ total, capacity, remaining: capacity - total });
console.log(total <= capacity ? 'Fits this toy budget' : 'Trim or retrieve less');
Knowledge check
What should you do before filling the entire context window with documents?
- Reserve output capacity and prioritize relevant evidence
- Duplicate every instruction
- Assume longer context always improves accuracy
Answer and explanation
Reserve output capacity and prioritize relevant evidence
Input and output constraints must be accounted for together. Relevant, non-conflicting evidence is more useful than indiscriminate volume.
Sources
- Lost in the Middle: How Language Models Use Long Contexts — Liu et al., 2023. Evidence that more context does not guarantee effective use of relevant information.
Continue learning
- What is an AI agent? — A language model proposes text. An agent system turns some of those proposals into actions, observes the result, and decides what comes next.
- The LLM core — A model predicts tokens from a context. Your application must turn that probabilistic output into a dependable interface.
- Context engineering — Context engineering is deciding what evidence and instructions the model gets, in what order, and within what budget.
- Prompting for agents — A useful prompt defines a job, the available evidence, the response contract, and what to do when the evidence is insufficient.