On-device agents
Local inference can change privacy, connectivity, and latency tradeoffs, but it brings device limits and model-distribution costs.
Local does not mean automatically private
A browser model can run on supported hardware through a runtime such as WebLLM using WebGPU. Initial model downloads, memory use, device compatibility, and battery impact matter. If the agent calls remote tools or sends telemetry, those data flows still leave the device.
How it works
Detect capabilities and present an honest fallback. Measure cold start separately from warm inference. Bound model download size and context use. Keep tool permissions explicit, and explain which operations are local versus remote. Test representative low-end devices instead of extrapolating from a development laptop.
A concrete example
A local support classifier may choose an intent offline, while order lookup still requires an authenticated network request. The classification step's local execution does not make the entire support workflow offline. This course's exercises run JavaScript locally; they do not download or run a language model.
Apply it to your assistant
Change orderLookup to local and explain what data would need to exist on the device. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
Draw the data-flow boundary and measure real device constraints before promising offline or private operation.
JavaScript exercise: On-device agents · code experiment
Change orderLookup to local and explain what data would need to exist on the device.
const stages = [{ name: 'intent', location: 'local' }, { name: 'orderLookup', location: 'remote' }, { name: 'format', location: 'local' }];
console.log({ fullyOffline: stages.every(stage => stage.location === 'local') });
console.log('Network boundaries:', stages.filter(stage => stage.location === 'remote'));
Knowledge check
A local model calls a remote order API. Is the complete workflow offline?
- Yes, because inference is local
- Yes, if the model is small
- No, the tool call still requires a network connection
Answer and explanation
No, the tool call still requires a network connection
Inference location and tool execution location are independent. A local model does not make remote dependencies disappear.
Sources
- WebLLM documentation — MLC AI, Living documentation. Practical documentation for running supported language models in the browser.
Continue learning
- On-device agents — Local inference can change privacy, connectivity, and latency tradeoffs, but it brings device limits and model-distribution costs.
- Choosing an agent framework — A framework should make your state, permissions, and failures easier to understand. Start from requirements, not a popularity list.
- Capstone: a support assistant — Connect the architecture to an evaluation plan. Your finished project should explain not only how it works, but why it is ready—or not ready—to ship.
- Reading agent case studies critically — A case study is evidence about a particular system under particular conditions. Learn to separate transferable ideas from headline claims.