Why agents change the token math
A chat assistant answers a question with one model call. An agent works a task: plan, call a tool, read the result, re-plan, retry the failed step, summarize — routinely 20–100 model calls per task, each carrying accumulated context. Token consumption runs 10–100× equivalent chat, and it's input-heavy (the agent re-reads its own transcript every step). This is why teams that casually shipped an agent feature discover it atop their API bill a month later — and why agents are the workload where local hardware's "loops are free" property is most valuable. An agent that retries a flaky scrape 50 times overnight costs pennies of electricity locally; the same persistence on Claude Sonnet pricing is a line item someone will question.
The compounding-error problem
Here's the honest physics working against local agents. If a model succeeds at an individual step 95% of the time, a 20-step chain completes cleanly about 36% of the time; at 99% per-step, 82%. Small per-step quality gaps compound into large end-to-end reliability gaps — which is exactly why frontier models, whose per-step judgment and tool-calling discipline remain ahead, dominate open-ended agent work (research tasks, multi-app workflows, anything requiring judgment about ambiguity). No amount of free electricity fixes an agent that wanders. Open models have closed much of this gap for scoped agents — fixed toolsets, clear success criteria, retry-friendly tasks — where Llama 3.3 70B and Qwen 3 32B-class models with well-engineered harnesses run reliably. But pretending a local 32B matches frontier per-step judgment on open-ended chains sells you a broken pipeline.
The router pattern
Production agent stacks in 2026 converge on splitting the roles:
- Planner/critic (frontier API): decomposes tasks, reviews results, handles ambiguity — low call volume, high judgment value. This is where the $3/$15 tokens earn their price.
- Executors (local): the high-volume grunt work — extraction, summarization per document, format conversion, code generation attempts, each tool-loop iteration. 80–95% of tokens, quality-tolerant, retry-friendly.
- Verification (local, cheap): schema checks, unit tests, regex validation — free locally, and the layer that converts unreliable steps into reliable pipelines.
The economics follow the split: at ~1.5M executor-tier tokens/day, a dual-RTX-3090 build ($2,870, 48 GB — runs Llama 3.3 70B at Q4) breaks even against Sonnet pricing in about 13 months, while the planner's API bill drops to tens of dollars. Rate limits also disappear from the executor tier — often the binding constraint before cost, since a fleet of concurrent agents exhausts API concurrency tiers quickly.
Practical notes from the trenches
Local agent stacks run through the same OpenAI-compatible endpoints as everything else (Ollama/vLLM), so frameworks like LangGraph or CrewAI point at localhost unchanged. Three hard-won tips: quantize conservatively (Q4_K_M is fine for chat, but agents feel quality loss in tool-call formatting — prefer Q5/Q6 for executor models if VRAM allows); cap loop budgets (local's free retries breed lazy harnesses — a loop that can't fail loudly will fail silently); and sandbox properly (an agent with shell access on your own network is your security problem — cloud providers' safety layers aren't in the path anymore). For burst capacity during development, rented GPUs beat owning a second box.