Coding is two different workloads
The "AI coding" bill conflates two very different things. Volume work — completions, docstrings, test scaffolds, "what does this error mean" — is enormous in token count (an editor assistant re-sends context on every keystroke pause; 3M tokens/day is normal for a heavy user) and modest in difficulty. Depth work — cross-file debugging, architecture trade-offs, unfamiliar-framework surgery — is maybe 5% of tokens and 50% of the value.
That split is the whole answer. Volume work is exactly what local models are now good at: Qwen 2.5 Coder 32B on a 24 GB card is a genuinely strong completion and everyday-questions model, and 7–14B coder models are more than adequate for autocomplete at speeds no cloud API can match, because there is no network in the loop. Depth work is where frontier models still clearly win, and where paying $3/$15 per million tokens (Anthropic pricing) is rational — you're buying hours of your own debugging time back.
The economics of the volume tier
At 3M tokens/day against Claude Sonnet, the assistant bill is roughly $594/month. An RTX 4090 workstation (~$2,590) running the volume tier locally consumes about $34/month of electricity at 14 hours of daily load. Even keeping a frontier key for the depth tier, moving the bulk offline breaks even in about 5 months — see the calculator below with your own volume. If your usage is lighter (say 500k tokens/day), break-even stretches to ~2.5 years and the case weakens to latency and privacy; measure your dashboard before buying.
One honest complication: coding contexts are large, and context is a local resource. A 32B model at Q4 with a 32k context fits a 24 GB card; push toward 128k and you need smarter context management (most editor integrations do retrieval-style trimming anyway) or more VRAM. Cloud models make long context somebody else's memory problem — a real advantage for repo-scale questions.
Latency: the underrated local win
Completion UX lives or dies at the 100ms scale. A local 7B coder model starts streaming in tens of milliseconds; a cloud round trip adds 100–500ms before the first token on a good day. Over hundreds of completions daily, that's the difference between the assistant feeling like part of the editor and feeling like a suggestion popup you wait for. Developers who try local completion rarely go back for that reason alone — quality parity at the completion tier arrived quietly in 2025.
The client-code problem
If you do contract work, your MSA very likely restricts sending client code to third parties, and "but the API terms say no training" does not amend a contract. A local model is the clean answer: NDA code never transits anyone's infrastructure. This single constraint pushes more professional devs to local coding setups than the cost math does — the privacy comparison covers the general case.
A working setup, concretely
Editor: Continue or Cline pointed at Ollama's OpenAI-compatible endpoint. Models: a 7B coder model for tab-completion (fast lane) + Qwen 2.5 Coder 32B for chat/refactor (quality lane) — Ollama hot-swaps them. Hardware: 16 GB VRAM minimum for the two-model setup, 24 GB comfortable (what fits your card). Keep the frontier API key wired into the same tools for the depth tier; the point is routing, not abstinence.