The solo dev's actual constraint
A solo developer's scarce resources are cash and focus, in that order. Cloud AI spends cash to save focus; local AI spends an evening of focus to stop spending cash. The right split depends on one number you should measure before deciding: your real daily token volume. Check your API dashboard or usage page. Most working devs who lean on AI heavily are surprised — coding assistants chew through hundreds of thousands of input tokens a day in context alone.
At 1M tokens/day against Claude Sonnet pricing ($3/$15 per 1M tokens), you're spending about $198/month. A used RTX 3090 build at ~$1,240 (current street prices) running Qwen 3 32B covers a large share of that workload for ~$12/month of electricity — break-even in about 7 months, and every month after that is nearly free capacity. If your measured volume is a tenth of that, the math says keep the API key and spend the $1,240 on something else. The calculator below takes your real numbers.
What you may already own
Before buying anything: if you have a gaming PC with a 12 GB+ GPU, your local option costs zero dollars. An RTX 3060 12GB runs Llama 3.1 8B and Qwen 3 14B-class models at conversational speed; an RTX 4070-class card runs them fast. Run the analyzer against your card, install Ollama, and you have a private, unmetered assistant tonight. This is the highest-ROI move in local AI: trying it on hardware you already own.
The hybrid stack that actually works
Almost no experienced solo dev is 100% local or 100% cloud. The stable equilibrium looks like:
- Local, default: code completion and quick questions (8–14B model in your editor via Ollama's OpenAI-compatible endpoint), summarizing docs, generating test boilerplate, RAG over your own notes and codebase (local RAG guide), anything involving client code you shouldn't ship to a third party.
- Cloud, deliberate: gnarly debugging sessions, architecture discussions, unfamiliar-domain questions — the 10–20% of prompts where frontier reasoning visibly earns its price. Pay per token via API; at deliberate-use volumes this is $10–30/month, not $200.
The psychological effect is underrated: when regenerations are free, you iterate more and settle for less slop. When every retry bills you, you subtly stop experimenting. Removing the meter from your inner loop changes how you work with these tools.
Time cost, honestly
Setup is an evening: install Ollama or LM Studio, pull two models, point your editor at the local endpoint. Ongoing maintenance is genuinely minimal — ollama pull when a model updates. Where time cost is real: chasing new model releases every week (fun, optional), fine-tuning (a project, not maintenance), and multi-GPU builds (don't, as your first move). If tinkering repels you, a Mac with 24 GB+ unified memory is the appliance version — silent, zero-config, runs 14B-class models well.
The failure mode to avoid is buying a $2,600 machine to save $30/month of API spend. Measure first; the numbers make the decision boring.