The only three numbers that matter
Every local-vs-cloud cost argument reduces to three numbers: your daily token volume, the per-token price of the cloud model you would otherwise use, and the up-front price of hardware capable of a comparable job. Everything else — electricity, amortization period, input/output mix — moves the answer by weeks, not months.
Cloud pricing is public and precise. As of our last verification, OpenAI charges $2.50 per million input tokens and $10.00 per million output tokens for GPT-4o, Anthropic charges $3.00/$15.00 for Claude Sonnet, and hosted open models are dramatically cheaper — Together.ai serves Llama 3.1 8B at $0.18 per million tokens in either direction. That spread matters: the honest cloud comparison for a local 8B model is the $0.18 hosted version of the same model, not a frontier model it can't match.
On the local side, the dominant cost is the GPU. A used RTX 3090 with 24 GB of VRAM — still the best value in local AI, as our GPU buyer's guide argues — anchors a complete build at about $1,240 (street prices, checked July 2026, see /build). That machine runs 30B-class models at Q4 quantization: genuinely useful quality, a tier below frontier.
The math, worked
Take a heavy user: 2 million tokens a day at a 70/30 input/output split.
- Cloud (GPT-4o): (1.4M × $2.50 + 0.6M × $10.00) / 1M ≈ $9.50/day → ~$285/month.
- Local (used RTX 3090 build): at ~45 tokens/sec for a 30B model, 2M tokens is ~12.3 hours of load. At 440 W under load and $0.15/kWh that is ~$24/month of electricity. Add the $1,240 build amortized over 24 months ($52/month) → ~$76/month all-in.
Break-even: $1,240 ÷ ($285 − $24) ≈ 4.8 months. After that, the marginal cost of every additional token is electricity — roughly $0.40 per million tokens, an order of magnitude below even budget cloud tiers.
Now the honest inversions. Drop the volume to 200k tokens/day and the cloud bill falls to ~$28/month; break-even stretches past three and a half years — longer than you should plan around a used GPU. Or keep the volume but switch the cloud comparison to GPT-4o Mini ($0.15/$0.60): the cloud bill is ~$17/month and local never breaks even on cost alone. If Mini-class quality genuinely covers your workload, cloud wins the pure cost argument at almost any volume.
What the calculator assumes (and what it ignores)
The calculator below uses the same math as this page — one shared module, not marketing arithmetic. Assumptions: 24-month amortization, zero resale value (conservative — 3090s resell well), electricity billed only for load time, and no cloud volume discounts. It ignores things that are real but unquantifiable per reader: your time to set up (an evening with Ollama, realistically), the value of offline capability, and the quality gap between a 30B open model and a frontier API — which for many workloads (summarization, extraction, internal chat, RAG over your own docs) is smaller than the price gap implies. See our /cost calculator for per-model comparisons across the whole library.
Quality per dollar, not just dollars
The trap in every cost comparison is holding quality constant when it isn't. A local Qwen 3 32B is not GPT-4o. But the right question is whether it clears your quality bar. If it does, you're comparing $0.40/M tokens against $4.75/M blended — and local wins by 10×. If it doesn't, no electricity math rescues a model that can't do the job; pay for the API. Run the model first (rent an hour — see how to use RunPod — or check what your GPU can run) before buying anything.