Compare the true cost of running any of 154+ local LLMs (across 55 GPUs) vs hosted cloud APIs (ChatGPT, Claude, Gemini, and more). Most power users break even in 3–8 months. Cloud prices as of 2026-07-30.
| GPT-5.5 (OpenAI) | $5.00 / 1M input tokens + $30.00 / 1M output |
| GPT-5.4 (OpenAI) | $2.50 / 1M input tokens + $15.00 / 1M output |
| Claude Opus 5 (Anthropic) | $5.00 / 1M input tokens + $25.00 / 1M output |
| Claude Sonnet 5 (Anthropic) | $2.00 / 1M input tokens + $10.00 / 1M output |
| Claude Haiku 4.5 (Anthropic) | $1.00 / 1M input tokens + $5.00 / 1M output |
| Gemini 3.1 Pro (Google) | $2.00 / 1M input tokens + $12.00 / 1M output |
| Llama 3.1 405B (Meta (Together.ai)) | $1.79 / 1M input tokens + $1.79 / 1M output |
| Llama 3.1 8B (Meta (Together.ai)) | $0.18 / 1M input tokens + $0.18 / 1M output |
| Mistral Large 2 (Mistral AI) | $2.00 / 1M input tokens + $6.00 / 1M output |
| Mistral Small 3 (Mistral AI) | $0.10 / 1M input tokens + $0.30 / 1M output |
| DeepSeek V3 (DeepSeek) | $0.27 / 1M input tokens + $1.10 / 1M output |
| DeepSeek R1 (DeepSeek) | $0.55 / 1M input tokens + $2.19 / 1M output |
For local inference, cost is driven by electricity use and GPU hardware amortization — typically a small fraction of a cent per 1M tokens once the GPU is paid for. The full calculator lets you pick any model from the library, any GPU, your electricity rate, and your currency (USD/EUR).