What it costs to generate a million tokens of qwen3-8b two ways: the electricity your machine burns doing it, and the hourly rate to rent the same class of card. Local figures use $0.15/kWh — set your own rate on the interactive page.
The honest framing, because the first number is startling: local electricity cost per token is remarkably low. It is also not the whole cost — it excludes the machine itself, amortisation, idle draw, power-supply inefficiency and every other component in the box. The cost calculator answers the question this page does not: when does buying the hardware pay for itself?
What this ranking rests on: 1 measured of 55 rows. 10 rows are modelled on an architecture whose efficiency constant was fitted from real runs; 44 rest on a constant nobody has calibrated. Every row below says which it is.
| Hardware | Speed | Power | Local $/1M at $0.15/kWh | Cloud $/1M | Evidence |
|---|---|---|---|---|---|
| 1. Apple M5 Max | 51 tok/s | 35 W SoC package | $0.028 | — | ESTIMATED · uncalibrated |
| 2. Apple M4 Max | 46 tok/s | 35 W SoC package | $0.031 | — | ESTIMATED · uncalibrated |
| 3. Apple M1 Max | 35 tok/s | 30 W SoC package | $0.035 | — | ESTIMATED · uncalibrated |
| 4. Apple M3 Ultra | 65 tok/s | 60 W SoC package | $0.038 | — | ESTIMATED · uncalibrated |
| 5. Apple M1 Ultra | 64 tok/s | 60 W SoC package | $0.039 | — | ESTIMATED · uncalibrated |
| 6. Apple M2 Ultra | 64 tok/s | 60 W SoC package | $0.039 | — | ESTIMATED · uncalibrated |
| 7. Apple M3 Max | 35 tok/s | 35 W SoC package | $0.041 | — | ESTIMATED · uncalibrated |
| 8. Apple M5 Pro | 28 tok/s | 30 W SoC package | $0.045 | — | ESTIMATED · uncalibrated |
| 9. Apple M2 Max | 35 tok/s | 40 W SoC package | $0.047 | — | ESTIMATED · uncalibrated |
| 10. Apple M4 Pro | 25 tok/s | 30 W SoC package | $0.050 | — | ESTIMATED · uncalibrated |
| 11. Apple M5 | 14 tok/s | 23 W SoC package | $0.066 | — | ESTIMATED · uncalibrated |
| 12. Apple M1 Pro | 19 tok/s | 30 W SoC package | $0.067 | — | ESTIMATED · uncalibrated |
| 13. Apple M2 Pro | 19 tok/s | 30 W SoC package | $0.067 | — | ESTIMATED · uncalibrated |
| 14. Apple M4 | 11 tok/s | 22 W SoC package | $0.080 | — | ESTIMATED · uncalibrated |
| 15. NVIDIA H100 80GB | 181 tok/s | 350 W board | $0.080 | $4.43 checked 2026-07 | ESTIMATED · uncalibrated |
| 16. NVIDIA A100 80GB | 146 tok/s | 300 W board | $0.086 | $2.65 checked 2026-07 | ESTIMATED · provisional fit (n=1) |
| 17. Apple M3 Pro | 14 tok/s | 30 W SoC package | $0.087 | — | ESTIMATED · uncalibrated |
| 18. Apple M2 | 10 tok/s | 20 W SoC package | $0.087 | — | ESTIMATED · uncalibrated |
| 19. Apple M3 | 10 tok/s | 22 W SoC package | $0.096 | — | ESTIMATED · uncalibrated |
| 20. NVIDIA GeForce RTX 5060 | 59 tok/s | 145 W board | $0.102 | — | ESTIMATED · uncalibrated |
On the SoC marker: unified-memory systems (Apple Silicon, DGX Spark, Strix Halo) report whole-package power — CPU, GPU and memory controller on one die — where a discrete card reports board power for the card alone. Neither figure is the wall draw of the complete machine. They are not the same measurement, so read the ranking as a real efficiency difference that is partly a difference in what is being counted.
Reference point. A mainstream hosted API charges materially more per million output tokens than any figure above — see the dated per-provider list on the cost calculator, last checked 2026-07-30. That comparison is the one every reader is making anyway, so it is better cited than guessed. It is also not like-for-like: a hosted flagship is a larger and more capable model than anything in this table.
How this is calculated. local $/1M = (watts / 1000) x kwh_rate x 1e6 / (tok_s x 3600) and cloud $/1M = hourly_rate x 1e6 / (tok_s x 3600). Worked checks: 450 W at $0.30/kWh and 100 tok/s is $0.375 per 1M; $0.34/hr at 100 tok/s is $0.94 per 1M. Cloud figures assume the rented card generates for the whole hour — idle time between requests is billed and is not in the number.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Also ranked: tokens per watt, performance vs cost vs power vs VRAM, or the full filterable database.