Cost per 1M generated tokens

What it costs to generate a million tokens of qwen3-8b two ways: the electricity your machine burns doing it, and the hourly rate to rent the same class of card. Local figures use $0.15/kWh — set your own rate on the interactive page.

The honest framing, because the first number is startling: local electricity cost per token is remarkably low. It is also not the whole cost — it excludes the machine itself, amortisation, idle draw, power-supply inefficiency and every other component in the box. The cost calculator answers the question this page does not: when does buying the hardware pay for itself?

What this ranking rests on: 1 measured of 55 rows. 10 rows are modelled on an architecture whose efficiency constant was fitted from real runs; 44 rest on a constant nobody has calibrated. Every row below says which it is.

HardwareSpeedPowerLocal $/1M at $0.15/kWhCloud $/1MEvidence
1. Apple M5 Max51 tok/s35 W SoC package$0.028ESTIMATED · uncalibrated
2. Apple M4 Max46 tok/s35 W SoC package$0.031ESTIMATED · uncalibrated
3. Apple M1 Max35 tok/s30 W SoC package$0.035ESTIMATED · uncalibrated
4. Apple M3 Ultra65 tok/s60 W SoC package$0.038ESTIMATED · uncalibrated
5. Apple M1 Ultra64 tok/s60 W SoC package$0.039ESTIMATED · uncalibrated
6. Apple M2 Ultra64 tok/s60 W SoC package$0.039ESTIMATED · uncalibrated
7. Apple M3 Max35 tok/s35 W SoC package$0.041ESTIMATED · uncalibrated
8. Apple M5 Pro28 tok/s30 W SoC package$0.045ESTIMATED · uncalibrated
9. Apple M2 Max35 tok/s40 W SoC package$0.047ESTIMATED · uncalibrated
10. Apple M4 Pro25 tok/s30 W SoC package$0.050ESTIMATED · uncalibrated
11. Apple M514 tok/s23 W SoC package$0.066ESTIMATED · uncalibrated
12. Apple M1 Pro19 tok/s30 W SoC package$0.067ESTIMATED · uncalibrated
13. Apple M2 Pro19 tok/s30 W SoC package$0.067ESTIMATED · uncalibrated
14. Apple M411 tok/s22 W SoC package$0.080ESTIMATED · uncalibrated
15. NVIDIA H100 80GB181 tok/s350 W board$0.080$4.43 checked 2026-07ESTIMATED · uncalibrated
16. NVIDIA A100 80GB146 tok/s300 W board$0.086$2.65 checked 2026-07ESTIMATED · provisional fit (n=1)
17. Apple M3 Pro14 tok/s30 W SoC package$0.087ESTIMATED · uncalibrated
18. Apple M210 tok/s20 W SoC package$0.087ESTIMATED · uncalibrated
19. Apple M310 tok/s22 W SoC package$0.096ESTIMATED · uncalibrated
20. NVIDIA GeForce RTX 506059 tok/s145 W board$0.102ESTIMATED · uncalibrated

On the SoC marker: unified-memory systems (Apple Silicon, DGX Spark, Strix Halo) report whole-package power — CPU, GPU and memory controller on one die — where a discrete card reports board power for the card alone. Neither figure is the wall draw of the complete machine. They are not the same measurement, so read the ranking as a real efficiency difference that is partly a difference in what is being counted.

Reference point. A mainstream hosted API charges materially more per million output tokens than any figure above — see the dated per-provider list on the cost calculator, last checked 2026-07-30. That comparison is the one every reader is making anyway, so it is better cited than guessed. It is also not like-for-like: a hosted flagship is a larger and more capable model than anything in this table.

How this is calculated. local $/1M = (watts / 1000) x kwh_rate x 1e6 / (tok_s x 3600) and cloud $/1M = hourly_rate x 1e6 / (tok_s x 3600). Worked checks: 450 W at $0.30/kWh and 100 tok/s is $0.375 per 1M; $0.34/hr at 100 tok/s is $0.94 per 1M. Cloud figures assume the rented card generates for the whole hour — idle time between requests is billed and is not in the number.

Buy This HardwareApple MacBook Pro M5 Max — 128 GB VRAM · 35 W board powerDeploy in the Cloud NowRTX 4090 on RunPod — from $0.34/hr · rate checked 2026-07

or compare on Vast.ai from $0.35/hr (typical low · varies)

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Also ranked: tokens per watt, performance vs cost vs power vs VRAM, or the full filterable database.