Tokens per watt, ranked

Speed leaderboards are everywhere. This one ranks qwen3-8b at Q4_K_M by how many tokens each machine produces per watt it draws. For a machine that lives on your desk, watts are the constraint: they set the heat, the heat sets the fan speed, the fan speed sets the noise, and the total decides whether your power supply has headroom. A card 20% faster that draws twice the power is not a better desk machine — it is a worse one that finishes sooner.

What this ranking rests on: 1 measured of 55 rows. 10 rows are modelled on an architecture whose efficiency constant was fitted from real runs; 44 rest on a constant nobody has calibrated. Every row below says which it is.

HardwareSpeedPowerTokens/wattEvidence
1. Apple M5 Max51 tok/s35 W SoC package1.50ESTIMATED · uncalibrated
2. Apple M4 Max46 tok/s35 W SoC package1.30ESTIMATED · uncalibrated
3. Apple M1 Max35 tok/s30 W SoC package1.20ESTIMATED · uncalibrated
4. Apple M1 Ultra64 tok/s60 W SoC package1.10ESTIMATED · uncalibrated
5. Apple M2 Ultra64 tok/s60 W SoC package1.10ESTIMATED · uncalibrated
6. Apple M3 Ultra65 tok/s60 W SoC package1.10ESTIMATED · uncalibrated
7. Apple M3 Max35 tok/s35 W SoC package1.00ESTIMATED · uncalibrated
8. Apple M5 Pro28 tok/s30 W SoC package0.93ESTIMATED · uncalibrated
9. Apple M2 Max35 tok/s40 W SoC package0.88ESTIMATED · uncalibrated
10. Apple M4 Pro25 tok/s30 W SoC package0.83ESTIMATED · uncalibrated
11. Apple M514 tok/s23 W SoC package0.63ESTIMATED · uncalibrated
12. Apple M1 Pro19 tok/s30 W SoC package0.62ESTIMATED · uncalibrated
13. Apple M2 Pro19 tok/s30 W SoC package0.62ESTIMATED · uncalibrated
14. Apple M411 tok/s22 W SoC package0.52ESTIMATED · uncalibrated
15. NVIDIA H100 80GB181 tok/s350 W board0.52ESTIMATED · uncalibrated
16. NVIDIA A100 80GB146 tok/s300 W board0.49ESTIMATED · provisional fit (n=1)
17. Apple M210 tok/s20 W SoC package0.48ESTIMATED · uncalibrated
18. Apple M3 Pro14 tok/s30 W SoC package0.48ESTIMATED · uncalibrated
19. Apple M310 tok/s22 W SoC package0.44ESTIMATED · uncalibrated
20. NVIDIA GeForce RTX 506059 tok/s145 W board0.41ESTIMATED · uncalibrated

On the SoC marker: unified-memory systems (Apple Silicon, DGX Spark, Strix Halo) report whole-package power — CPU, GPU and memory controller on one die — where a discrete card reports board power for the card alone. Neither figure is the wall draw of the complete machine. They are not the same measurement, so read the ranking as a real efficiency difference that is partly a difference in what is being counted.

How this is calculated. tokens_per_watt = decode_tok_s / watts. Watts resolve in one order and one only: a measured draw with its method where a run reports one, otherwise the hardware's spec power — board power for a discrete card, SoC package power for a unified-memory system — otherwise nothing at all. We never estimate a power figure for hardware that lacks one.

Buy This HardwareApple MacBook Pro M5 Max — 128 GB VRAM · 35 W board powerDeploy in the Cloud NowRTX 4090 on RunPod — from $0.34/hr · rate checked 2026-07

or compare on Vast.ai from $0.35/hr (typical low · varies)

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Also ranked: cost per 1M generated tokens, performance vs cost vs power vs VRAM, or the full filterable database.