Speed leaderboards are everywhere. This one ranks qwen3-8b at Q4_K_M by how many tokens each machine produces per watt it draws. For a machine that lives on your desk, watts are the constraint: they set the heat, the heat sets the fan speed, the fan speed sets the noise, and the total decides whether your power supply has headroom. A card 20% faster that draws twice the power is not a better desk machine — it is a worse one that finishes sooner.
What this ranking rests on: 1 measured of 55 rows. 10 rows are modelled on an architecture whose efficiency constant was fitted from real runs; 44 rest on a constant nobody has calibrated. Every row below says which it is.
| Hardware | Speed | Power | Tokens/watt | Evidence |
|---|---|---|---|---|
| 1. Apple M5 Max | 51 tok/s | 35 W SoC package | 1.50 | ESTIMATED · uncalibrated |
| 2. Apple M4 Max | 46 tok/s | 35 W SoC package | 1.30 | ESTIMATED · uncalibrated |
| 3. Apple M1 Max | 35 tok/s | 30 W SoC package | 1.20 | ESTIMATED · uncalibrated |
| 4. Apple M1 Ultra | 64 tok/s | 60 W SoC package | 1.10 | ESTIMATED · uncalibrated |
| 5. Apple M2 Ultra | 64 tok/s | 60 W SoC package | 1.10 | ESTIMATED · uncalibrated |
| 6. Apple M3 Ultra | 65 tok/s | 60 W SoC package | 1.10 | ESTIMATED · uncalibrated |
| 7. Apple M3 Max | 35 tok/s | 35 W SoC package | 1.00 | ESTIMATED · uncalibrated |
| 8. Apple M5 Pro | 28 tok/s | 30 W SoC package | 0.93 | ESTIMATED · uncalibrated |
| 9. Apple M2 Max | 35 tok/s | 40 W SoC package | 0.88 | ESTIMATED · uncalibrated |
| 10. Apple M4 Pro | 25 tok/s | 30 W SoC package | 0.83 | ESTIMATED · uncalibrated |
| 11. Apple M5 | 14 tok/s | 23 W SoC package | 0.63 | ESTIMATED · uncalibrated |
| 12. Apple M1 Pro | 19 tok/s | 30 W SoC package | 0.62 | ESTIMATED · uncalibrated |
| 13. Apple M2 Pro | 19 tok/s | 30 W SoC package | 0.62 | ESTIMATED · uncalibrated |
| 14. Apple M4 | 11 tok/s | 22 W SoC package | 0.52 | ESTIMATED · uncalibrated |
| 15. NVIDIA H100 80GB | 181 tok/s | 350 W board | 0.52 | ESTIMATED · uncalibrated |
| 16. NVIDIA A100 80GB | 146 tok/s | 300 W board | 0.49 | ESTIMATED · provisional fit (n=1) |
| 17. Apple M2 | 10 tok/s | 20 W SoC package | 0.48 | ESTIMATED · uncalibrated |
| 18. Apple M3 Pro | 14 tok/s | 30 W SoC package | 0.48 | ESTIMATED · uncalibrated |
| 19. Apple M3 | 10 tok/s | 22 W SoC package | 0.44 | ESTIMATED · uncalibrated |
| 20. NVIDIA GeForce RTX 5060 | 59 tok/s | 145 W board | 0.41 | ESTIMATED · uncalibrated |
On the SoC marker: unified-memory systems (Apple Silicon, DGX Spark, Strix Halo) report whole-package power — CPU, GPU and memory controller on one die — where a discrete card reports board power for the card alone. Neither figure is the wall draw of the complete machine. They are not the same measurement, so read the ranking as a real efficiency difference that is partly a difference in what is being counted.
How this is calculated. tokens_per_watt = decode_tok_s / watts. Watts resolve in one order and one only: a measured draw with its method where a run reports one, otherwise the hardware's spec power — board power for a discrete card, SoC package power for a unified-memory system — otherwise nothing at all. We never estimate a power figure for hardware that lacks one.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Also ranked: cost per 1M generated tokens, performance vs cost vs power vs VRAM, or the full filterable database.