Four axes, weighted by what you value

Speed, running cost, efficiency and memory for qwen3-8b, each normalised across the hardware shown, so you can see where a machine is strong and where it gives something up.

There is no score here until you make one. A single blended number with weights we picked would look objective while hiding an opinion you never saw. The default is the four axes; the ranking appears once you set the sliders, and your weights go in the URL so a link you share shows what you valued. Three starting points are offered — "I run this all day", "I want the biggest model I can", "I want it fast and I don't care" — and each is labelled as the opinion it is.

What this ranking rests on: 1 measured of 55 rows. 10 rows are modelled on an architecture whose efficiency constant was fitted from real runs; 44 rest on a constant nobody has calibrated. Every row below says which it is.

How this is calculated. axis = (value - min) / (max - min) across the hardware shown, with cost inverted first so higher is better on all four. An axis every row agrees on scores 0 rather than 1: it carries no information, and scoring it full would let a flat axis dominate your composite. The composite is SUM(axis x your_weight) / SUM(your_weights), and it does not exist until you supply weights.

Buy This HardwareNVIDIA GeForce RTX 4090 24GB — 24 GB VRAM · 450 W board powerDeploy in the Cloud NowRTX 4090 on RunPod — from $0.34/hr · rate checked 2026-07

or compare on Vast.ai from $0.35/hr (typical low · varies)

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Also ranked: tokens per watt, cost per 1M generated tokens, or the full filterable database.