Methodology
Every number on this site is either a measurement with a source, or the output of one of the two formulas below. This page publishes both, the constants they use, and the cases where we know the model is wrong.
1. VRAM
What must be resident in memory for the model to load at all. This is an accounting calculation, not a performance model.
vram_gb = weights_gb + kv_cache_gb + overhead_gb weights_gb = totalParamsB x bits_per_weight / 8 <- TOTAL parameters kv_cache_gb = 2 x layers x kv_heads x head_dim x ctx x kv_bytes / 1e9 overhead_gb = 0.8 framework + CUDA context
Bits per weight
Effective on-disk width including block scales and metadata, which is why the 4-bit formats land above 4.0.
2. Memory-bandwidth utilization
A decode step is memory-bound: the weights that are read per token have to cross the memory bus, and no real runtime sustains 100% of a card's theoretical peak. Assuming 1.0 would overstate throughput by 25-40%. This site fits a per-architecture efficiency to published llama-bench runs rather than assuming a single figure, because sustained bandwidth is a property of the memory subsystem and not of LLMs. The fitted values are in the table below. End-to-end, after the fixed framework cost is included, the effective utilization these produce lands between roughly 0.45 and 0.85 depending on card and model size.
3. The small-model ceiling
Kernel launch, graph dispatch and sampling cost the same whether the model is 0.5B or 70B. As the weights shrink, that fixed cost comes to dominate and throughput asymptotes to 1000/t_fixed rather than rising without limit. This site models that as a measured per-backend constant (t_fixed) instead of a hardcoded ceiling, so the flattening is an emergent property of the backend and different small models still return different numbers. A hard clamp was removed precisely because it made a 1B model and a 109B model report the same value, and a test fails the build if one reappears.
4. Confidence labels
No number is published bare. Every throughput figure carries one of these:
5. Throughput, calibration and limitations
The decode formula, its fitted constants, every calibration point with its residual, and the cases where the model is known to be wrong.
t_token_ms = (W_active + KV(ctx)) / (BW × η_arch) × 1000 memory streaming
+ W_offloaded / (BW_host × η_host) × 1000 offload penalty
+ t_fixed_ms framework floor
tokens_per_sec = 1000 / t_token_ms- W_active
- Quantized bytes of the weights actually read per token. For a dense model this is the whole file. For a Mixture-of-Experts model it is only the routed experts, which is why a 109B MoE can decode as fast as a 17B dense model. Note this is the SPEED input only — the FIT check uses total weights, because every expert must be resident.
- KV(ctx)
- 2 × n_layers × n_kv_heads × head_dim × ctx × bytes_per_element, from each model's published config.json. The KV cache is re-read every token, so it adds to the streaming cost and grows linearly with context. This is the single biggest omission in a plain roofline model: an 8B Q4 loses roughly half its throughput between 4K and 65K context purely because of this term. Models using Multi-head Latent Attention (DeepSeek, Kimi) cache one compressed latent instead of K and V, which is 10-20× smaller, and are computed separately.
- η_arch
- Fraction of theoretical memory bandwidth a decode loop actually sustains, fitted per GPU architecture. Sustained bandwidth is a property of the memory subsystem, not of the model being run. Ada and Ampere sit about 0.30 apart in our fit, so a single global constant would misattribute an architectural difference to every model on the page.
- t_fixed
- Per-token latency that does not scale with model size: kernel launch, graph dispatch, sampling. Fitted per backend. This is why a 0.5B model does not run 40× faster than a 20B one on the same card. It also replaces the hardcoded 400 tok/s ceiling this page used to carry: throughput now approaches 1000/t_fixed as weights shrink, so the ceiling is a consequence of the backend rather than a number someone chose.
- W_offloaded / BW_host
- Weights that do not fit in VRAM, read from system RAM by the CPU. llama.cpp does not stream non-resident weights to the GPU across PCIe — it computes those layers on the CPU, reading them from system DRAM, and only the activation tensor crosses the link. So the bottleneck is host memory bandwidth, not link bandwidth. On unified-memory systems (Apple, DGX Spark, Strix Halo) there is no second tier at all and this term is zero.
Efficiency by GPU architecture (η)
| NVIDIA Blackwell | 0.850 | assumed | assumption — no calibration data for this architecture |
|---|---|---|---|
| NVIDIA Ada Lovelace | 0.927 | fitted | solved from 12 measurements |
| NVIDIA Ampere | 0.622 | fitted | solved from 1 measurement |
| NVIDIA Hopper | 0.850 | assumed | assumption — no calibration data for this architecture |
| AMD RDNA 4 | 0.700 | assumed | assumption — no calibration data for this architecture |
| AMD RDNA 3 / 3.5 | 0.650 | assumed | assumption — no calibration data for this architecture |
| Intel Xe2 (Battlemage) | 0.600 | assumed | assumption — no calibration data for this architecture |
| Apple Silicon | 0.548 | fitted | solved from 1 measurement |
| Custom hardware (unclassified) | 0.750 | assumed | assumption — no calibration data for this architecture |
| CPU (system RAM) | 0.350 | assumed | assumption — no calibration data for this architecture |
Fixed per-token cost by backend
| CUDA | 2.25 ms | fitted | solved from 13 measurements |
|---|---|---|---|
| ROCM | 3.50 ms | assumed | assumption — no calibration data for this backend |
| METAL | 3.00 ms | assumed | not identifiable from 1 measurement — efficiency and fixed cost trade off exactly on a single point |
| VULKAN | 4.00 ms | assumed | assumption — no calibration data for this backend |
| CPU | 5.00 ms | assumed | assumption — no calibration data for this backend |
Calibration set and residuals
| Measurement | Measured | Predicted | Error | Source |
|---|---|---|---|---|
| rtx-4090 / qwen3-30b-a3b / Q4_K_XL / 4096 ctx | 195.78 | 218.15 | +11.43% | source |
| rtx-4090 / qwen3-30b-a3b / Q4_K_XL / 57344 ctx | 74.64 | 98.18 | +31.53% | source |
| rtx-4090 / qwen3-8b / Q4_K_M / 4096 ctx | 141.3 | 122.09 | -13.59% | source |
| rtx-4090 / qwen3-8b / Q4_K_M / 16384 ctx | 108 | 98.72 | -8.59% | source |
| rtx-4090 / qwen3-8b / Q4_K_M / 32768 ctx | 82.3 | 78.65 | -4.44% | source |
| rtx-4090 / qwen3-8b / Q4_K_M / 65536 ctx | 56.1 | 55.91 | -0.34% | source |
| rtx-4090 / qwen3-8b / Q4_K_M / 131072 ctx | 33.8 | 35.43 | +4.81% | source |
| rtx-4090 / llama-3.1-8b / Q4_K_XL / 4096 ctxinferred identity | 131 | 124.33 | -5.09% | source |
| rtx-4090 / llama-3.1-8b / Q4_K_XL / 65536 ctxinferred identity | 53.07 | 60.02 | +13.09% | source |
| rtx-4090 / qwen3-14b / Q4_K_XL / 4096 ctxinferred identity | 82.82 | 79.58 | -3.92% | source |
| rtx-4090 / qwen3-14b / Q4_K_XL / 65536 ctxinferred identity | 38.73 | 42.85 | +10.63% | source |
| rtx-4090 / qwen-2.5-7b / Q4_K_M / 4096 ctxinferred identity | 135 | 134.99 | -0.01% | source |
| rtx-3090 / qwen-2.5-7b / Q4_K_M / 4096 ctxinferred identity | 95 | 95 | -0.01% | source |
| apple-m3-max / qwen-2.5-7b / Q4_K_M / 4096 ctxinferred identity | 40 | 40 | 0% | source |
Measurements we chose not to fit, and why
- Llama 4 Scout, Unsloth 1.78-bit dynamic GGUF, RTX 4090, ~20 tok/s — The only published offload data point available, but under-specified: no stated context length, host RAM speed, or GPU/CPU layer split, each of which moves the result by more than the figure itself. Fitting the offload term to it would encode three assumptions as one measurement. The offload term is therefore left unfitted and offloaded estimates report calibrated:false. source
- mustafa.net 70B Q2 rows (4090 18 tok/s, 3090 10 tok/s) — A 70B at Q2 does not fit in 24 GB, so these rows exercise the offload path with an unstated CPU/GPU split. Same problem as the Scout row. source
- M3 Max 64GB, 70B, 5 tok/s — Quantization not stated, and at 5 tok/s on a 400 GB/s part the model is almost certainly spilling past the Apple usable-memory fraction, making the residency assumption unverifiable. source
Known limitations
Small-active-parameter MoE at long context is over-predicted
Qwen3-30B-A3B at 57K context on an RTX 4090 predicts about 98 tok/s against 74.6 measured, a 31% error — the single worst point in the calibration set and the only one above 25%. The same model at 4K is within 11%, and the dense Qwen3-8B series holds within 5% out to 128K, so this is not a general long-context failure. Likely causes are expert-routed reads sustaining less bandwidth than sequential dense reads, and the card running near capacity. With only two MoE points in the set, adding a term to correct it would be fitting a parameter to one observation.
Apple Silicon rests on a single measurement, and two sources disagree
An M3 Max at 400 GB/s measured at 40 tok/s and an M4 Max at 546 GB/s published at 75-90 tok/s imply a 1.9-2.25× speedup from a 1.37× bandwidth increase. No single Apple efficiency constant satisfies both. We fit to the M3 Max point, which means the M4 Max held-out check fails — our prediction of about 52 tok/s falls below the published range. Either one source is wrong or Apple generations need separate constants, and one measurement cannot settle it.
The offload path is not calibrated at all
Every published llama-bench figure we could source runs fully resident at -ngl 99, so no measurement constrains the offload term. Its constants (89.6 GB/s host DRAM at 35% efficiency) are assumptions, and any row whose weights spill to system RAM is reported as uncalibrated with a wide band. Treat offloaded figures as order-of-magnitude only.
Most architectures have no measurements behind them
Blackwell, Hopper, RDNA 3, RDNA 4 and Intel Xe2 have no calibration rows. Their efficiency constants are seeded assumptions, flagged as unfitted throughout the site, and estimates on those cards carry a deliberately wider band. Bandwidth ratios between cards of the same architecture should still be reliable; absolute figures on an unfitted architecture should not be trusted to better than about a third.
Roughly half the model library has no verified architecture
90 of 222 library variants have a config.json we could verify, and their KV figures are exact. The rest are unreleased, speculative, or explicitly unverified listings, and their KV cache is inferred from parameter count by a documented heuristic. Those rows report their KV provenance as estimated rather than measured.
The residency rule is slightly too conservative at very long context
Qwen3-8B Q4_K_M at 128K context on a 24 GB RTX 4090 computes to 25.1 GB required — weights plus a 19.3 GB KV cache plus framework overhead — so we predict it will not load. Hardware Corner measured it running at 33.8 tok/s. Flash attention and a smaller real KV footprint close a roughly 4% gap our arithmetic does not model. Where a measurement exists past a predicted memory cliff, the page shows it and says the rule is wrong rather than hiding the point; trust the measurement.
Prompt processing is calibrated on one architecture only
Prefill (and therefore time-to-first-token) is modelled as a multiple of decode speed: prefill_tok_s = decode_tok_s x R_arch x (active/total params). R is solved from three published prompt-processing figures, all on an RTX 4090, giving 72.8 with a 5.4% mean error. Every other architecture inherits that value and reports uncalibrated, so TTFT off Ada is a rough guide rather than an estimate. Apple Silicon is the weakest case: it has far less compute per unit of bandwidth than a discrete NVIDIA part, so its true prefill advantage is almost certainly smaller than the figure shown.
Four calibration rows do not name their exact checkpoint
Some sources report only a size class ("7B", "14B"). Those rows are fitted against a same-geometry stand-in and tagged as inferred identity, and the fit report gives MAPE both with and without them so the headline figure can be read either way.
Measured rows are compiled from published benchmarks by Hardware Corner, Mustafa.net and credited per row. The dataset is published under CC BY 4.0 at /measured-benchmarks.json.