Compare Local Models

Pick up to 4 models and see how they perform on your hardware at every quantization level.

Select models above to compare

Pick 2–4 models using the dropdowns, then choose your hardware.

VRAM · Disk · tok/s · Fit · Ollama pull command

How we estimate this

VRAM = quantized weights + KV cache at your context + runtime overhead. Weights are sized by TOTAL parameters: every expert of a Mixture-of-Experts model must sit in memory.

Speed is modelled from memory bandwidth using the ACTIVE parameters (only the experts read per token), with an efficiency fitted per GPU architecture. Estimates, not benchmarks.

The published formulas and constants →