Compare Local Models
Pick up to 4 models and see how they perform on your hardware at every quantization level.
Select models above to compare
Pick 2–4 models using the dropdowns, then choose your hardware.
VRAM · Disk · tok/s · Fit · Ollama pull command
How we estimate this
VRAM = quantized weights + KV cache at your context + runtime overhead. Weights are sized by TOTAL parameters: every expert of a Mixture-of-Experts model must sit in memory.
Speed is modelled from memory bandwidth using the ACTIVE parameters (only the experts read per token), with an efficiency fitted per GPU architecture. Estimates, not benchmarks.
More ways to choose a model
Models section
Tools