Free tool · per-quant memory estimates

VRAM calculator

How much GPU memory a model needs: quantized weights, KV cache at your context length, runtime overhead and agent scaffolding, at every quantization from Q2 to FP16.

How the estimate works

Weights

Parameters × bits per weight ÷ 8. At Q4_K_M (4.83 bits) a 32B model is about 19 GB before anything else loads.

KV cache

Grows linearly with context and with every concurrent user. Quantizing it to Q8 halves it with little quality loss; Q4 halves it again.

Overhead

About 0.8 GB for the runtime, CUDA context and buffers. On Apple Silicon only about 75% of unified memory is usable by the GPU by default.

Agent scaffolding

Tool schemas, system prompts and scratchpads add a fraction on top. Leave it off for plain chat.

Each links to the model page, with per-quant memory and speed.