LLM VRAM calculator
How much GPU memory a model needs: quantized weights, KV cache at your context length, runtime overhead and agent scaffolding, at every quantization from Q2 to FP16.
How the estimate works
Weights
Parameters × bits per weight ÷ 8. At Q4_K_M (4.83 bits) a 32B model is about 19 GB before anything else loads.
KV cache
Grows linearly with context and with every concurrent user. Quantizing it to Q8 halves it with little quality loss; Q4 halves it again.
Overhead
About 0.8 GB for the runtime, CUDA context and buffers. On Apple Silicon only about 75% of unified memory is usable by the GPU by default.
Agent scaffolding
Tool schemas, system prompts and scratchpads add a fraction on top. Leave it off for plain chat.
Popular models
Each links to the model page, with per-quant memory and speed.
VRAM calculator FAQ
How much VRAM do 7B, 14B, 32B and 70B models need at Q4_K_M?
Weights plus runtime overhead, before any context: 7B (Qwen 2.5 7B): about 5.4 GB; 14B (Qwen3 14B): about 9.7 GB; 32B (Qwen3 32B): about 20.6 GB; 70B (Llama 3.3 70B): about 43.1 GB. Add the KV cache for your context length on top. Mixture-of-experts models need memory for ALL their parameters, not just the active ones.
Why does context length change how much VRAM I need?
Every token in the context keeps keys and values in the KV cache, so it grows linearly with context. For Llama 3.1 8B the cache is 1.1 GB at 8K tokens and 4.3 GB at 32K, against 4.8 GB of Q4_K_M weights.
What is the difference between model weights and the KV cache?
Weights are the model itself: a fixed size set by the parameter count and the quantization. The KV cache is working memory for the conversation: it starts near zero and grows with every token of context and every concurrent user.
How much memory does the runtime itself use?
About 0.8 GB for the CUDA or Metal context and buffers. On Apple Silicon only part of unified memory is available to the GPU by default, so a Mac needs headroom beyond the model size.
Is CPU offload worth it when a model does not fit?
It lets a model run, but every layer left in system RAM is read over a much slower bus on every token, so speed drops sharply. A smaller quantization or a smaller model that fits entirely in VRAM is usually faster. Mixture-of-experts models offload best, because only a few experts are read per token.
Keep going
Tools