VRAM calculator
How much GPU memory a model needs: quantized weights, KV cache at your context length, runtime overhead and agent scaffolding, at every quantization from Q2 to FP16.
How the estimate works
Weights
Parameters × bits per weight ÷ 8. At Q4_K_M (4.83 bits) a 32B model is about 19 GB before anything else loads.
KV cache
Grows linearly with context and with every concurrent user. Quantizing it to Q8 halves it with little quality loss; Q4 halves it again.
Overhead
About 0.8 GB for the runtime, CUDA context and buffers. On Apple Silicon only about 75% of unified memory is usable by the GPU by default.
Agent scaffolding
Tool schemas, system prompts and scratchpads add a fraction on top. Leave it off for plain chat.
Popular models
Each links to the model page, with per-quant memory and speed.
BitNet b1.58 3BMicrosoftDeepSeek R1 Distill Llama 8BDeepSeekDeepSeek R1 Distill Qwen 32BDeepSeekDeepSeek R1 Distill Qwen 14BDeepSeekDeepSeek R1 (671B)DeepSeekLlama 3.3 70B InstructMetaLlama 3.1 8B InstructMetaNemotron 70B InstructNVIDIACommand R (35B)CohereCommand R+ (104B)CoherePhi-4 (14B)MicrosoftPhi 3.5 MiniMicrosoftQwen 2.5 Coder 32BAlibaba CloudQwen 2.5 14B InstructAlibaba CloudQwen 2.5 7B InstructAlibaba CloudQwen 2.5 72B InstructAlibaba CloudGemma 2 9B ITGoogleMistral Small 3 (24B)Mistral AIMistral NeMo 12BMistral AIYi 1.5 34B Chat01.AIYi 1.5 9B Chat01.AIGranite 3.0 8B InstructIBMLlama 4 Scout 17BMetaLlama 4 Maverick 17BMeta
Keep going
Tools