Best Local LLM by VRAM Tier (2026): 4GB / 8GB / 16GB / 24GB+ Guide

Written by Jakub Rusinowski · Last updated July 21, 2026

Not sure which model to run? This guide cuts through the noise. We recommend the single best model per VRAM tier — the one most users should start with.

In This Guide

Not sure which model to run? This guide cuts through the noise. We recommend the single best model per VRAM tier — the one most users should start with.

4 GB VRAM (Budget GPUs / iGPUs)

Best pick: Phi-4 Mini (3.8B, Q4_K_M)

Runner-up: Gemma 3 4B (ollama run gemma3:4b) — better for creative writing, 3.5 GB.

What to expect: Capable assistant for everyday chat, coding help, and writing. Don't expect 70B-level reasoning, but for 4 GB, this is remarkable.

8 GB VRAM (RTX 4060, RX 7600, laptop GPUs)

Best pick: Llama 3.1 8B (Q4_K_M)

Runner-up for coding: Qwen 2.5 Coder 7B (ollama run qwen2.5-coder:7b) — best pure-coding 8B model.

Runner-up for reasoning: DeepSeek R1 8B Distill (ollama run deepseek-r1:8b) — near-GPT-4-level reasoning in 8B.

12 GB VRAM (RTX 4070, RTX 3060 12GB, RX 7700 XT)

Best pick: Qwen 3 8B (Q8_0) — run at higher quality since you have headroom

Or try: Gemma 3 12B Q4 (ollama run gemma3:12b) — Google's best mid-size model.

16 GB VRAM (RTX 4060 Ti 16GB, RTX 4070 Ti, RX 7800 XT)

Best pick: Qwen 3 14B (Q4_K_M)

Runner-up for coding: Qwen 2.5 Coder 14B (ollama run qwen2.5-coder:14b)

Best for long documents: Mistral Small 3.1 (ollama run mistral-small) — 128K context window.

24 GB VRAM (RTX 4090, RTX 3090, RX 7900 XTX)

Best pick: DeepSeek R1 32B (Q4_K_M)

Alternative: Qwen 3 32B (ollama run qwen3:32b) — also fits, stronger at general tasks.

32 GB+ VRAM (RTX 5090, Apple M4 Max, Professional GPUs)

Best pick: Llama 4 Scout (Q4_K_M)

For 70B models: Llama 3.3 70B Q4_K_M needs 42 GB — works on Apple M2/M4 Ultra (128–192 GB unified memory) or dual RTX 4090/5090 setups.

The Quick Reference Table

VRAMBest ModelOllama CommandSpeed
4 GBPhi-4 Miniollama run phi4-mini~40 t/s
8 GBLlama 3.1 8Bollama run llama3.1:8b~70 t/s
12 GBQwen 3 8B Q8ollama run qwen3:8b~80 t/s
16 GBQwen 3 14Bollama run qwen3:14b~85 t/s
24 GBDeepSeek R1 32Bollama run deepseek-r1:32b~50 t/s
32 GB+Llama 4 Scoutollama run llama4:scout175 t/s

→ Check your exact GPU compatibility | → Browse all models

← All Guides | Check GPU Compatibility