Best LLMs for 32 GB VRAM

作者： Jakub Rusinowski · 最后更新： 2026年7月30日

These are the strongest local models that fit entirely in 32 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.

GPUs at This Tier

Ranked Models

Qwen 2.5 Family — Qwen 2.5 Coder 32B	Q4_K_M · 19.5 GB · ~6 tok/s on Apple M4
Llama 3.3 — Llama 3.3 70B Instruct	Q2_K_XS (Tight) · 26 GB · ~5 tok/s on Apple M4
Qwen 3 — Qwen 3 32B	Q4_K_M · 19.5 GB · ~6 tok/s on Apple M4
DeepSeek R1 — DeepSeek R1 Distill Qwen 32B	Q4_K_M · 19.5 GB · ~6 tok/s on Apple M4
Kimi K2.5 / K2.6 — Kimi K2.6	Q4_K_M · 19 GB · ~6 tok/s on Apple M4
Qwen 2.5 Family — Qwen 2.5 14B Instruct	Q4_K_M · 9.5 GB · ~13 tok/s on Apple M4
Kimi K2.5 / K2.6 — Kimi K2.5	Q4_K_M · 19 GB · ~6 tok/s on Apple M4
Qwen 3.5 (Legacy Listing — Unverified) — Qwen 3.5 122B-A10B (MoE)	Q4_K_M · 13.5 GB · ~108 tok/s on Apple M4
Gemma 4 (Legacy Listing — Unverified) — Gemma 4 27B ⭐	Q4_K_M · 14 GB · ~9 tok/s on Apple M4
Qwen 3 — Qwen 3 14B	Q4_K_M · 9.5 GB · ~13 tok/s on Apple M4
Qwen 3.7 — Qwen 3.7 35B-A3B	Q4_K_M · 21 GB · ~67 tok/s on Apple M4
Llama 4 — Llama 4 Maverick 17B	Q4_K_M · 24 GB · ~118 tok/s on Apple M4
Codestral — Codestral 22B	Q4_K_M · 13 GB · ~9 tok/s on Apple M4
Qwen 3.6 — Qwen 3.6 35B-A3B	Q4_K_M · 21 GB · ~67 tok/s on Apple M4
DeepSeek R1 — DeepSeek R1 Distill Qwen 14B	Q4_K_M · 9.2 GB · ~13 tok/s on Apple M4

FAQ

What LLMs run well with 32 GB VRAM?

Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1, Kimi K2.5 / K2.6 all fit in 32 GB VRAM.

Which GPUs have 32 GB VRAM?

Apple M4, Apple M5, NVIDIA GeForce RTX 5090, Apple M2 Pro.

Can-I-Run Pages Near 32 GB

Adjacent VRAM Tiers

Buying Guide

Best GPU Buyer Guide 2026

← All VRAM Tiers | Check Your Hardware