Best LLMs for 32 GB VRAM
作者: Jakub Rusinowski · 最后更新: 2026年4月16日
These are the strongest local models that fit entirely in 32 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.
GPUs at This Tier
Ranked Models
| Qwen 2.5 Family — Qwen 2.5 Coder 32B | Q4_K_M · 19.32 GB · ~17 tok/s on Intel Arc Pro B65 |
| Llama 3.3 — Llama 3.3 70B Instruct | Q2_K_XS (Tight) · 20.212500000000002 GB · ~16 tok/s on Intel Arc Pro B65 |
| Qwen 3 — Qwen 3 32B | Q4_K_M · 19.802999999999997 GB · ~16 tok/s on Intel Arc Pro B65 |
| DeepSeek R1 — DeepSeek R1 Distill Qwen 32B | Q4_K_M · 19.32 GB · ~17 tok/s on Intel Arc Pro B65 |
| Qwen 2.5 Family — Qwen 2.5 14B Instruct | Q4_K_M · 8.4525 GB · ~34 tok/s on Intel Arc Pro B65 |
| Qwen 3 — Qwen 3 14B | Q4_K_M · 8.935500000000001 GB · ~33 tok/s on Intel Arc Pro B65 |
| Codestral — Codestral 22B | Q4_K_M · 13.40325 GB · ~23 tok/s on Intel Arc Pro B65 |
| Qwen 3.6 — Qwen 3.6 35B-A3B | Q4_K_M · 21.13125 GB · ~87 tok/s on Intel Arc Pro B65 |
| DeepSeek R1 — DeepSeek R1 Distill Qwen 14B | Q4_K_M · 8.4525 GB · ~34 tok/s on Intel Arc Pro B65 |
| Phi-4 Family — Phi-4 (14B) | Q4_K_M · 8.4525 GB · ~34 tok/s on Intel Arc Pro B65 |
| Gemma 4 — Gemma 4 31B | Q4_K_M · 18.71625 GB · ~17 tok/s on Intel Arc Pro B65 |
| Gemma 3 — Gemma 3 27B Instruct | Q4_K_M · 16.30125 GB · ~17 tok/s on Intel Arc Pro B65 |
| GLM-4.7 / GLM-Z1 — GLM-Z1 32B (Reasoning) | Q4_K_M · 19.32 GB · ~17 tok/s on Intel Arc Pro B65 |
| Qwen 3.5 — Qwen 3.5 35B-A3B | Q4_K_M · 21.13125 GB · ~87 tok/s on Intel Arc Pro B65 |
| Mistral Small 3.1 — Mistral Small 3.1 24B | Q4_K_M · 14.248500000000002 GB · ~22 tok/s on Intel Arc Pro B65 |
购买此硬件 Apple Mac mini M4 (16GB) — 32 GB VRAM · 22 W board power立即云端部署 RunPod 上的 NVIDIA A40 — 低至 $0.44/小时 · 价格核实于 2026-08
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
FAQ
What LLMs run well with 32 GB VRAM?
Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1, Qwen 2.5 Family all fit in 32 GB VRAM.
Which GPUs have 32 GB VRAM?
Intel Arc Pro B65, AMD Ryzen AI 9 HX 370, Intel Arc Pro B70, Apple M4.
Can-I-Run Pages Near 32 GB
- DeepSeek R1 on NVIDIA GeForce RTX 5090
- DeepSeek R1 on Apple M4
- DeepSeek R1 on Apple M2 Pro
- DeepSeek R1 on Apple M1 Pro
- DeepSeek R1 on Apple M5
- Llama 3.3 on NVIDIA GeForce RTX 5090
- Llama 3.3 on Apple M4
- Llama 3.3 on Apple M2 Pro