Best LLMs for 64 GB VRAM

作者: Jakub Rusinowski · 最后更新: 2026年4月16日

These are the strongest local models that fit entirely in 64 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.

GPUs at This Tier

Ranked Models

Qwen 2.5 Family — Qwen 2.5 72B InstructQ4_K_M · 43.47 GB · ~4 tok/s on Apple M5 Pro
Qwen 2.5 Family — Qwen 2.5 Coder 32BQ4_K_M · 19.32 GB · ~8 tok/s on Apple M5 Pro
Llama 3.3 — Llama 3.3 70B InstructQ2_K_XS (Tight) · 20.212500000000002 GB · ~8 tok/s on Apple M5 Pro
Qwen 3 — Qwen 3 32BQ4_K_M · 19.802999999999997 GB · ~8 tok/s on Apple M5 Pro
DeepSeek R1 — DeepSeek R1 Distill Qwen 32BQ4_K_M · 19.32 GB · ~8 tok/s on Apple M5 Pro
Qwen 2.5 Family — Qwen 2.5 14B InstructQ4_K_M · 8.4525 GB · ~17 tok/s on Apple M5 Pro
Nemotron 70B — Nemotron 70B InstructQ4_K_M · 42.62475 GB · ~4 tok/s on Apple M5 Pro
Qwen 3 — Qwen 3 14BQ4_K_M · 8.935500000000001 GB · ~17 tok/s on Apple M5 Pro
Codestral — Codestral 22BQ4_K_M · 13.40325 GB · ~11 tok/s on Apple M5 Pro
Qwen 3.6 — Qwen 3.6 35B-A3BQ4_K_M · 21.13125 GB · ~52 tok/s on Apple M5 Pro
DeepSeek R1 — DeepSeek R1 Distill Qwen 14BQ4_K_M · 8.4525 GB · ~17 tok/s on Apple M5 Pro
Phi-4 Family — Phi-4 (14B)Q4_K_M · 8.4525 GB · ~17 tok/s on Apple M5 Pro
Gemma 4 — Gemma 4 31BQ4_K_M · 18.71625 GB · ~8 tok/s on Apple M5 Pro
Qwen 2.5 VL — Qwen 2.5 VL 72B InstructQ4_K_M · 44.315250000000006 GB · ~4 tok/s on Apple M5 Pro
Qwen3-Coder — Qwen3-Coder 80B-A3B (MoE)Q4_K_M · 48.3 GB

FAQ

What LLMs run well with 64 GB VRAM?

Qwen 2.5 Family, Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1 all fit in 64 GB VRAM.

Which GPUs have 64 GB VRAM?

Apple M5 Pro, Apple M1 Max, NVIDIA A100 80GB (PCIe), NVIDIA H100 80GB (PCIe).

Can-I-Run Pages Near 64 GB

Adjacent VRAM Tiers

No discrete GPU?

Buying Guide

← All VRAM Tiers | Check Your Hardware