Best LLMs for 24 GB VRAM

作者: Jakub Rusinowski · 最后更新: 2026年4月16日

These are the strongest local models that fit entirely in 24 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.

GPUs at This Tier

Ranked Models

Qwen 2.5 Family — Qwen 2.5 Coder 32BQ4_K_M · 19.32 GB · ~34 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Llama 3.3 — Llama 3.3 70B InstructQ2_K_XS (Tight) · 20.212500000000002 GB · ~33 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Qwen 3 — Qwen 3 32BQ4_K_M · 19.802999999999997 GB · ~34 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
DeepSeek R1 — DeepSeek R1 Distill Qwen 32BQ4_K_M · 19.32 GB · ~34 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Qwen 2.5 Family — Qwen 2.5 14B InstructQ4_K_M · 8.4525 GB · ~69 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Qwen 3 — Qwen 3 14BQ4_K_M · 8.935500000000001 GB · ~67 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Codestral — Codestral 22BQ4_K_M · 13.40325 GB · ~47 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Qwen 3.6 — Qwen 3.6 35B-A3BQ4_K_M · 21.13125 GB · ~171 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
DeepSeek R1 — DeepSeek R1 Distill Qwen 14BQ4_K_M · 8.4525 GB · ~70 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Phi-4 Family — Phi-4 (14B)Q4_K_M · 8.4525 GB · ~69 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Gemma 4 — Gemma 4 31BQ4_K_M · 18.71625 GB · ~36 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Gemma 3 — Gemma 3 27B InstructQ4_K_M · 16.30125 GB · ~34 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
GLM-4.7 / GLM-Z1 — GLM-Z1 32B (Reasoning)Q4_K_M · 19.32 GB · ~35 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Qwen 3.5 — Qwen 3.5 35B-A3BQ4_K_M · 21.13125 GB · ~171 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
Mistral Small 3.1 — Mistral Small 3.1 24BQ4_K_M · 14.248500000000002 GB · ~46 tok/s on NVIDIA GeForce RTX 5090 Laptop GPU
购买此硬件 AMD Radeon RX 7900 XTX 24GB — 24 GB VRAM · 355 W board power立即云端部署 RunPod 上的 RTX 4090 — 低至 $0.34/小时 · 价格核实于 2026-07

或在 Vast.ai 比较,低至 $0.35/小时 (typical low · varies)

作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。

FAQ

What LLMs run well with 24 GB VRAM?

Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1, Qwen 2.5 Family all fit in 24 GB VRAM.

Which GPUs have 24 GB VRAM?

NVIDIA GeForce RTX 5090 Laptop GPU, AMD Radeon RX 7900 XTX, Apple M2, Apple M3.

Can-I-Run Pages Near 24 GB

Adjacent VRAM Tiers

No discrete GPU?

Buying Guide

← All VRAM Tiers | Check Your Hardware