Best LLMs for 48 GB VRAM

作者: Jakub Rusinowski · 最后更新: 2026年4月16日

These are the strongest local models that fit entirely in 48 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.

GPUs at This Tier

Ranked Models

Qwen 2.5 Family — Qwen 2.5 72B InstructQ4_K_M · 43.47 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 2.5 Family — Qwen 2.5 Coder 32BQ4_K_M · 19.32 GB · ~40 tok/s on NVIDIA RTX 6000 Ada Generation
Llama 3.3 — Llama 3.3 70B InstructQ2_K_XS (Tight) · 20.212500000000002 GB · ~38 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3 — Qwen 3 32BQ4_K_M · 19.802999999999997 GB · ~39 tok/s on NVIDIA RTX 6000 Ada Generation
DeepSeek R1 — DeepSeek R1 Distill Qwen 32BQ4_K_M · 19.32 GB · ~40 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 2.5 Family — Qwen 2.5 14B InstructQ4_K_M · 8.4525 GB · ~79 tok/s on NVIDIA RTX 6000 Ada Generation
Nemotron 70B — Nemotron 70B InstructQ4_K_M · 42.62475 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3 — Qwen 3 14BQ4_K_M · 8.935500000000001 GB · ~77 tok/s on NVIDIA RTX 6000 Ada Generation
Codestral — Codestral 22BQ4_K_M · 13.40325 GB · ~54 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 3.6 — Qwen 3.6 35B-A3BQ4_K_M · 21.13125 GB · ~187 tok/s on NVIDIA RTX 6000 Ada Generation
DeepSeek R1 — DeepSeek R1 Distill Qwen 14BQ4_K_M · 8.4525 GB · ~80 tok/s on NVIDIA RTX 6000 Ada Generation
Phi-4 Family — Phi-4 (14B)Q4_K_M · 8.4525 GB · ~79 tok/s on NVIDIA RTX 6000 Ada Generation
Gemma 4 — Gemma 4 31BQ4_K_M · 18.71625 GB · ~41 tok/s on NVIDIA RTX 6000 Ada Generation
Qwen 2.5 VL — Qwen 2.5 VL 72B InstructQ4_K_M · 44.315250000000006 GB · ~19 tok/s on NVIDIA RTX 6000 Ada Generation
Gemma 3 — Gemma 3 27B InstructQ4_K_M · 16.30125 GB · ~40 tok/s on NVIDIA RTX 6000 Ada Generation
购买此硬件 Apple MacBook Pro M5 Pro — 64 GB VRAM · 30 W board power立即云端部署 RunPod 上的 NVIDIA A40 — 低至 $0.44/小时 · 价格核实于 2026-08

或在 Vast.ai 比较

作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。

FAQ

What LLMs run well with 48 GB VRAM?

Qwen 2.5 Family, Qwen 2.5 Family, Llama 3.3, Qwen 3, DeepSeek R1 all fit in 48 GB VRAM.

Which GPUs have 48 GB VRAM?

NVIDIA RTX 6000 Ada Generation, NVIDIA L40S, Apple M5 Pro, Apple M1 Max.

Can-I-Run Pages Near 48 GB

Adjacent VRAM Tiers

No discrete GPU?

Buying Guide

← All VRAM Tiers | Check Your Hardware