Best LLMs for 12 GB VRAM
作者: Jakub Rusinowski · 最后更新: 2026年7月15日
These are the strongest local models that fit entirely in 12 GB of VRAM, ranked by capability, with the quantization level and estimated tokens/sec needed to fit.
GPUs at This Tier
Ranked Models
| Qwen 2.5 Family — Qwen 2.5 14B Instruct | Q4_K_M · 8.4525 GB · ~54 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Qwen 3 — Qwen 3 14B | Q4_K_M · 8.935500000000001 GB · ~52 tok/s on NVIDIA GeForce RTX 3080 12GB |
| DeepSeek R1 — DeepSeek R1 Distill Qwen 14B | Q4_K_M · 8.4525 GB · ~55 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Phi-4 Family — Phi-4 (14B) | Q4_K_M · 8.4525 GB · ~54 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Granite 3.0 — Granite 3.0 8B Instruct | Q4_K_M · 4.83 GB · ~84 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Bonsai 27B — Ternary Bonsai 27B | Ternary (1.58-bit, ~1.71 bpw) · 5.77125 GB · ~72 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Qwen 2.5 Family — Qwen 2.5 7B Instruct | Q4_K_M · 4.5885 GB · ~93 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Qwen 3 — Qwen 3 8B | Q4_K_M · 4.950749999999999 GB · ~83 tok/s on NVIDIA GeForce RTX 3080 12GB |
| DeepSeek R1 — DeepSeek R1 Distill Llama 8B | Q4_K_M · 4.83 GB · ~85 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Mistral Family — Mistral NeMo 12B | Q4_K_M · 7.245 GB · ~62 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Gemma 3 — Gemma 3 12B Instruct | Q4_K_M · 7.245 GB · ~56 tok/s on NVIDIA GeForce RTX 3080 12GB |
| StarCoder 2 — StarCoder 2 15B | Q4_K_M · 9.358125 GB · ~52 tok/s on NVIDIA GeForce RTX 3080 12GB |
| InternLM 3 — InternLM 3 8B Instruct | Q4_K_M · 5.313000000000001 GB · ~84 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Qwen 3.5 — Qwen 3.5 9B | Q4_K_M · 5.43375 GB · ~78 tok/s on NVIDIA GeForce RTX 3080 12GB |
| Bonsai 27B — 1-bit Bonsai 27B | 1-bit (binary, ~1.125 bpw) · 3.796875 GB · ~95 tok/s on NVIDIA GeForce RTX 3080 12GB |
购买此硬件 Intel Arc B580 12GB — 12 GB VRAM · 190 W board power立即云端部署 RunPod 上的 RTX 4090 — 低至 $0.34/小时 · 价格核实于 2026-07
或在 Vast.ai 比较,低至 $0.35/小时 (typical low · varies)
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
FAQ
What LLMs run well with 12 GB VRAM?
Qwen 2.5 Family, Qwen 3, DeepSeek R1, Phi-4 Family, Granite 3.0 all fit in 12 GB VRAM.
Which GPUs have 12 GB VRAM?
NVIDIA GeForce RTX 3080 12GB, Intel Arc B580, NVIDIA GeForce RTX 3060 (12GB), AMD Radeon RX 7700 XT.
Can-I-Run Pages Near 12 GB
- DeepSeek R1 on NVIDIA GeForce RTX 5070
- DeepSeek R1 on NVIDIA GeForce RTX 4070 Ti
- DeepSeek R1 on NVIDIA GeForce RTX 4070 Super
- DeepSeek R1 on NVIDIA GeForce RTX 4070
- DeepSeek R1 on NVIDIA GeForce RTX 3060 (12GB)
- DeepSeek R1 on Intel Arc B580
- Llama 3.3 on NVIDIA GeForce RTX 5070
- Llama 3.3 on NVIDIA GeForce RTX 4070 Ti