作者: Jakub Rusinowski · 最后更新: 2026年7月12日
最受欢迎的入门AI GPU。12 GB VRAM可运行7–8B模型并支持大上下文。二手价150–220美元。本地AI的绝佳起步显卡。
| VRAM | 12 GB |
| Memory Bandwidth | 360 GB/s |
| TDP | 170 W |
| Architecture | Ampere GA106 |
| Release Year | 2021 |
| MSRP at Launch | $329 |
| Inference Speed (Llama 3.1 8B Q4_K_M) | 26–50 tok/s (estimated) |
| Inference Speed (Llama 3.3 70B Q4_K_M) | Does not fit — needs ~44 GB of 12 GB usable |
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
All models below run comfortably in 12 GB VRAM with Q4_K_M quantization.
| Llama 3.1 Family | Llama 3.1 8B Instruct · 6 GB VRAM · Q4_K_M · ollama run llama3.1 |
| Llama 3.2 Family | Llama 3.2 11B Vision Instruct · 7 GB VRAM · Q4_K_M · llama-3-2 |
| Qwen 2.5 Family | Qwen 2.5 14B Instruct · 9 GB VRAM · Q4_K_M · ollama run qwen2.5:14b |
| Gemma 2 Family | Gemma 2 9B IT · 6 GB VRAM · Q4_K_M · ollama run gemma2 |
| Phi-4 Mini | Phi-4 Mini (3.8B) · 3 GB VRAM · Q4_K_M · ollama run phi4-mini |
| Mistral Family | Mistral NeMo 12B · 8 GB VRAM · Q4_K_M · ollama run mistral-nemo |
| SmolLM2 | SmolLM2 1.7B Instruct · 2 GB VRAM · Q4_K_M · ollama run smollm2:1.7b |
Install Ollama then run the recommended model for this GPU:
ollama run llama3.1:8b
Yes — the NVIDIA GeForce RTX 3060 (12GB) has 12 GB VRAM and runs 最受欢迎的入门AI GPU。12 GB VRAM可运行7–8B模型并支持大上下文。二手价150–220美元。本地AI的绝佳起步显卡。
The NVIDIA GeForce RTX 3060 (12GB) is estimated to run Llama 3.1 8B at 26–50 tok/s with Q4_K_M quantization. Llama 3.3 70B does not fit: it needs about 44 GB against 12 GB usable. These are modelled estimates, not measurements — see /en/methodology.
With 12 GB you can run: Llama 3.1 Family, Llama 3.2 Family, Qwen 2.5 Family, Gemma 2 Family, Phi-4 Mini. Use Ollama for the easiest setup: ollama run llama3.1:8b.
← All GPU Reviews | Check Your Hardware | Full Benchmarks | Can I Run It?