Written by Jakub Rusinowski · Last updated July 12, 2026
Best all-round pick: Qwen 3 8B
The NVIDIA GeForce RTX 4070 has 12 GB of VRAM, of which about 12 GB is available to a model. The largest model it can hold is Cosmos 3 Nano (16B, 10.5 GB at Q4_K_M).
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| Qwen 3 8B | 94.5 | 5.8 GB | 125K | Apache 2.0 |
| Qwen 3.5 14B | 94.4 | 9.3 GB | 125K | Apache 2.0 |
| Gemma 3 12B Instruct | 93.6 | 8 GB | 128K | Gemma |
| Qwen 3.5 14B | 93.5 | 9.3 GB | 125K | Apache 2.0 |
| Gemma 4 12B | 93.1 | 8 GB | 125K | Gemma License (commercial OK) |
| GLM-6 9B | 92.2 | 6.2 GB | 125K | MIT |
| Workload | Recommended on this GPU |
|---|---|
| Coding | Qwen 3 14B (95.9) Qwen3-Coder 8B (94.9) DeepSeek R1 Distill Qwen 14B (94.2) |
| General assistant | Qwen 3 8B (94.5) Qwen 3.5 14B (94.4) Gemma 3 12B Instruct (93.6) |
| Reasoning | Qwen 3 14B (96) DeepSeek R1 Distill Qwen 14B (94.4) Qwen 3.5 14B (94) |
| RAG | Qwen 3 14B (84.4) Qwen 2.5 14B Instruct (81.8) Qwen 3.5 14B (81.5) |
| Agents | GLM-6 9B (86) GLM-5 9B (84.6) Qwen 3.5 14B (84.6) |
| Vision | Qwen 3.5 14B (95.7) Gemma 4 12B (92.3) Gemma 4 12B (Unified) (92) |
These models are strong picks generally but exceed the 12 GB this card makes available.
| Model | Needs at Q4_K_M | Short by |
|---|---|---|
| Gemma 4 27B ⭐ | 17.2 GB | ~5.2 GB |
| Mistral Small 3.1 24B | 15.1 GB | ~3.1 GB |
| Qwen 3.7 35B-A3B | 22 GB | ~10 GB |
| Gemma 4 31B | 19.6 GB | ~7.6 GB |
Qwen 3 8B is the strongest all-round pick that fits its 12 GB.
Cosmos 3 Nano — 16B parameters, needing 10.5 GB at Q4_K_M.
About 12 GB of its 12 GB, before the desktop and runtime overhead are accounted for.