Written by Jakub Rusinowski · Last updated July 21, 2026
Best all-round pick: GPT-oss 120B
The NVIDIA RTX PRO 6000 Blackwell has 96 GB of VRAM, of which about 96 GB is available to a model. The largest model it can hold is Devstral-2 123B (123B, 75.1 GB at Q4_K_M).
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| GPT-oss 120B | 96.4 | 73.3 GB | 125K | Apache-2.0 |
| Nemotron 70B Instruct | 96.2 | 43.4 GB | 125K | Llama Community |
| Qwen 3.5 122B-A10B (MoE) | 94.4 | 74.5 GB | 125K | Apache 2.0 |
| GLM-5.1 72B | 92.8 | 44.3 GB | 125K | MIT |
| Qwen 3.5 122B-A10B | 92.8 | 74.5 GB | 128K | Apache-2.0 |
| Command R+ (104B) | 92.6 | 63.6 GB | 125K | CC-BY-NC |
| Workload | Recommended on this GPU |
|---|---|
| Coding | Devstral-2 123B (98.7) Qwen3-Coder 80B-A3B (MoE) (97.6) Qwen 3.5 122B-A10B (MoE) (96.8) |
| General assistant | GPT-oss 120B (96.4) Nemotron 70B Instruct (96.2) Qwen 3.5 122B-A10B (MoE) (94.4) |
| Reasoning | GLM-5.1 72B (99.3) Qwen 3.5 122B-A10B (MoE) (97.1) Gemma 4 27B ⭐ (96.4) |
| RAG | Llama 4 Scout 17B (94.4) Gemma 4 31B (93.5) Llama 4.5 Scout (93.1) |
| Agents | Ternary Bonsai 27B (89.6) Qwen 3.5 122B-A10B (MoE) (87.7) GLM-5.1 72B (87) |
| Vision | Qwen 3.5 72B (98.2) Llama 3.2 90B Vision Instruct (96.8) Qwen 3.7 35B-A3B (93.9) |
GPT-oss 120B is the strongest all-round pick that fits its 96 GB.
Devstral-2 123B — 123B parameters, needing 75.1 GB at Q4_K_M.
About 96 GB of its 96 GB, before the desktop and runtime overhead are accounted for.