Best Local LLMs for the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)
Written by Jakub Rusinowski · Last updated October 7, 2026
Best all-round pick: GPT-OSS 120B
The NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) has 96 GB of VRAM, of which about 96 GB is available to a model. The largest model it can hold is Devstral-2 123B (123B, 75.1 GB at Q4_K_M).
Best models overall on the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| GPT-OSS 120B | 96.4 | 71.3 GB | 128K | Apache-2.0 |
| Qwen 3.5 122B-A10B (MoE) | 94.4 | 74.5 GB | 125K | Apache 2.0 |
| GLM-5.1 72B | 92.8 | 44.3 GB | 125K | MIT |
| Qwen 3.5 122B-A10B | 92.8 | 74.5 GB | 128K | Apache-2.0 |
| Command R+ (104B) | 92.6 | 63.6 GB | 125K | CC-BY-NC |
| Qwen 3.5 72B | 92.4 | 44.3 GB | 125K | Apache 2.0 |
Best model by what you are doing
| Workload | Recommended on this GPU |
|---|---|
| Coding | Devstral-2 123B (99.1) Qwen3-Coder-Next (80B-A3B MoE) (98) Qwen 3.5 122B-A10B (MoE) (96.8) |
| General assistant | GPT-OSS 120B (96.4) Qwen 3.5 122B-A10B (MoE) (94.4) GLM-5.1 72B (92.8) |
| Reasoning | GLM-5.1 72B (99.3) Qwen 3.5 122B-A10B (MoE) (97.1) Gemma 4 27B ⭐ (96.4) |
| RAG | Llama 4 Scout 17B (94.4) Gemma 4 31B (93.5) Llama 4.5 Scout (93.1) |
| Agents | Devstral-2 123B (95.3) Qwen3-Coder-Next (80B-A3B MoE) (94.3) Kolibri 1 (93.5) |
| Vision | Qwen 3.5 72B (98.2) Llama 3.2 90B Vision Instruct (96.8) Qwen 3.7 35B-A3B (93.9) |
How these numbers are calculated
- Usable memory is 96 GB of the card's 96 GB.
- Fit is measured at Q4_K_M — quantized weights plus KV cache plus 0.8 GB runtime overhead.
- Within this card's budget, models are ranked on how well they use the memory available, not on being small — so the recommendation scales with the hardware.
FAQ
What is the best LLM for the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)?
GPT-OSS 120B is the strongest all-round pick that fits its 96 GB.
What is the largest model the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) can run?
Devstral-2 123B — 123B parameters, needing 75.1 GB at Q4_K_M.
How much of the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)'s memory can a model actually use?
About 96 GB of its 96 GB, before the desktop and runtime overhead are accounted for.
Similar GPUs
By Workload
- Best local LLMs for coding
- Best local LLMs for general assistant
- Best local LLMs for reasoning
- Best local LLMs for rag
More
- NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) specifications
- Best LLMs for 80 GB VRAM
- Check your own hardware
← All GPUs | NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) specs