Best Local LLMs for the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)

Written by Jakub Rusinowski · Last updated October 7, 2026

Best all-round pick: GPT-OSS 120B

The NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) has 96 GB of VRAM, of which about 96 GB is available to a model. The largest model it can hold is Devstral-2 123B (123B, 75.1 GB at Q4_K_M).

Best models overall on the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)

ModelScoreMemory at Q4_K_MContextLicence
GPT-OSS 120B96.471.3 GB128KApache-2.0
Qwen 3.5 122B-A10B (MoE)94.474.5 GB125KApache 2.0
GLM-5.1 72B92.844.3 GB125KMIT
Qwen 3.5 122B-A10B92.874.5 GB128KApache-2.0
Command R+ (104B)92.663.6 GB125KCC-BY-NC
Qwen 3.5 72B92.444.3 GB125KApache 2.0

Best model by what you are doing

WorkloadRecommended on this GPU
CodingDevstral-2 123B (99.1)
Qwen3-Coder-Next (80B-A3B MoE) (98)
Qwen 3.5 122B-A10B (MoE) (96.8)
General assistantGPT-OSS 120B (96.4)
Qwen 3.5 122B-A10B (MoE) (94.4)
GLM-5.1 72B (92.8)
ReasoningGLM-5.1 72B (99.3)
Qwen 3.5 122B-A10B (MoE) (97.1)
Gemma 4 27B ⭐ (96.4)
RAGLlama 4 Scout 17B (94.4)
Gemma 4 31B (93.5)
Llama 4.5 Scout (93.1)
AgentsDevstral-2 123B (95.3)
Qwen3-Coder-Next (80B-A3B MoE) (94.3)
Kolibri 1 (93.5)
VisionQwen 3.5 72B (98.2)
Llama 3.2 90B Vision Instruct (96.8)
Qwen 3.7 35B-A3B (93.9)

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)?

GPT-OSS 120B is the strongest all-round pick that fits its 96 GB.

What is the largest model the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) can run?

Devstral-2 123B — 123B parameters, needing 75.1 GB at Q4_K_M.

How much of the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition)'s memory can a model actually use?

About 96 GB of its 96 GB, before the desktop and runtime overhead are accounted for.

Similar GPUs

By Workload

More

← All GPUs | NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) specs