NVIDIADesktop GPUCurrent

NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) for local LLMs

Written by Jakub Rusinowski · Last updated

With 96 GB of GDDR7 at 1,792 GB/s, the RTX PRO 6000 Blackwell (Max-Q Edition) runs 127 catalogued models at Q4_K_M with 8K context. The largest that fits is Devstral-2 123B (~75.1 GB), and the top pick is Devstral-2 123B at 12–26 tok/s.

The 96 GB, 1,792 GB/s memory system of the RTX PRO 6000 in a 300 W, blower-style card for multi-GPU workstations. Board power is not stored: only NVIDIA's page states it.

Models that run on the RTX PRO 6000 Blackwell (Max-Q Edition)

Q4_K_M, 8K context, 96 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Devstral-2 123B
Devstral
75.1 GB12–26 tok/s
Qwen 3.5 122B-A10B
Qwen 3.5
74.5 GB80–167 tok/s
Nemotron 3 Super 120B-A12B
Nemotron 3 Super
73.3 GB73–152 tok/s
Mistral Small 4 119B-A6.5B
Mistral Small 4
72.6 GB51–105 tok/s
GPT-OSS 120B
GPT-OSS
71.9 GB139–289 tok/s
Llama 4.5 Scout
Llama 4.5
66.6 GB60–125 tok/s
Llama 4 Scout 17B
Llama 4
68.2 GB65–134 tok/s
Llama 3.2 Vision 90B
Llama 3.2 Vision
57.8 GB16–34 tok/s
Qwen3-Coder-Next (80B-A3B MoE)
Qwen3-Coder
49.1 GB128–267 tok/s
Kolibri 1
Kolibri
48.4 GB174–362 tok/s
Llama 3.3 70B Instruct
Llama 3.3
45.7 GB20–43 tok/s
Cogito v1 70B
Cogito v1
45.7 GB20–43 tok/s
Showing 12 of 127

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Buy this hardware NVIDIA DGX Spark (128GB) — 128 GB VRAM · 150 W board powerDeploy in the cloud now RunPod

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
96 GB GDDR7
Memory bandwidth
1,792 GB/s
Memory bus
512-bit
Architecture
Blackwell GB202
Series
RTX PRO Blackwell
Release year
2025
Compute backends
CUDA, VULKAN
Usable for models
96 GB
Best for96 GB workstation AImulti-GPU workstations

Similar GPUs

Frequently asked questions

Can the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) run local LLMs?

Yes. With 96 GB (96 GB usable by a model) it runs 127 of the catalogued models at Q4_K_M with 8K context; the largest is Devstral-2 123B, needing about 75.1 GB.

How fast is the NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) for AI inference?

It is estimated to run Llama 3.1 8B at 113–234 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 21–44 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 96 GB?

Among the best that fit: Devstral-2 123B, Qwen 3.5 122B-A10B, Nemotron 3 Super 120B-A12B, Mistral Small 4 119B-A6.5B, GPT-OSS 120B.