NVIDIA RTX PRO 6000 Blackwell for local LLMs
Written by Jakub Rusinowski · Last updated
With 96 GB at 1,792 GB/s, the RTX PRO 6000 Blackwell runs 130 catalogued models at Q4_K_M with 8K context. The largest that fits is Devstral-2 123B (~75.1 GB), and the top pick is Devstral-2 123B at 12–26 tok/s.
The flagship workstation GPU of the Blackwell generation: RTX 5090 silicon with 96 GB ECC GDDR7 at 1.79 TB/s. Fits Llama 3.3 70B at Q4 with room for long contexts and concurrent users, or GPT-oss 120B on a single card. The Workstation Edition (600W) needs a serious PSU; the Max-Q variant caps at 300W for quieter builds.
Models that run on the RTX PRO 6000 Blackwell
Q4_K_M, 8K context, 96 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 96 GB | VRAM | Speed |
|---|---|---|---|
| Devstral-2 123B Devstral | 75.1 GB | 12–26 tok/s | |
| Qwen 3.5 122B-A10B Qwen 3.5 | 74.5 GB | 80–167 tok/s | |
| Nemotron 3 Super 120B-A12B Nemotron 3 Super | 73.3 GB | 73–152 tok/s | |
| Mistral Small 4 119B-A6.5B Mistral Small 4 | 72.6 GB | 51–105 tok/s | |
| GPT-OSS 120B GPT-OSS | 71.9 GB | 139–289 tok/s | |
| Llama 4.5 Scout Llama 4.5 | 66.6 GB | 60–125 tok/s | |
| Llama 4 Scout 17B Llama 4 | 68.2 GB | 65–134 tok/s | |
| Llama 3.2 Vision 90B Llama 3.2 Vision | 57.8 GB | 16–34 tok/s | |
| Qwen3-Coder-Next (80B-A3B MoE) Qwen3-Coder | 49.1 GB | 128–267 tok/s | |
| Kolibri 1 Kolibri | 48.4 GB | 174–362 tok/s | |
| Llama 3.3 70B Instruct Llama 3.3 | 45.7 GB | 20–43 tok/s | |
| Cogito v1 70B Cogito v1 | 45.7 GB | 20–43 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-07-21.
- Memory
- 96 GB
- Memory bandwidth
- 1,792 GB/s
- Architecture
- Blackwell GB202
- Series
- Pro Workstation
- Board power
- 600 W
- Release year
- 2025
- Launch price
- $8,565
- Usable for models
- 96 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA RTX PRO 6000 Blackwell run local LLMs?
Yes. With 96 GB (96 GB usable by a model) it runs 130 of the catalogued models at Q4_K_M with 8K context; the largest is Devstral-2 123B, needing about 75.1 GB.
How fast is the NVIDIA RTX PRO 6000 Blackwell for AI inference?
It is estimated to run Llama 3.1 8B at 113–234 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 21–44 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 96 GB?
Among the best that fit: Devstral-2 123B, Qwen 3.5 122B-A10B, Nemotron 3 Super 120B-A12B, Mistral Small 4 119B-A6.5B, GPT-OSS 120B. The quickest start is Ollama: ollama run llama3.3:70b.
Keep going
Can I run it on the RTX PRO 6000 Blackwell?
More for this card
Tools