NVIDIADesktop GPUCurrent

NVIDIA RTX PRO 6000 Blackwell for local LLMs

Written by Jakub Rusinowski · Last updated

With 96 GB at 1,792 GB/s, the RTX PRO 6000 Blackwell runs 130 catalogued models at Q4_K_M with 8K context. The largest that fits is Devstral-2 123B (~75.1 GB), and the top pick is Devstral-2 123B at 12–26 tok/s.

The flagship workstation GPU of the Blackwell generation: RTX 5090 silicon with 96 GB ECC GDDR7 at 1.79 TB/s. Fits Llama 3.3 70B at Q4 with room for long contexts and concurrent users, or GPT-oss 120B on a single card. The Workstation Edition (600W) needs a serious PSU; the Max-Q variant caps at 300W for quieter builds.

Models that run on the RTX PRO 6000 Blackwell

Q4_K_M, 8K context, 96 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Devstral-2 123B
Devstral
75.1 GB12–26 tok/s
Qwen 3.5 122B-A10B
Qwen 3.5
74.5 GB80–167 tok/s
Nemotron 3 Super 120B-A12B
Nemotron 3 Super
73.3 GB73–152 tok/s
Mistral Small 4 119B-A6.5B
Mistral Small 4
72.6 GB51–105 tok/s
GPT-OSS 120B
GPT-OSS
71.9 GB139–289 tok/s
Llama 4.5 Scout
Llama 4.5
66.6 GB60–125 tok/s
Llama 4 Scout 17B
Llama 4
68.2 GB65–134 tok/s
Llama 3.2 Vision 90B
Llama 3.2 Vision
57.8 GB16–34 tok/s
Qwen3-Coder-Next (80B-A3B MoE)
Qwen3-Coder
49.1 GB128–267 tok/s
Kolibri 1
Kolibri
48.4 GB174–362 tok/s
Llama 3.3 70B Instruct
Llama 3.3
45.7 GB20–43 tok/s
Cogito v1 70B
Cogito v1
45.7 GB20–43 tok/s
Showing 12 of 130

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Buy this hardware NVIDIA RTX PRO 6000 Blackwell 96GB — 96 GB VRAM · 600 W board powerDeploy in the cloud now RunPod

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
NVIDIA RTX PRO 6000 Blackwell 96GB
96 GB VRAM · 600 W board power
2026 prices are volatile — check the current listing.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-07-21.

Memory
96 GB
Memory bandwidth
1,792 GB/s
Architecture
Blackwell GB202
Series
Pro Workstation
Board power
600 W
Release year
2025
Launch price
$8,565
Usable for models
96 GB
Best for70B models at speedMoE modelsdepartment-scale servingfine-tuning
Recommended first model (Ollama)

Similar GPUs

Frequently asked questions

Can the NVIDIA RTX PRO 6000 Blackwell run local LLMs?

Yes. With 96 GB (96 GB usable by a model) it runs 130 of the catalogued models at Q4_K_M with 8K context; the largest is Devstral-2 123B, needing about 75.1 GB.

How fast is the NVIDIA RTX PRO 6000 Blackwell for AI inference?

It is estimated to run Llama 3.1 8B at 113–234 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 21–44 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 96 GB?

Among the best that fit: Devstral-2 123B, Qwen 3.5 122B-A10B, Nemotron 3 Super 120B-A12B, Mistral Small 4 119B-A6.5B, GPT-OSS 120B. The quickest start is Ollama: ollama run llama3.3:70b.