NVIDIADesktop GPUCurrent

NVIDIA RTX PRO 4000 Blackwell for local LLMs

Written by Jakub Rusinowski · Last updated

With 24 GB of GDDR7 at 672 GB/s, the RTX PRO 4000 Blackwell runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Qwen 3 32B (~22.8 GB), and the top pick is Laguna XS 2.1 33B-A3B at 75–156 tok/s.

24 GB of GDDR7 with ECC on a 192-bit bus, 672 GB/s, single slot. Board power is not stored: NVIDIA states 145 W while two other references give 140 W.

Models that run on the RTX PRO 4000 Blackwell

Q4_K_M, 8K context, 24 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB75–156 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB43–90 tok/s
Olmo 3 32B
Olmo 3
20.1 GB17–34 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB71–148 tok/s
Gemma 4 31B
Gemma 4
19.5 GB17–35 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB91–189 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB91–189 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB76–158 tok/s
Granite 4.1 30B
IBM Granite 4.1
18.9 GB18–36 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB76–158 tok/s
North Mini Code 1.0 30B-A3B
North Mini Code
18.9 GB76–158 tok/s
Granite 4.2 30B
IBM Granite 4.2
18.9 GB18–36 tok/s
Showing 12 of 109

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
24 GB GDDR7
Memory bandwidth
672 GB/s
Memory bus
192-bit
Architecture
Blackwell GB203
Series
RTX PRO Blackwell
Release year
2025
Compute backends
CUDA, VULKAN
Usable for models
24 GB
Best for24 GB single-slot32B at Q4

Similar GPUs

Frequently asked questions

Can the NVIDIA RTX PRO 4000 Blackwell run local LLMs?

Yes. With 24 GB (24 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Qwen 3 32B, needing about 22.8 GB.

How fast is the NVIDIA RTX PRO 4000 Blackwell for AI inference?

It is estimated to run Llama 3.1 8B at 56–116 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 24 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 24 GB?

Among the best that fit: Laguna XS 2.1 33B-A3B, Granite 4.0 Small-H 32B-A9B, Olmo 3 32B, Nemotron-Cascade 2 30B-A3B, Gemma 4 31B.