NVIDIA RTX PRO 5000 Blackwell (72 GB) for local LLMs
Written by Jakub Rusinowski · Last updated
With 72 GB of GDDR7 at 1,344 GB/s, the RTX PRO 5000 Blackwell (72 GB) runs 123 catalogued models at Q4_K_M with 8K context. The largest that fits is GPT-OSS 120B (~71.9 GB), and the top pick is Llama 3.2 Vision 90B at 12–26 tok/s.
72 GB of GDDR7 with ECC at the same 1,344 GB/s and 300 W as the 48 GB version. More memory, not more speed.
Models that run on the RTX PRO 5000 Blackwell (72 GB)
Q4_K_M, 8K context, 72 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 72 GB | VRAM | Speed |
|---|---|---|---|
| Llama 3.2 Vision 90B Llama 3.2 Vision | 57.8 GB | 12–26 tok/s | |
| Qwen3-Coder-Next (80B-A3B MoE) Qwen3-Coder | 49.1 GB | 108–225 tok/s | |
| Kolibri 1 Kolibri | 48.4 GB | 154–320 tok/s | |
| Llama 3.3 70B Instruct Llama 3.3 | 45.7 GB | 16–32 tok/s | |
| Cogito v1 70B Cogito v1 | 45.7 GB | 16–32 tok/s | |
| Apertus 70B Apertus | 43.1 GB | 16–33 tok/s | |
| Cosmos 3 Super Cosmos 3 | 39.4 GB | 31–64 tok/s | |
| Qwen 3.6 35B-A3B Qwen 3.6 | 21.9 GB | 119–247 tok/s | |
| Nex-N2.5 mini Nex-N2.5 | 21.9 GB | 119–247 tok/s | |
| Nex-N2 mini Nex-N2 | 21.9 GB | 119–247 tok/s | |
| Qwen 3.5 35B-A3B Qwen 3.5 | 21.9 GB | 119–247 tok/s | |
| Laguna XS 2.1 33B-A3B Poolside Laguna XS 2.1 | 20.7 GB | 119–248 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.77/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 72 GB GDDR7
- Memory bandwidth
- 1,344 GB/s
- Memory bus
- 384-bit
- Architecture
- Blackwell GB202
- Series
- RTX PRO Blackwell
- Board power
- 300 W
- Release year
- 2025
- Compute backends
- CUDA, VULKAN
- Usable for models
- 72 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA RTX PRO 5000 Blackwell (72 GB) run local LLMs?
Yes. With 72 GB (72 GB usable by a model) it runs 123 of the catalogued models at Q4_K_M with 8K context; the largest is GPT-OSS 120B, needing about 71.9 GB.
How fast is the NVIDIA RTX PRO 5000 Blackwell (72 GB) for AI inference?
It is estimated to run Llama 3.1 8B at 94–194 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 16–33 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 72 GB?
Among the best that fit: Llama 3.2 Vision 90B, Qwen3-Coder-Next (80B-A3B MoE), Kolibri 1, Llama 3.3 70B Instruct, Cogito v1 70B.