NVIDIALaptop GPUCurrent

NVIDIA RTX PRO 5000 Blackwell Laptop GPU for local LLMs

Written by Jakub Rusinowski · Last updated

With 24 GB of GDDR7 at 896 GB/s, the RTX PRO 5000 Blackwell Laptop GPU runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Qwen 3 32B (~22.8 GB), and the top pick is Laguna XS 2.1 33B-A3B at 92–192 tok/s.

24 GB of GDDR7 on a 256-bit bus, 896 GB/s, in mobile workstations. Power depends on the chassis and is not stored.

Models that run on the RTX PRO 5000 Blackwell Laptop GPU

Q4_K_M, 8K context, 24 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB92–192 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB55–114 tok/s
Olmo 3 32B
Olmo 3
20.1 GB22–45 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB88–182 tok/s
Gemma 4 31B
Gemma 4
19.5 GB22–46 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB110–228 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB110–228 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB93–194 tok/s
Granite 4.1 30B
IBM Granite 4.1
18.9 GB23–48 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB93–194 tok/s
North Mini Code 1.0 30B-A3B
North Mini Code
18.9 GB93–194 tok/s
Granite 4.2 30B
IBM Granite 4.2
18.9 GB23–48 tok/s
Showing 12 of 109

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
24 GB GDDR7
Memory bandwidth
896 GB/s
Memory bus
256-bit
Architecture
Blackwell GB203
Series
RTX PRO Blackwell (Laptop)
Release year
2025
Compute backends
CUDA
Usable for models
24 GB
Best for24 GB laptop AI

Similar GPUs

Frequently asked questions

Can the NVIDIA RTX PRO 5000 Blackwell Laptop GPU run local LLMs?

Yes. With 24 GB (24 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Qwen 3 32B, needing about 22.8 GB.

How fast is the NVIDIA RTX PRO 5000 Blackwell Laptop GPU for AI inference?

It is estimated to run Llama 3.1 8B at 70–145 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 24 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 24 GB?

Among the best that fit: Laguna XS 2.1 33B-A3B, Granite 4.0 Small-H 32B-A9B, Olmo 3 32B, Nemotron-Cascade 2 30B-A3B, Gemma 4 31B.