NVIDIADesktop GPUCurrent

NVIDIA RTX 5000 Ada Generation for local LLMs

Written by Jakub Rusinowski · Last updated

With 32 GB of GDDR6 at 576 GB/s, the RTX 5000 Ada Generation runs 110 catalogued models at Q4_K_M with 8K context. The largest that fits is Gemma 3 27B Instruct (~25.4 GB), and the top pick is Qwen 3.6 35B-A3B at 75–144 tok/s.

32 GB of GDDR6 with ECC, 576 GB/s, 250 W.

Models that run on the RTX 5000 Ada Generation

Q4_K_M, 8K context, 32 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Qwen 3.6 35B-A3B
Qwen 3.6
21.9 GB75–144 tok/s
Nex-N2.5 mini
Nex-N2.5
21.9 GB75–144 tok/s
Nex-N2 mini
Nex-N2
21.9 GB75–144 tok/s
Qwen 3.5 35B-A3B
Qwen 3.5
21.9 GB75–144 tok/s
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB75–145 tok/s
Qwen 3 32B
Qwen 3
22.8 GB16–30 tok/s
Aya Expanse 32B
Aya Expanse
21.6 GB17–32 tok/s
EXAONE 3.5 32B
EXAONE 3.5
22.3 GB16–31 tok/s
Cogito v1 32B
Cogito v1
22.3 GB16–31 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB43–83 tok/s
Olmo 3 32B
Olmo 3
20.1 GB16–31 tok/s
DeepSeek R1 Distill Qwen 32B
DeepSeek R1
22.3 GB16–31 tok/s
Showing 12 of 110

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Buy this hardware NVIDIA GeForce RTX 5090 32GB — 32 GB VRAM · 575 W board powerDeploy in the cloud now NVIDIA A40 on RunPod — from $0.44/hr · rate checked 2026-08

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
32 GB GDDR6
Memory bandwidth
576 GB/s
Memory bus
256-bit
Architecture
Ada Lovelace AD102
Series
RTX Ada Generation
Board power
250 W
Release year
2023
Compute backends
CUDA, VULKAN
Usable for models
32 GB
Best for32 GB workstation AI

Similar GPUs

Frequently asked questions

Can the NVIDIA RTX 5000 Ada Generation run local LLMs?

Yes. With 32 GB (32 GB usable by a model) it runs 110 of the catalogued models at Q4_K_M with 8K context; the largest is Gemma 3 27B Instruct, needing about 25.4 GB.

How fast is the NVIDIA RTX 5000 Ada Generation for AI inference?

It is estimated to run Llama 3.1 8B at 56–107 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 32 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 32 GB?

Among the best that fit: Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B, Laguna XS 2.1 33B-A3B.