AMDDesktop GPUCurrent

AMD Radeon PRO W7900 for local LLMs

Written by Jakub Rusinowski · Last updated

With 48 GB of GDDR6 at 864 GB/s, the PRO W7900 runs 115 catalogued models at Q4_K_M with 8K context. The largest that fits is Llama 3.3 70B Instruct (~45.7 GB), and the top pick is Cosmos 3 Super at 15–32 tok/s.

48 GB of GDDR6 with ECC, 864 GB/s, 295 W. AMD's 48 GB workstation card; check ROCm support for your OS before buying.

Models that run on the PRO W7900

Q4_K_M, 8K context, 48 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Cosmos 3 Super
Cosmos 3
39.4 GB15–32 tok/s
Qwen 3.6 35B-A3B
Qwen 3.6
21.9 GB65–134 tok/s
Nex-N2.5 mini
Nex-N2.5
21.9 GB65–134 tok/s
Nex-N2 mini
Nex-N2
21.9 GB65–134 tok/s
Qwen 3.5 35B-A3B
Qwen 3.5
21.9 GB65–134 tok/s
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB65–135 tok/s
Qwen 3 32B
Qwen 3
22.8 GB15–32 tok/s
Aya Expanse 32B
Aya Expanse
21.6 GB16–33 tok/s
EXAONE 3.5 32B
EXAONE 3.5
22.3 GB16–32 tok/s
Cogito v1 32B
Cogito v1
22.3 GB16–32 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB39–82 tok/s
Olmo 3 32B
Olmo 3
20.1 GB16–33 tok/s
Showing 12 of 115

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Buy this hardware NVIDIA DGX Spark (128GB) — 128 GB VRAM · 150 W board powerDeploy in the cloud now NVIDIA A40 on RunPod — from $0.44/hr · rate checked 2026-08

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
48 GB GDDR6
Memory bandwidth
864 GB/s
Memory bus
384-bit
Architecture
RDNA 3 Navi 31
Series
Radeon PRO W7000
Board power
295 W
Release year
2023
Compute backends
ROCM, VULKAN
Usable for models
48 GB
Best for48 GB workstation AIROCm

Similar GPUs

Frequently asked questions

Can the AMD Radeon PRO W7900 run local LLMs?

Yes. With 48 GB (48 GB usable by a model) it runs 115 of the catalogued models at Q4_K_M with 8K context; the largest is Llama 3.3 70B Instruct, needing about 45.7 GB.

How fast is the AMD Radeon PRO W7900 for AI inference?

It is estimated to run Llama 3.1 8B at 50–103 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 8.0–17 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 48 GB?

Among the best that fit: Cosmos 3 Super, Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B.