AMDDesktop GPUCurrent

AMD Radeon RX 7600 XT for local LLMs

Written by Jakub Rusinowski · Last updated

With 16 GB of GDDR6 at 288 GB/s, the RX 7600 XT runs 74 catalogued models at Q4_K_M with 8K context. The largest that fits is OLMo 2 13B Instruct (~15.8 GB), and the top pick is GPT-OSS 20B at 38–78 tok/s.

16 GB of GDDR6 on a 128-bit bus, but only 288 GB/s: it holds 14B-class models that a 12 GB card cannot, and runs them slowly.

Models that run on the RX 7600 XT

Q4_K_M, 8K context, 16 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
GPT-OSS 20B
GPT-OSS
13.8 GB38–78 tok/s
Cosmos 3 Nano
Cosmos 3
10.5 GB18–36 tok/s
StarCoder 2 15B
StarCoder 2
10.8 GB11–24 tok/s
Qwen 3 14B
Qwen 3
11.1 GB11–23 tok/s
Phi-4 (14B)
Phi-4 Family
10.9 GB11–23 tok/s
Cogito v1 14B
Cogito v1
10.9 GB11–24 tok/s
Ministral 3 14B
Ministral 3
9.3 GB12–24 tok/s
DeepSeek R1 Distill Qwen 14B
DeepSeek R1
10.6 GB12–24 tok/s
Gemma 4 12B (Unified)
Gemma 4
8 GB13–27 tok/s
Bielik PL 11B v3.0 Instruct
Bielik
7.4 GB14–29 tok/s
Llama 3.2 Vision 11B
Llama 3.2 Vision
8.5 GB14–30 tok/s
Falcon 3 10B Instruct
Falcon 3
8.4 GB15–31 tok/s
Showing 12 of 74

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
16 GB GDDR6
Memory bandwidth
288 GB/s
Memory bus
128-bit
Architecture
RDNA 3 Navi 33
Series
RX 7000-series
Board power
190 W
Release year
2024
Compute backends
ROCM, VULKAN
Usable for models
16 GB
Best for16 GB on a budget14B models

Similar GPUs

Frequently asked questions

Can the AMD Radeon RX 7600 XT run local LLMs?

Yes. With 16 GB (16 GB usable by a model) it runs 74 of the catalogued models at Q4_K_M with 8K context; the largest is OLMo 2 13B Instruct, needing about 15.8 GB.

How fast is the AMD Radeon RX 7600 XT for AI inference?

It is estimated to run Llama 3.1 8B at 20–42 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 16 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 16 GB?

Among the best that fit: GPT-OSS 20B, Cosmos 3 Nano, StarCoder 2 15B, Qwen 3 14B, Phi-4 (14B).