AMDDesktop GPUCurrent

AMD Radeon RX 9070 for local LLMs

Written by Jakub Rusinowski · Last updated

With 16 GB at 640 GB/s, the RX 9070 runs 75 catalogued models at Q4_K_M with 8K context. The largest that fits is OLMo 2 13B Instruct (~15.8 GB), and the top pick is EuroLLM 22B at 18–37 tok/s.

Same Navi 48 chip and 16 GB VRAM as the 9070 XT but with lower clocks and bandwidth (640 GB/s). Still handles all 13–14B models comfortably via ROCm or Vulkan.

Models that run on the RX 9070

Q4_K_M, 8K context, 16 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
EuroLLM 22B
EuroLLM
14.1 GB18–37 tok/s
GPT-OSS 20B
GPT-OSS
13.8 GB70–146 tok/s
InternLM 3 20B Instruct
InternLM 3
12.9 GB19–40 tok/s
Cosmos 3 Nano
Cosmos 3
10.5 GB37–77 tok/s
StarCoder 2 15B
StarCoder 2
10.8 GB25–52 tok/s
Qwen 3 14B
Qwen 3
11.1 GB25–51 tok/s
Phi-4 (14B)
Phi-4 Family
10.9 GB25–52 tok/s
Cogito v1 14B
Cogito v1
10.9 GB25–52 tok/s
Ministral 3 14B
Ministral 3
9.3 GB26–53 tok/s
DeepSeek R1 Distill Qwen 14B
DeepSeek R1
10.6 GB26–53 tok/s
Gemma 4 12B (Unified)
Gemma 4
8 GB29–60 tok/s
Bielik PL 11B v3.0 Instruct
Bielik
7.4 GB31–64 tok/s
Showing 12 of 75

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 9070 16GB
16 GB VRAM · 220 W board power
2026 prices are volatile — check the current listing.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-07-12.

Memory
16 GB
Memory bandwidth
640 GB/s
Architecture
RDNA 4 Navi 48
Series
RX 9000-series
Board power
220 W
Release year
2025
Launch price
$549
Usable for models
16 GB
Best for14B modelsAMD 2025 mid-flagship16GB value
Recommended first model (Ollama)

Similar GPUs

Frequently asked questions

Can the AMD Radeon RX 9070 run local LLMs?

Yes. With 16 GB (16 GB usable by a model) it runs 75 of the catalogued models at Q4_K_M with 8K context; the largest is OLMo 2 13B Instruct, needing about 15.8 GB.

How fast is the AMD Radeon RX 9070 for AI inference?

It is estimated to run Llama 3.1 8B at 42–87 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 16 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 16 GB?

Among the best that fit: EuroLLM 22B, GPT-OSS 20B, InternLM 3 20B Instruct, Cosmos 3 Nano, StarCoder 2 15B. The quickest start is Ollama: ollama run qwen3:14b.