AMDDesktop GPUCurrent

AMD Radeon RX 9070 GRE for local LLMs

Written by Jakub Rusinowski · Last updated

With 12 GB of GDDR6 at 432 GB/s, the RX 9070 GRE runs 67 catalogued models at Q4_K_M with 8K context. The largest that fits is Gemma 3 12B Instruct (~11.1 GB), and the top pick is Cosmos 3 Nano at 27–56 tok/s.

12 GB of GDDR6 on a 192-bit bus. AMD states 432 GB/s; a figure of about 482 GB/s seen in one launch article is not supported by AMD's page or by the bus arithmetic.

Models that run on the RX 9070 GRE

Q4_K_M, 8K context, 12 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Cosmos 3 Nano
Cosmos 3
10.5 GB27–56 tok/s
Ministral 3 14B
Ministral 3
9.3 GB18–37 tok/s
DeepSeek R1 Distill Qwen 14B
DeepSeek R1
10.6 GB18–38 tok/s
Gemma 4 12B (Unified)
Gemma 4
8 GB20–42 tok/s
Mistral NeMo 12B
Mistral Family
9.4 GB20–42 tok/s
Bielik PL 11B v3.0 Instruct
Bielik
7.4 GB22–45 tok/s
Llama 3.2 Vision 11B
Llama 3.2 Vision
8.5 GB22–46 tok/s
Llama 3.2 11B Vision Instruct
Llama 3.2 Family
8.5 GB22–46 tok/s
Falcon 3 10B Instruct
Falcon 3
8.4 GB23–47 tok/s
Qwen 3.5 9B
Qwen 3.5
6.2 GB26–53 tok/s
GLM-4 9B
GLM-4.7 / GLM-Z1
6.2 GB26–53 tok/s
GLM-4.6V-Flash 9B
GLM-4.6V
6.2 GB26–53 tok/s
Showing 12 of 67

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
12 GB GDDR6
Memory bandwidth
432 GB/s
Memory bus
192-bit
Architecture
RDNA 4 Navi 48
Series
RX 9000-series
Board power
220 W
Release year
2025
Compute backends
ROCM, VULKAN
Usable for models
12 GB
Best for12 GBRDNA 4

Similar GPUs

Frequently asked questions

Can the AMD Radeon RX 9070 GRE run local LLMs?

Yes. With 12 GB (12 GB usable by a model) it runs 67 of the catalogued models at Q4_K_M with 8K context; the largest is Gemma 3 12B Instruct, needing about 11.1 GB.

How fast is the AMD Radeon RX 9070 GRE for AI inference?

It is estimated to run Llama 3.1 8B at 31–64 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 12 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 12 GB?

Among the best that fit: Cosmos 3 Nano, Ministral 3 14B, DeepSeek R1 Distill Qwen 14B, Gemma 4 12B (Unified), Mistral NeMo 12B.