NVIDIA RTX A6000 for local LLMs
Written by Jakub Rusinowski · Last updated
With 48 GB of GDDR6 at 768 GB/s, the RTX A6000 runs 115 catalogued models at Q4_K_M with 8K context. The largest that fits is Llama 3.3 70B Instruct (~45.7 GB), and the top pick is Cosmos 3 Super at 14–28 tok/s.
48 GB of GDDR6 with ECC, 768 GB/s, 300 W. The usual route to 48 GB on one used card.
Models that run on the RTX A6000
Q4_K_M, 8K context, 48 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 48 GB | VRAM | Speed |
|---|---|---|---|
| Cosmos 3 Super Cosmos 3 | 39.4 GB | 14–28 tok/s | |
| Qwen 3.6 35B-A3B Qwen 3.6 | 21.9 GB | 69–132 tok/s | |
| Nex-N2.5 mini Nex-N2.5 | 21.9 GB | 69–132 tok/s | |
| Nex-N2 mini Nex-N2 | 21.9 GB | 69–132 tok/s | |
| Qwen 3.5 35B-A3B Qwen 3.5 | 21.9 GB | 69–132 tok/s | |
| Laguna XS 2.1 33B-A3B Poolside Laguna XS 2.1 | 20.7 GB | 69–133 tok/s | |
| Qwen 3 32B Qwen 3 | 22.8 GB | 14–27 tok/s | |
| Aya Expanse 32B Aya Expanse | 21.6 GB | 15–29 tok/s | |
| EXAONE 3.5 32B EXAONE 3.5 | 22.3 GB | 15–28 tok/s | |
| Cogito v1 32B Cogito v1 | 22.3 GB | 15–28 tok/s | |
| Granite 4.0 Small-H 32B-A9B IBM Granite 4.0 | 20.1 GB | 39–75 tok/s | |
| Olmo 3 32B Olmo 3 | 20.1 GB | 15–28 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 48 GB GDDR6
- Memory bandwidth
- 768 GB/s
- Memory bus
- 384-bit
- Architecture
- Ampere GA102
- Series
- RTX A-series
- Board power
- 300 W
- Release year
- 2020
- Launch price
- $4,649
- Compute backends
- CUDA, VULKAN
- Usable for models
- 48 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA RTX A6000 run local LLMs?
Yes. With 48 GB (48 GB usable by a model) it runs 115 of the catalogued models at Q4_K_M with 8K context; the largest is Llama 3.3 70B Instruct, needing about 45.7 GB.
How fast is the NVIDIA RTX A6000 for AI inference?
It is estimated to run Llama 3.1 8B at 51–98 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 7.3–14 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 48 GB?
Among the best that fit: Cosmos 3 Super, Qwen 3.6 35B-A3B, Nex-N2.5 mini, Nex-N2 mini, Qwen 3.5 35B-A3B.