Best Local LLMs for the NVIDIA RTX A6000
Written by Jakub Rusinowski · Last updated October 7, 2026
Best all-round pick: GLM-5.1 72B
The NVIDIA RTX A6000 has 48 GB of VRAM, of which about 48 GB is available to a model. The largest model it can hold is Qwen 2.5 VL 72B Instruct (73B, 45.1 GB at Q4_K_M).
Best models overall on the NVIDIA RTX A6000
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| GLM-5.1 72B | 94.7 | 44.3 GB | 125K | MIT |
| Qwen 3.5 72B | 94.4 | 44.3 GB | 125K | Apache 2.0 |
| Gemma 4 27B ⭐ | 94 | 17.1 GB | 125K | Gemma License (commercial OK) |
| Llama 3.3 70B Instruct | 93.8 | 43.1 GB | 128K | Llama Community |
| Qwen 3.7 35B-A3B | 93.2 | 21.9 GB | 256K | Apache-2.0 |
| Qwen3.8 27B | 92.7 | 17.6 GB | 256K | Apache-2.0 |
Best model by what you are doing
| Workload | Recommended on this GPU |
|---|---|
| Coding | Qwen 3.7 35B-A3B (96.9) Qwen 3.5 72B (96.6) Qwen 3.6 35B-A3B (96.2) |
| General assistant | GLM-5.1 72B (94.7) Qwen 3.5 72B (94.4) Gemma 4 27B ⭐ (94) |
| Reasoning | GLM-5.1 72B (100) Gemma 4 27B ⭐ (98.1) GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (97.7) |
| RAG | Gemma 4 31B (95.6) Qwen 3.6 27B (94.5) Qwen 3.7 35B-A3B (93.8) |
| Agents | Ternary Bonsai 27B (91.6) Nemotron-Cascade 2 30B-A3B (91) Nemotron 3.5 Lightning 30B-A3B (90.6) |
| Vision | Qwen 3.5 72B (99.7) Qwen 3.7 35B-A3B (96.5) Qwen 3.6 35B-A3B (95.8) |
How these numbers are calculated
- Usable memory is 48 GB of the card's 48 GB.
- Fit is measured at Q4_K_M — quantized weights plus KV cache plus 0.8 GB runtime overhead.
- Within this card's budget, models are ranked on how well they use the memory available, not on being small — so the recommendation scales with the hardware.
FAQ
What is the best LLM for the NVIDIA RTX A6000?
GLM-5.1 72B is the strongest all-round pick that fits its 48 GB.
What is the largest model the NVIDIA RTX A6000 can run?
Qwen 2.5 VL 72B Instruct — 73B parameters, needing 45.1 GB at Q4_K_M.
How much of the NVIDIA RTX A6000's memory can a model actually use?
About 48 GB of its 48 GB, before the desktop and runtime overhead are accounted for.
Similar GPUs
- Best models for the NVIDIA RTX 6000 Ada Generation
- Best models for the NVIDIA L40S
- Best models for the NVIDIA RTX PRO 5000 Blackwell (48 GB)
- Best models for the NVIDIA L40
By Workload
- Best local LLMs for coding
- Best local LLMs for general assistant
- Best local LLMs for reasoning
- Best local LLMs for rag