Best Local LLMs for the NVIDIA GeForce RTX 5090
Written by Jakub Rusinowski · Last updated September 29, 2026
Best all-round pick: Gemma 4 27B ⭐
The NVIDIA GeForce RTX 5090 has 32 GB of VRAM, of which about 32 GB is available to a model. The largest model it can hold is K2 Horizon MoVA 36B-A4B (37B, 23.4 GB at Q4_K_M).
Best models overall on the NVIDIA GeForce RTX 5090
| Model | Score | Memory at Q4_K_M | Context | Licence |
|---|---|---|---|---|
| Gemma 4 27B ⭐ | 96.5 | 17.1 GB | 125K | Gemma License (commercial OK) |
| Qwen3.8 27B | 95.2 | 17.6 GB | 256K | Apache-2.0 |
| Qwen 3.7 35B-A3B | 95.2 | 21.9 GB | 256K | Apache-2.0 |
| Qwen 3 32B | 95.1 | 20.6 GB | 125K | Apache 2.0 |
| Gemma 4 31B | 94.8 | 19.5 GB | 250K | Apache-2.0 |
| Qwen 3.6 35B-A3B | 94.6 | 21.9 GB | 256K | Apache-2.0 |
Best model by what you are doing
| Workload | Recommended on this GPU |
|---|---|
| Coding | Qwen 3.7 35B-A3B (98.5) Qwen 3 32B (98) Qwen 3.6 35B-A3B (97.8) |
| General assistant | Gemma 4 27B ⭐ (96.5) Qwen3.8 27B (95.2) Qwen 3.7 35B-A3B (95.2) |
| Reasoning | Gemma 4 27B ⭐ (99.6) GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (99.2) Qwen 3 32B (98.3) |
| RAG | Gemma 4 31B (97.5) Qwen 3.6 27B (96.2) Qwen3.8 27B (95.4) |
| Agents | Ternary Bonsai 27B (93.3) Nemotron-Cascade 2 30B-A3B (92.8) Nemotron 3.5 Lightning 30B-A3B (92.5) |
| Vision | Qwen 3.7 35B-A3B (98.1) Gemma 4 31B (97.8) Qwen 3.6 35B-A3B (97.4) |
What the NVIDIA GeForce RTX 5090 cannot run
These models are strong picks generally but exceed the 32 GB this card makes available.
| Model | Needs at Q4_K_M | Short by |
|---|---|---|
| Llama 3.3 70B Instruct | 43.1 GB | ~11.1 GB |
How these numbers are calculated
- Usable memory is 32 GB of the card's 32 GB.
- Fit is measured at Q4_K_M — quantized weights plus KV cache plus 0.8 GB runtime overhead.
- Within this card's budget, models are ranked on how well they use the memory available, not on being small — so the recommendation scales with the hardware.
FAQ
What is the best LLM for the NVIDIA GeForce RTX 5090?
Gemma 4 27B ⭐ is the strongest all-round pick that fits its 32 GB.
What is the largest model the NVIDIA GeForce RTX 5090 can run?
K2 Horizon MoVA 36B-A4B — 37B parameters, needing 23.4 GB at Q4_K_M.
How much of the NVIDIA GeForce RTX 5090's memory can a model actually use?
About 32 GB of its 32 GB, before the desktop and runtime overhead are accounted for.
Compatibility Checks for the NVIDIA GeForce RTX 5090
- Can I run DeepSeek R1 on the NVIDIA GeForce RTX 5090?
- Can I run Llama 3.3 on the NVIDIA GeForce RTX 5090?
- Can I run Nemotron 70B on the NVIDIA GeForce RTX 5090?
- Can I run Command R Family on the NVIDIA GeForce RTX 5090?
- Can I run Qwen 2.5 Family on the NVIDIA GeForce RTX 5090?
Similar GPUs
- Best models for the NVIDIA RTX PRO 4500 Blackwell (Workstation Edition)
- Best models for the NVIDIA RTX 5000 Ada Generation
- Best models for the NVIDIA GeForce RTX 5090 Laptop GPU
- Best models for the NVIDIA GeForce RTX 3090 Ti
By Workload
- Best local LLMs for coding
- Best local LLMs for general assistant
- Best local LLMs for reasoning
- Best local LLMs for rag