Best Local LLMs for the NVIDIA GeForce RTX 5090

Written by Jakub Rusinowski · Last updated July 12, 2026

Best all-round pick: Gemma 4 27B ⭐

The NVIDIA GeForce RTX 5090 has 32 GB of VRAM, of which about 32 GB is available to a model. The largest model it can hold is Command R (35B) (35B, 21.9 GB at Q4_K_M).

Best models overall on the NVIDIA GeForce RTX 5090

ModelScoreMemory at Q4_K_MContextLicence
Gemma 4 27B ⭐96.517.1 GB125KGemma License (commercial OK)
Qwen 3.7 35B-A3B95.221.9 GB256KApache-2.0
Qwen 3 32B95.120.6 GB125KApache 2.0
Gemma 4 31B94.819.5 GB250KApache-2.0
Qwen 3.6 35B-A3B94.621.9 GB256KApache-2.0
Kimi K2.694.320.1 GB125KKimi License (research)

Best model by what you are doing

WorkloadRecommended on this GPU
CodingKimi K2.5 (99.7)
Qwen 3.7 35B-A3B (98.5)
Qwen 3 32B (98)
General assistantGemma 4 27B ⭐ (96.5)
Qwen 3.7 35B-A3B (95.2)
Qwen 3 32B (95.1)
ReasoningGemma 4 27B ⭐ (99.6)
GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (99.2)
Qwen 3 32B (98.3)
RAGGemma 4 31B (97.5)
Qwen 3.6 27B (96.2)
Qwen 3.7 35B-A3B (95.2)
AgentsTernary Bonsai 27B (93.3)
Poolside Laguna XS 2.1 Laguna XS 2.1 33B-A3B (90.4)
GLM-4.7-Flash 30B-A3B (90.3)
VisionQwen 3.7 35B-A3B (98.1)
Gemma 4 31B (97.8)
Qwen 3.6 35B-A3B (97.4)

What the NVIDIA GeForce RTX 5090 cannot run

These models are strong picks generally but exceed the 32 GB this card makes available.

ModelNeeds at Q4_K_MShort by
Nemotron 70B Instruct43.5 GB~11.5 GB
Llama 3.3 70B Instruct43.1 GB~11.1 GB

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA GeForce RTX 5090?

Gemma 4 27B ⭐ is the strongest all-round pick that fits its 32 GB.

What is the largest model the NVIDIA GeForce RTX 5090 can run?

Command R (35B) — 35B parameters, needing 21.9 GB at Q4_K_M.

How much of the NVIDIA GeForce RTX 5090's memory can a model actually use?

About 32 GB of its 32 GB, before the desktop and runtime overhead are accounted for.

Compatibility Checks for the NVIDIA GeForce RTX 5090

Similar GPUs

By Workload

More

← All GPUs | NVIDIA GeForce RTX 5090 specs