Best Local LLMs for the NVIDIA L40S

Written by Jakub Rusinowski · Last updated July 21, 2026

Best all-round pick: Nemotron 70B Instruct

The NVIDIA L40S has 48 GB of VRAM, of which about 48 GB is available to a model. The largest model it can hold is Qwen 2.5 VL 72B Instruct (73B, 45.1 GB at Q4_K_M).

Best models overall on the NVIDIA L40S

ModelScoreMemory at Q4_K_MContextLicence
Nemotron 70B Instruct98.343.4 GB125KLlama Community
GLM-5.1 72B94.744.3 GB125KMIT
Qwen 3.5 72B94.444.3 GB125KApache 2.0
Gemma 4 27B ⭐9417.1 GB125KGemma License (commercial OK)
Llama 3.3 70B Instruct93.843.1 GB128KLlama Community
Qwen 2.5 72B Instruct93.344.3 GB128KApache-2.0

Best model by what you are doing

WorkloadRecommended on this GPU
CodingKimi K2.5 (97.7)
Qwen 2.5 72B Instruct (97.6)
Qwen 3.7 35B-A3B (96.9)
General assistantNemotron 70B Instruct (98.3)
GLM-5.1 72B (94.7)
Qwen 3.5 72B (94.4)
ReasoningGLM-5.1 72B (100)
Gemma 4 27B ⭐ (98.1)
GLM-4.7 / GLM-Z1 GLM-Z1 32B (Reasoning) (97.7)
RAGGemma 4 31B (95.6)
Qwen 3.6 27B (94.5)
Qwen 3.7 35B-A3B (93.8)
AgentsTernary Bonsai 27B (91.6)
Poolside Laguna XS 2.1 Laguna XS 2.1 33B-A3B (88.7)
Kimi K2.5 (88.5)
VisionQwen 3.5 72B (99.7)
Qwen 3.7 35B-A3B (96.5)
Qwen 3.6 35B-A3B (95.8)

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA L40S?

Nemotron 70B Instruct is the strongest all-round pick that fits its 48 GB.

What is the largest model the NVIDIA L40S can run?

Qwen 2.5 VL 72B Instruct — 73B parameters, needing 45.1 GB at Q4_K_M.

How much of the NVIDIA L40S's memory can a model actually use?

About 48 GB of its 48 GB, before the desktop and runtime overhead are accounted for.

Compatibility Checks for the NVIDIA L40S

Similar GPUs

By Workload

More

← All GPUs | NVIDIA L40S specs