NVIDIADesktop GPUCurrent

NVIDIA L4 for local LLMs

Written by Jakub Rusinowski · Last updated

With 24 GB of GDDR6 at 300 GB/s, the L4 runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Qwen 3 32B (~22.8 GB), and the top pick is Laguna XS 2.1 33B-A3B at 44–85 tok/s.

24 GB of GDDR6 at 300 GB/s in a 72 W, single-slot, low-profile card. Plenty of memory, little bandwidth. Common in cloud instances.

Models that run on the L4

Q4_K_M, 8K context, 24 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB44–85 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB24–46 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB42–80 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB56–107 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB56–107 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB45–87 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB45–87 tok/s
North Mini Code 1.0 30B-A3B
North Mini Code
18.9 GB45–87 tok/s
Nemotron 3.5 Lightning 30B-A3B
Nemotron 3.5
18.9 GB45–87 tok/s
Trinity Mini
Trinity
16.9 GB74–142 tok/s
Gemma 4 26B-A4B
Gemma 4
16.5 GB40–77 tok/s
GPT-OSS 20B
GPT-OSS
13.8 GB59–114 tok/s
Showing 12 of 109

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
24 GB GDDR6
Memory bandwidth
300 GB/s
Memory bus
192-bit
Architecture
Ada Lovelace AD104
Series
Data center
Board power
72 W
Release year
2023
Compute backends
CUDA
Usable for models
24 GB
Best forrented 24 GBlow power inference

Similar GPUs

Frequently asked questions

Can the NVIDIA L4 run local LLMs?

Yes. With 24 GB (24 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Qwen 3 32B, needing about 22.8 GB.

How fast is the NVIDIA L4 for AI inference?

It is estimated to run Llama 3.1 8B at 32–61 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 24 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 24 GB?

Among the best that fit: Laguna XS 2.1 33B-A3B, Granite 4.0 Small-H 32B-A9B, Nemotron-Cascade 2 30B-A3B, Qwen 3 30B-A3B (MoE), Qwen3-Coder 30B-A3B (MoE).