NVIDIA Tesla P40 for local LLMs
Written by Jakub Rusinowski · Last updated
With 24 GB of GDDR5 at 346 GB/s, the Tesla P40 runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Qwen 3 32B (~22.8 GB), and the top pick is Laguna XS 2.1 33B-A3B.
24 GB of GDDR5, about 346 GB/s (the two references give 345.6 and 347.1), 250 W. Passively cooled with no display output, so it needs server airflow or an aftermarket fan. Pascal: no tensor cores and very slow FP16. No speed calibration for Pascal: fit verdicts only.
Models that run on the Tesla P40
Q4_K_M, 8K context, 24 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 24 GB | VRAM | Speed |
|---|---|---|---|
| Laguna XS 2.1 33B-A3B Poolside Laguna XS 2.1 | 20.7 GB | — | |
| Granite 4.0 Small-H 32B-A9B IBM Granite 4.0 | 20.1 GB | — | |
| Nemotron-Cascade 2 30B-A3B Nemotron Cascade 2 | 19.9 GB | — | |
| Qwen 3 30B-A3B (MoE) Qwen 3 | 20 GB | — | |
| Qwen3-Coder 30B-A3B (MoE) Qwen3-Coder | 20 GB | — | |
| GLM-4.7-Flash 30B-A3B GLM-4.7 / GLM-Z1 | 18.9 GB | — | |
| Nemotron 3 Nano Omni 30B-A3B Nemotron 3 Nano Omni | 18.9 GB | — | |
| North Mini Code 1.0 30B-A3B North Mini Code | 18.9 GB | — | |
| Nemotron 3.5 Lightning 30B-A3B Nemotron 3.5 | 18.9 GB | — | |
| Trinity Mini Trinity | 16.9 GB | — | |
| Gemma 4 26B-A4B Gemma 4 | 16.5 GB | — | |
| Devstral Small 2505 24B Devstral | 16.6 GB | — |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 24 GB GDDR5
- Memory bandwidth
- 346 GB/s
- Memory bus
- 384-bit
- Architecture
- Pascal GP102
- Series
- Tesla
- Board power
- 250 W
- Release year
- 2016
- Compute backends
- CUDA, VULKAN
- Usable for models
- 24 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA Tesla P40 run local LLMs?
Yes. With 24 GB (24 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Qwen 3 32B, needing about 22.8 GB.
How fast is the NVIDIA Tesla P40 for AI inference?
This site has no speed calibration for the Pascal GP102 architecture, so it publishes fit verdicts for this card but no tokens-per-second estimate. Llama 3.3 70B does not fit: it needs about 44 GB against 24 GB usable.
What LLMs can I run on 24 GB?
Among the best that fit: Laguna XS 2.1 33B-A3B, Granite 4.0 Small-H 32B-A9B, Nemotron-Cascade 2 30B-A3B, Qwen 3 30B-A3B (MoE), Qwen3-Coder 30B-A3B (MoE).