NVIDIA GeForce RTX 5080 Laptop GPU for local LLMs
Written by Jakub Rusinowski · Last updated
With 16 GB of GDDR7 at 896 GB/s, the RTX 5080 Laptop GPU runs 74 catalogued models at Q4_K_M with 8K context. The largest that fits is OLMo 2 13B Instruct (~15.8 GB), and the top pick is EuroLLM 22B at 30–62 tok/s.
16 GB of GDDR7 on a 256-bit bus, 896 GB/s. Not the desktop RTX 5080 (960 GB/s, 360 W). 80-150 W by chassis.
Models that run on the RTX 5080 Laptop GPU
Q4_K_M, 8K context, 16 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 16 GB | VRAM | Speed |
|---|---|---|---|
| EuroLLM 22B EuroLLM | 14.1 GB | 30–62 tok/s | |
| GPT-OSS 20B GPT-OSS | 13.8 GB | 115–240 tok/s | |
| InternLM 3 20B Instruct InternLM 3 | 12.9 GB | 32–67 tok/s | |
| Cosmos 3 Nano Cosmos 3 | 10.5 GB | 62–129 tok/s | |
| StarCoder 2 15B StarCoder 2 | 10.8 GB | 42–88 tok/s | |
| Qwen 3 14B Qwen 3 | 11.1 GB | 41–86 tok/s | |
| Phi-4 (14B) Phi-4 Family | 10.9 GB | 42–87 tok/s | |
| Cogito v1 14B Cogito v1 | 10.9 GB | 42–87 tok/s | |
| Ministral 3 14B Ministral 3 | 9.3 GB | 43–89 tok/s | |
| DeepSeek R1 Distill Qwen 14B DeepSeek R1 | 10.6 GB | 43–89 tok/s | |
| Gemma 4 12B (Unified) Gemma 4 | 8 GB | 48–100 tok/s | |
| Bielik PL 11B v3.0 Instruct Bielik | 7.4 GB | 51–107 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 16 GB GDDR7
- Memory bandwidth
- 896 GB/s
- Memory bus
- 256-bit
- Architecture
- Blackwell GB203
- Series
- RTX 50-series (Laptop)
- Board power
- 150 W
- Laptop TGP range
- 80–150 W
- Release year
- 2025
- Compute backends
- CUDA
- Usable for models
- 16 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA GeForce RTX 5080 Laptop GPU run local LLMs?
Yes. With 16 GB (16 GB usable by a model) it runs 74 of the catalogued models at Q4_K_M with 8K context; the largest is OLMo 2 13B Instruct, needing about 15.8 GB.
How fast is the NVIDIA GeForce RTX 5080 Laptop GPU for AI inference?
It is estimated to run Llama 3.1 8B at 70–145 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 16 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 16 GB?
Among the best that fit: EuroLLM 22B, GPT-OSS 20B, InternLM 3 20B Instruct, Cosmos 3 Nano, StarCoder 2 15B.