NVIDIA GeForce RTX 3080 Ti Laptop GPU for local LLMs
Written by Jakub Rusinowski · Last updated
With 16 GB of GDDR6 at 512 GB/s, the RTX 3080 Ti Laptop GPU runs 74 catalogued models at Q4_K_M with 8K context. The largest that fits is OLMo 2 13B Instruct (~15.8 GB), and the top pick is EuroLLM 22B at 14–27 tok/s.
16 GB of GDDR6, 512 GB/s: the 16 GB laptop GPU most often found used. Not the 12 GB desktop RTX 3080 Ti. Power depends on the chassis and is not stored.
Models that run on the RTX 3080 Ti Laptop GPU
Q4_K_M, 8K context, 16 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 16 GB | VRAM | Speed |
|---|---|---|---|
| EuroLLM 22B EuroLLM | 14.1 GB | 14–27 tok/s | |
| GPT-OSS 20B GPT-OSS | 13.8 GB | 66–127 tok/s | |
| InternLM 3 20B Instruct InternLM 3 | 12.9 GB | 15–29 tok/s | |
| Cosmos 3 Nano Cosmos 3 | 10.5 GB | 31–60 tok/s | |
| StarCoder 2 15B StarCoder 2 | 10.8 GB | 20–39 tok/s | |
| Qwen 3 14B Qwen 3 | 11.1 GB | 20–38 tok/s | |
| Phi-4 (14B) Phi-4 Family | 10.9 GB | 20–39 tok/s | |
| Cogito v1 14B Cogito v1 | 10.9 GB | 20–39 tok/s | |
| Ministral 3 14B Ministral 3 | 9.3 GB | 21–40 tok/s | |
| DeepSeek R1 Distill Qwen 14B DeepSeek R1 | 10.6 GB | 21–40 tok/s | |
| Gemma 4 12B (Unified) Gemma 4 | 8 GB | 24–45 tok/s | |
| Bielik PL 11B v3.0 Instruct Bielik | 7.4 GB | 25–49 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 16 GB GDDR6
- Memory bandwidth
- 512 GB/s
- Memory bus
- 256-bit
- Architecture
- Ampere GA103
- Series
- RTX 30-series (Laptop)
- Release year
- 2022
- Compute backends
- CUDA
- Usable for models
- 16 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA GeForce RTX 3080 Ti Laptop GPU run local LLMs?
Yes. With 16 GB (16 GB usable by a model) it runs 74 of the catalogued models at Q4_K_M with 8K context; the largest is OLMo 2 13B Instruct, needing about 15.8 GB.
How fast is the NVIDIA GeForce RTX 3080 Ti Laptop GPU for AI inference?
It is estimated to run Llama 3.1 8B at 36–69 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 16 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 16 GB?
Among the best that fit: EuroLLM 22B, GPT-OSS 20B, InternLM 3 20B Instruct, Cosmos 3 Nano, StarCoder 2 15B.