NVIDIA GeForce RTX 2080 Ti for local LLMs
Written by Jakub Rusinowski · Last updated
With 11 GB of GDDR6 at 616 GB/s, the RTX 2080 Ti runs 65 catalogued models at Q4_K_M with 8K context. The largest that fits is Phi-4 (14B) (~10.9 GB), and the top pick is Ministral 3 14B.
11 GB of GDDR6 on a 352-bit bus, 616 GB/s. Turing: FP16 tensor cores but no BF16 and no FP8. The site has no speed calibration for Turing, so this card gets fit verdicts but no tok/s estimate.
Models that run on the RTX 2080 Ti
Q4_K_M, 8K context, 11 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 11 GB | VRAM | Speed |
|---|---|---|---|
| Ministral 3 14B Ministral 3 | 9.3 GB | — | |
| Gemma 4 12B (Unified) Gemma 4 | 8 GB | — | |
| Mistral NeMo 12B Mistral Family | 9.4 GB | — | |
| Bielik PL 11B v3.0 Instruct Bielik | 7.4 GB | — | |
| Llama 3.2 Vision 11B Llama 3.2 Vision | 8.5 GB | — | |
| Llama 3.2 11B Vision Instruct Llama 3.2 Family | 8.5 GB | — | |
| Falcon 3 10B Instruct Falcon 3 | 8.4 GB | — | |
| Qwen 3.5 9B Qwen 3.5 | 6.2 GB | — | |
| GLM-4 9B GLM-4.7 / GLM-Z1 | 6.2 GB | — | |
| GLM-4.6V-Flash 9B GLM-4.6V | 6.2 GB | — | |
| EuroLLM 9B EuroLLM | 6.2 GB | — | |
| InternLM 3 8B Instruct InternLM 3 | 6.5 GB | — |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 11 GB GDDR6
- Memory bandwidth
- 616 GB/s
- Memory bus
- 352-bit
- Architecture
- Turing TU102
- Series
- RTX 20-series
- Release year
- 2018
- Launch price
- $999
- Compute backends
- CUDA, VULKAN
- Usable for models
- 11 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA GeForce RTX 2080 Ti run local LLMs?
Yes. With 11 GB (11 GB usable by a model) it runs 65 of the catalogued models at Q4_K_M with 8K context; the largest is Phi-4 (14B), needing about 10.9 GB.
How fast is the NVIDIA GeForce RTX 2080 Ti for AI inference?
This site has no speed calibration for the Turing TU102 architecture, so it publishes fit verdicts for this card but no tokens-per-second estimate. Llama 3.3 70B does not fit: it needs about 44 GB against 11 GB usable.
What LLMs can I run on 11 GB?
Among the best that fit: Ministral 3 14B, Gemma 4 12B (Unified), Mistral NeMo 12B, Bielik PL 11B v3.0 Instruct, Llama 3.2 Vision 11B.