NVIDIADesktop GPUPrevious generation

NVIDIA GeForce GTX 1080 Ti for local LLMs

Written by Jakub Rusinowski · Last updated

With 11 GB of GDDR5X at 484 GB/s, the GTX 1080 Ti runs 65 catalogued models at Q4_K_M with 8K context. The largest that fits is Phi-4 (14B) (~10.9 GB), and the top pick is Ministral 3 14B.

11 GB of GDDR5X, 484 GB/s. Pascal has no tensor cores and very slow FP16, so it runs quantized models through llama.cpp but nothing that needs modern kernels. No speed calibration for Pascal: fit verdicts only.

Models that run on the GTX 1080 Ti

Q4_K_M, 8K context, 11 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Ministral 3 14B
Ministral 3
9.3 GB—
Gemma 4 12B (Unified)
Gemma 4
8 GB—
Mistral NeMo 12B
Mistral Family
9.4 GB—
Bielik PL 11B v3.0 Instruct
Bielik
7.4 GB—
Llama 3.2 Vision 11B
Llama 3.2 Vision
8.5 GB—
Llama 3.2 11B Vision Instruct
Llama 3.2 Family
8.5 GB—
Falcon 3 10B Instruct
Falcon 3
8.4 GB—
Qwen 3.5 9B
Qwen 3.5
6.2 GB—
GLM-4 9B
GLM-4.7 / GLM-Z1
6.2 GB—
GLM-4.6V-Flash 9B
GLM-4.6V
6.2 GB—
EuroLLM 9B
EuroLLM
6.2 GB—
InternLM 3 8B Instruct
InternLM 3
6.5 GB—
Showing 12 of 65

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
11 GB GDDR5X
Memory bandwidth
484 GB/s
Memory bus
352-bit
Architecture
Pascal GP102
Series
GTX 10-series
Board power
250 W
Release year
2017
Launch price
$699
Compute backends
CUDA, VULKAN
Usable for models
11 GB
Best forused 11 GBlegacy rigs

Similar GPUs

Frequently asked questions

Can the NVIDIA GeForce GTX 1080 Ti run local LLMs?

Yes. With 11 GB (11 GB usable by a model) it runs 65 of the catalogued models at Q4_K_M with 8K context; the largest is Phi-4 (14B), needing about 10.9 GB.

How fast is the NVIDIA GeForce GTX 1080 Ti for AI inference?

This site has no speed calibration for the Pascal GP102 architecture, so it publishes fit verdicts for this card but no tokens-per-second estimate. Llama 3.3 70B does not fit: it needs about 44 GB against 11 GB usable.

What LLMs can I run on 11 GB?

Among the best that fit: Ministral 3 14B, Gemma 4 12B (Unified), Mistral NeMo 12B, Bielik PL 11B v3.0 Instruct, Llama 3.2 Vision 11B.