NVIDIA H100 80GB (SXM) for local LLMs
Written by Jakub Rusinowski · Last updated
With 80 GB of HBM3 at 3,350 GB/s, the H100 80GB (SXM) runs 127 catalogued models at Q4_K_M with 8K context. The largest that fits is Devstral-2 123B (~75.1 GB), and the top pick is GPT-OSS 120B at 184–381 tok/s.
The SXM module: 80 GB of HBM3 at 3.35 TB/s, up to 700 W. The catalog's other H100 record is the PCIe card at 2,000 GB/s.
Models that run on the H100 80GB (SXM)
Q4_K_M, 8K context, 80 GB usable. Ranked by quality and speed.
| Model | Memory · marker = 80 GB | VRAM | Speed |
|---|---|---|---|
| GPT-OSS 120B GPT-OSS | 71.9 GB | 184–381 tok/s | |
| Llama 4.5 Scout Llama 4.5 | 66.6 GB | 95–198 tok/s | |
| Llama 4 Scout 17B Llama 4 | 68.2 GB | 101–210 tok/s | |
| Llama 3.2 Vision 90B Llama 3.2 Vision | 57.8 GB | 29–61 tok/s | |
| Qwen3-Coder-Next (80B-A3B MoE) Qwen3-Coder | 49.1 GB | 173–360 tok/s | |
| Kolibri 1 Kolibri | 48.4 GB | 214–444 tok/s | |
| Llama 3.3 70B Instruct Llama 3.3 | 45.7 GB | 36–75 tok/s | |
| Cogito v1 70B Cogito v1 | 45.7 GB | 36–75 tok/s | |
| Apertus 70B Apertus | 43.1 GB | 36–75 tok/s | |
| Cosmos 3 Super Cosmos 3 | 39.4 GB | 66–137 tok/s | |
| Qwen 3.6 35B-A3B Qwen 3.6 | 21.9 GB | 184–381 tok/s | |
| Nex-N2.5 mini Nex-N2.5 | 21.9 GB | 184–381 tok/s |
Buy it or rent the same memory
Buy the card, or rent a GPU with the same memory by the hour to try models first.
or compare on Vast.ai from $0.77/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Speed vs other GPUs
Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated
Specifications
Specs last updated 2026-10-07.
- Memory
- 80 GB HBM3
- Memory bandwidth
- 3,350 GB/s
- Memory bus
- 5120-bit
- Architecture
- Hopper GH100
- Series
- Data center
- Board power
- 700 W
- Release year
- 2022
- Compute backends
- CUDA
- Usable for models
- 80 GB
Similar GPUs
Frequently asked questions
Can the NVIDIA H100 80GB (SXM) run local LLMs?
Yes. With 80 GB (80 GB usable by a model) it runs 127 of the catalogued models at Q4_K_M with 8K context; the largest is Devstral-2 123B, needing about 75.1 GB.
How fast is the NVIDIA H100 80GB (SXM) for AI inference?
It is estimated to run Llama 3.1 8B at 157–327 tok/s at Q4_K_M. Llama 3.3 70B is estimated at 37–77 tok/s. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.
What LLMs can I run on 80 GB?
Among the best that fit: GPT-OSS 120B, Llama 4.5 Scout, Llama 4 Scout 17B, Llama 3.2 Vision 90B, Qwen3-Coder-Next (80B-A3B MoE).