GPT-OSS 20B — VRAM, speed & local setup
Written by Jakub Rusinowski · Last updated
The size of the gpt-oss pair most people actually run. Shipped in MXFP4 precision — the MoE weights are quantized to ~4.25 bits per parameter — so it loads in roughly 14 GB and runs comfortably on a 16GB card, far below what a 20B parameter count would normally imply. Ollama supports MXFP4 natively, with no extra conversion step. 128K context, Apache 2.0, and o3-mini-class reasoning.
GPT-OSS 20B needs about 13 GB of VRAM at MXFP4 — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
VRAM and speed by quantization
Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.
| Quant | Memory | VRAM | Speed (est.) | Fit |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 8.1 GB | ~245 tok/s | Fits | |
| Q3_K_M 3.41 bpw | 10.1 GB | ~220 tok/s | Fits | |
| Q4_K_M 4.83 bpw | 13.8 GB | ~186 tok/s | Fits | |
| Q5_K_M 5.67 bpw | 16 GB | ~170 tok/s | Fits | |
| Q6_K 6.56 bpw | 18.3 GB | ~156 tok/s | Fits | |
| Q8_0 8.50 bpw | 23.4 GB | ~132 tok/s | Tight | |
| F16 16.00 bpw | 43 GB | ~12 tok/s | Offload |
Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. GPT-OSS 20B VRAM calculator →
Get GPT-OSS 20B running
The cheapest catalogued GPU that runs GPT-OSS 20B is the AMD Radeon RX 9060 XT 16GB (16 GB).
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
How to run GPT-OSS 20B
Install Ollama, then run:
ollama run gpt-oss:20bSpecifications
Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.
Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.
- Parameters
- 20 Billion
- Context window
- 128,000
- Architecture
- Mixture-of-Experts (MXFP4)
- Provider
- OpenAI
- Licence
- Apache 2.0
- Specified at
- MXFP4
- System RAM
- 32 GB
- Record updated
- 2026-08-15
Commercial use permitted. No usage restrictions beyond attribution.
Quality and use cases
Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.
Can I run GPT-OSS 20B on my GPU?
- GPT-OSS 20B on AMD Radeon RX 7800 XT
- GPT-OSS 20B on AMD Radeon RX 7900 XT
- GPT-OSS 20B on AMD Radeon RX 7900 XTX
- GPT-OSS 20B on AMD Radeon RX 9060 XT 16GB
- GPT-OSS 20B on AMD Radeon RX 9060 XT 8GB
- GPT-OSS 20B on AMD Radeon RX 9070
- GPT-OSS 20B on AMD Radeon RX 9070 XT
- GPT-OSS 20B on AMD Ryzen AI Max+ 395
- GPT-OSS 20B on Apple M1
- GPT-OSS 20B on Apple M1 Ultra
- GPT-OSS 20B on Apple M2
- GPT-OSS 20B on Apple M2 Max
- GPT-OSS 20B on Apple M3
- GPT-OSS 20B on Apple M4 Max
- GPT-OSS 20B on Apple M4 Pro
- GPT-OSS 20B on Apple M5 Max
- GPT-OSS 20B on Intel Arc B570
- GPT-OSS 20B on Intel Arc B580
- GPT-OSS 20B on NVIDIA A100 80GB (PCIe)
- GPT-OSS 20B on NVIDIA DGX Spark
- GPT-OSS 20B on NVIDIA GeForce RTX 3060 (12GB)
- GPT-OSS 20B on NVIDIA GeForce RTX 3070
- GPT-OSS 20B on NVIDIA GeForce RTX 3070 Ti
- GPT-OSS 20B on NVIDIA GeForce RTX 3080 (10GB)
- GPT-OSS 20B on NVIDIA GeForce RTX 3090
- GPT-OSS 20B on NVIDIA GeForce RTX 4060
- GPT-OSS 20B on NVIDIA GeForce RTX 4060 Ti 16GB
- GPT-OSS 20B on NVIDIA GeForce RTX 4070
- GPT-OSS 20B on NVIDIA GeForce RTX 4070 Super
- GPT-OSS 20B on NVIDIA GeForce RTX 4070 Ti
- GPT-OSS 20B on NVIDIA GeForce RTX 4070 Ti Super
- GPT-OSS 20B on NVIDIA GeForce RTX 4080
- GPT-OSS 20B on NVIDIA GeForce RTX 4080 Super
- GPT-OSS 20B on NVIDIA GeForce RTX 4090
- GPT-OSS 20B on NVIDIA GeForce RTX 5060
- GPT-OSS 20B on NVIDIA GeForce RTX 5060 Ti 16GB
- GPT-OSS 20B on NVIDIA GeForce RTX 5060 Ti 8GB
- GPT-OSS 20B on NVIDIA GeForce RTX 5070
- GPT-OSS 20B on NVIDIA GeForce RTX 5070 Ti
- GPT-OSS 20B on NVIDIA GeForce RTX 5080
- GPT-OSS 20B on NVIDIA H100 80GB (PCIe)
- GPT-OSS 20B on NVIDIA RTX PRO 6000 Blackwell
Other GPT-OSS sizes
GPT-OSS 20B — frequently asked questions
How much VRAM does GPT-OSS 20B need?
About 13 GB at MXFP4 — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does GPT-OSS 20B run on an RTX 4090 (24 GB)?
Yes. GPT-OSS 20B needs about 13 GB at MXFP4, inside a 24 GB card, at an estimated 186 tokens/sec.
How do I run GPT-OSS 20B locally?
Install Ollama and run `ollama run gpt-oss:20b`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.
What other sizes does GPT-OSS come in?
GPT-OSS 120B (71 GB), GPT-OSS 20B (13 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.