OpenAI20B~13 GB VRAM at MXFP4

GPT-OSS 20B — VRAM, speed & local setup

Written by Jakub Rusinowski · Last updated

The size of the gpt-oss pair most people actually run. Shipped in MXFP4 precision — the MoE weights are quantized to ~4.25 bits per parameter — so it loads in roughly 14 GB and runs comfortably on a 16GB card, far below what a 20B parameter count would normally imply. Ollama supports MXFP4 natively, with no extra conversion step. 128K context, Apache 2.0, and o3-mini-class reasoning.

GPT-OSS 20B needs about 13 GB of VRAM at MXFP4 — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Logic84
Creative80
Coding82

VRAM and speed by quantization

Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.

QuantVRAMSpeed (est.)Fit
Q2_K
2.63 bpw
8.1 GB~245 tok/sFits
Q3_K_M
3.41 bpw
10.1 GB~220 tok/sFits
Q4_K_M
4.83 bpw
13.8 GB~186 tok/sFits
Q5_K_M
5.67 bpw
16 GB~170 tok/sFits
Q6_K
6.56 bpw
18.3 GB~156 tok/sFits
Q8_0
8.50 bpw
23.4 GB~132 tok/sTight
F16
16.00 bpw
43 GB~12 tok/sOffload

Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. GPT-OSS 20B VRAM calculator →

Get GPT-OSS 20B running

The cheapest catalogued GPU that runs GPT-OSS 20B is the AMD Radeon RX 9060 XT 16GB (16 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 9060 XT 16GB
16 GB VRAM · 160 W board power
2026 prices are volatile — check the current listing.

How to run GPT-OSS 20B

Install Ollama, then run:

ollama run gpt-oss:20b
Weights on Hugging Face: openai/gpt-oss-20b ↗

Specifications

Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.

Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.

Parameters
20 Billion
Context window
128,000
Architecture
Mixture-of-Experts (MXFP4)
Provider
OpenAI
Licence
Apache 2.0
Specified at
MXFP4
System RAM
32 GB
Record updated
2026-08-15
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Quality and use cases

Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.

Best forreasoninggeneral purposelocal firstprivacy sensitive

Can I run GPT-OSS 20B on my GPU?

Other GPT-OSS sizes

GPT-OSS 20B — frequently asked questions

How much VRAM does GPT-OSS 20B need?

About 13 GB at MXFP4 — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does GPT-OSS 20B run on an RTX 4090 (24 GB)?

Yes. GPT-OSS 20B needs about 13 GB at MXFP4, inside a 24 GB card, at an estimated 186 tokens/sec.

How do I run GPT-OSS 20B locally?

Install Ollama and run `ollama run gpt-oss:20b`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.

What other sizes does GPT-OSS come in?

GPT-OSS 120B (71 GB), GPT-OSS 20B (13 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.