RTX 5080 16GB runs GPT-OSS 20B — but there's less headroom than you'd think

Sized at Q4_K_M with a 8K context: weights, KV cache and framework overhead, against what RTX 5080 16GB leaves free. GPT-OSS 20B at Q4_K_M needs 13.3 GB and you have 16 GB available, so it fits with 2.7 GB to spare. Expect 37.1–77 tok/s (estimated).

What to do

Your machine already runs this

GPT-OSS 20B at Q4_K_M needs 13.3 GB and you have 16 GB available, so it fits with 2.7 GB to spare.

What we checked
  • VRAM required at Q4_K_M and 8K context: 13.3 GB
  • VRAM available on your hardware: 16 GB

Your machine is not exactly this one. Run this for your exact setup — the form opens pre-filled with RTX 5080 16GB and GPT-OSS 20B.

Questions people ask about this pairing

Can an RTX 5080 16GB run GPT-OSS 20B?

Yes. GPT-OSS 20B at Q4_K_M needs 13.3 GB and you have 16 GB available, so it fits with 2.7 GB to spare.

How fast will GPT-OSS 20B actually run?

Around 37.1–77 tok/s (estimated) on this setup at Q4_K_M. Decode speed on a local model is set by memory bandwidth rather than compute, so the figure moves with the card's bandwidth and the bytes read per token — not with its price.

Does a longer context change what GPT-OSS 20B needs here?

Yes, and it is the figure people forget. The 13.3 GB above already includes the KV cache at 8K; that cache grows roughly linearly with context, so doubling the window adds real gigabytes rather than a rounding error. If you plan to work with long documents on RTX 5080 16GB, size for the context you will actually use, not the default.

Related

Data behind this page last checked 2026-08-15.