RTX 5080 16GB can't get to Llama 4 Scout 17B at any price. Here's what actually can.
Llama 4 Scout 17B at Q4_K_M needs 68.2 GB once weights, KV cache at 8K context and framework overhead are counted. RTX 5080 16GB offers 16 GB, leaving you 52.2 GB short. 128 GB of unified memory leaves about 96 GB for a model after the OS takes its share, which holds Llama 4 Scout 17B at Q4_K_M.
What to do
128 GB of unified memory leaves about 96 GB for a model after the OS takes its share, which holds Llama 4 Scout 17B at Q4_K_M.
- Memory is fixed at purchase on these machines — choose the capacity you will still want in three years
- Unified memory: 128 GB, of which roughly 96 GB is usable by a model
- Model requires 68.2 GB at Q4_K_M
Other ways to get there
- Add 96GB DDR5 kit (2x48GB) — You are 52.2 GB short in VRAM, and 96 GB of system RAM gives llama.cpp somewhere to put the overflow layers without changing your graphics card.
- Rent a GPU by the hour instead — You want this daily, so renting is the way to try Llama 4 Scout 17B before spending on hardware — not the long-run answer if the usage holds.
- Consider waiting this market out — DDR5 kit prices are roughly 3-4x their mid-2025 level, and the shortage is expected to run into 2027. Relief is not expected before 2027-Q4. Meanwhile you have a way to run this that costs nothing extra.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Your machine is not exactly this one. Run this for your exact setup — the form opens pre-filled with RTX 5080 16GB and Llama 4 Scout 17B.
Questions people ask about this pairing
Can an RTX 5080 16GB run Llama 4 Scout 17B?
Not as it stands. Llama 4 Scout 17B needs 68.2 GB at Q4_K_M and this machine has 16 GB available, a shortfall of 52.2 GB. 128 GB of unified memory leaves about 96 GB for a model after the OS takes its share, which holds Llama 4 Scout 17B at Q4_K_M.
Would more system RAM fix this?
Here, yes — up to a point. You are 52.2 GB short in VRAM, and 96 GB of system RAM gives llama.cpp somewhere to put the overflow layers without changing your graphics card. It works because the shortfall is small enough that the layers living in system RAM do not dominate each token.
Should I wait for prices to come down?
DDR5 kit prices are roughly 3-4x their mid-2025 level, and the shortage is expected to run into 2027. Relief is not expected before 2027-Q4. The current estimate for relief is 2027-Q4. If you can run this some other way meanwhile, waiting is defensible; if you cannot, the part still does the job today.
Does a longer context change what Llama 4 Scout 17B needs here?
Yes, and it is the figure people forget. The 68.2 GB above already includes the KV cache at 8K; that cache grows roughly linearly with context, so doubling the window adds real gigabytes rather than a rounding error. If you plan to work with long documents on RTX 5080 16GB, size for the context you will actually use, not the default.
Related
- Llama 4 Scout 17B — full specs and VRAM by quantization
- NVIDIA GeForce RTX 5080 — specifications and what else it runs
- VRAM calculator for Llama 4 Scout 17B
- RTX 5080 16GB → Mistral Small 3.1 24B: change a setting
- RTX 5080 16GB → Qwen 3 30B-A3B (MoE): change a setting
- RTX 5070 12GB → Llama 4 Scout 17B: a different machine
Data behind this page last checked 2026-09-29.