Mistral Small 3.1 24B at Q4_K_M needs 16.4 GB once weights, KV cache at 8K context and framework overhead are counted. 16GB RAM, no graphics card offers 0 GB, leaving you 16.4 GB short. 24 GB holds Mistral Small 3.1 24B at Q4_K_M outright, closing the 16.4 GB you are short by. Expect 23.6–45.3 tok/s (estimated).
24 GB holds Mistral Small 3.1 24B at Q4_K_M outright, closing the 16.4 GB you are short by.
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
Your machine is not exactly this one. Run this for your exact setup — the form opens pre-filled with 16GB RAM, no graphics card and Mistral Small 3.1 24B.
Not as it stands. Mistral Small 3.1 24B needs 16.4 GB at Q4_K_M and this machine has 0 GB available, a shortfall of 16.4 GB. 24 GB holds Mistral Small 3.1 24B at Q4_K_M outright, closing the 16.4 GB you are short by.
With no GPU to hold any of the model, more system RAM would let it load but it would run entirely on the CPU — too slow to use for anything interactive. The bottleneck here is the missing GPU, not the memory.
Check your power supply before ordering: this card needs a unit rated 650 W or better (350 W board power plus the rest of the system, with the 30% headroom a PSU needs to run efficiently and age well). That figure is the card's board power plus a typical system draw, with the 30% headroom a supply needs to run efficiently and age well — not the bare minimum it would boot at.
It is the standing value answer for this capacity, and it is the reason this recommendation exists — but it is used-market only, so there is no warranty and no way to know how the card was treated. Buy from a seller with returns, and test it under sustained load in the first week.
New GPU prices are far above MSRP because the memory on the board costs several times what it did in 2025. Relief is not expected before 2027-Q4. The current estimate for relief is 2027-Q4. If you can run this some other way meanwhile, waiting is defensible; if you cannot, the part still does the job today.
Around 23.6–45.3 tok/s (estimated) on this setup at Q4_K_M. Decode speed on a local model is set by memory bandwidth rather than compute, so the figure moves with the card's bandwidth and the bytes read per token — not with its price.
Yes, and it is the figure people forget. The 16.4 GB above already includes the KV cache at 8K; that cache grows roughly linearly with context, so doubling the window adds real gigabytes rather than a rounding error. If you plan to work with long documents on 16GB RAM, no graphics card, size for the context you will actually use, not the default.
Data behind this page last checked 2026-08-26.