Tell us the model you want to run and what you own. You get a ranked set of routes from here to there — including the ones that cost nothing, which win whenever they are correct.
We size the model the way every other page here does: total parameters at your quantization, plus the KV cache for your context length, plus framework overhead. Mixture-of-Experts models are sized by TOTAL parameters, because every expert must be resident even though only some are read per token.
We compare that against what your machine can actually offer — a card's usable VRAM, or an Apple or mini-PC unified pool minus the share the OS keeps.
Then we rank every route that closes the gap. A route that costs nothing outranks one that costs money whenever it works: if a quantization at or above Q4_K_M fits, or a smaller model in the same family fits, or your card plus system RAM holds it at a usable speed, we recommend buying nothing.
You will not see a price next to a shop link here. Amazon's terms only permit showing a price beside a link into their store when it is fetched live from their product API, which we do not have — so rather than a number that may be months stale, we show the specification and send you to the retailer, who knows today's price. Budget is your ceiling, not our guess at anyone's price.
Memory and graphics cards are unusually expensive in 2026: AI datacentre demand redirected the capacity that used to make consumer parts. Where that affects the part you would need, and you can run the model some other way meanwhile, we say so and give the date relief is expected, with our sources. A rumour is labelled a rumour, and we tell you not to defer a purchase on one.