Arc B570 10GB can run Phi-4 (14B), but not the way you'd expect. Here's the change that makes it work.
Phi-4 (14B) at Q4_K_M needs 10.9 GB once weights, KV cache at 8K context and framework overhead are counted. Arc B570 10GB offers 10 GB, leaving you 0.9 GB short. You are 0.9 GB short in VRAM but have 25.6 GB of usable system RAM, so llama.cpp can hold the overflow layers in RAM instead of refusing to load. Expect 8.8–18.2 tok/s (estimated).
What to do
You are 0.9 GB short in VRAM but have 25.6 GB of usable system RAM, so llama.cpp can hold the overflow layers in RAM instead of refusing to load.
- Use llama.cpp or Ollama, which support partial GPU offload
- Keep at least 0.9 GB of system RAM free while the model is loaded
- Deficit: 0.9 GB
- System RAM usable for offload: 25.6 GB
- Offload throughput checked against the 5 tok/s usable floor
Other ways to get there
- Replace your card with an RTX 4060 Ti 16GB — 16 GB holds Phi-4 (14B) at Q4_K_M outright, closing the 0.9 GB you are short by.
- Add an RTX 4060 Ti 16GB alongside your current card — Your 10 GB plus this card's 16 GB clears the 10.9 GB this model needs, without replacing what you already own.
- Add 64GB DDR5 kit (2x32GB) — You are 0.9 GB short in VRAM, and 64 GB of system RAM gives llama.cpp somewhere to put the overflow layers without changing your graphics card.
Your machine is not exactly this one. Run this for your exact setup — the form opens pre-filled with Arc B570 10GB and Phi-4 (14B).
Questions people ask about this pairing
Can an Arc B570 10GB run Phi-4 (14B)?
Not as it stands. Phi-4 (14B) needs 10.9 GB at Q4_K_M and this machine has 10 GB available, a shortfall of 0.9 GB. You are 0.9 GB short in VRAM but have 25.6 GB of usable system RAM, so llama.cpp can hold the overflow layers in RAM instead of refusing to load.
Would more system RAM fix this?
Here, yes — up to a point. You are 0.9 GB short in VRAM, and 64 GB of system RAM gives llama.cpp somewhere to put the overflow layers without changing your graphics card. It works because the shortfall is small enough that the layers living in system RAM do not dominate each token.
Should I wait for prices to come down?
New GPU prices are far above MSRP because the memory on the board costs several times what it did in 2025. Relief is not expected before 2027-Q4. The current estimate for relief is 2027-Q4. If you can run this some other way meanwhile, waiting is defensible; if you cannot, the part still does the job today.
How fast will Phi-4 (14B) actually run?
Around 8.8–18.2 tok/s (estimated) on this setup at Q4_K_M. Decode speed on a local model is set by memory bandwidth rather than compute, so the figure moves with the card's bandwidth and the bytes read per token — not with its price.
Does a longer context change what Phi-4 (14B) needs here?
Yes, and it is the figure people forget. The 10.9 GB above already includes the KV cache at 8K; that cache grows roughly linearly with context, so doubling the window adds real gigabytes rather than a rounding error. If you plan to work with long documents on Arc B570 10GB, size for the context you will actually use, not the default.
Related
- Phi-4 (14B) — full specs and VRAM by quantization
- Intel Arc B570 — specifications and what else it runs
- Can I run Phi-4 (14B) on an Intel Arc B570?
- VRAM calculator for Phi-4 (14B)
- MacBook Air M2, 8GB → Phi-4 (14B): a different machine
- RTX 5070 12GB → Phi-4 (14B): keep what you have
Data behind this page last checked 2026-08-26.