Sized at Q4_K_M with a 8K context: weights, KV cache and framework overhead, against what RTX 3060 12GB leaves free. Phi-4 (14B) at Q4_K_M needs 10.9 GB and you have 12 GB available, so it fits with 1.1 GB to spare. Expect 14.4–27.7 tok/s (estimated).
Phi-4 (14B) at Q4_K_M needs 10.9 GB and you have 12 GB available, so it fits with 1.1 GB to spare.
Your machine is not exactly this one. Run this for your exact setup — the form opens pre-filled with RTX 3060 12GB and Phi-4 (14B).
Yes. Phi-4 (14B) at Q4_K_M needs 10.9 GB and you have 12 GB available, so it fits with 1.1 GB to spare.
Around 14.4–27.7 tok/s (estimated) on this setup at Q4_K_M. Decode speed on a local model is set by memory bandwidth rather than compute, so the figure moves with the card's bandwidth and the bytes read per token — not with its price.
Yes, and it is the figure people forget. The 10.9 GB above already includes the KV cache at 8K; that cache grows roughly linearly with context, so doubling the window adds real gigabytes rather than a rounding error. If you plan to work with long documents on RTX 3060 12GB, size for the context you will actually use, not the default.
Data behind this page last checked 2026-07-12.