Phi-4 (14B) — VRAM, Speed & Local Setup

Written by Jakub Rusinowski · Last updated January 6, 2025

Model libraryPhi-4 Family → Phi-4 (14B)

A powerhouse for 12GB+ VRAM GPUs. Exceptional reasoning capabilities derived from synthetic textbook quality data.

Phi-4 (14B) needs about 9 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Specifications

Parameters14 Billion
Context window16,000
ArchitectureDense
ProviderMicrosoft
LicenceMIT
Specified atQ4_K_M
System RAM16 GB
Record updated2025-01-06

Licence

MITcommercial use permitted. Commercial use permitted. No usage restrictions beyond attribution.

VRAM and Speed by Quantization

Modelled on a reference NVIDIA RTX 4090 (24 GB), at 8K context. Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.

QuantWeightsVRAM neededEst. speedFit on 24 GB
Q2_K4.6 GB7.1 GB~106 tok/s (est.)Fits comfortably
Q3_K_M6.0 GB8.4 GB~89 tok/s (est.)Fits comfortably
Q4_K_M8.5 GB10.9 GB~69 tok/s (est.)Fits comfortably
Q5_K_M9.9 GB12.4 GB~61 tok/s (est.)Fits comfortably
Q6_K11.5 GB14.0 GB~54 tok/s (est.)Fits comfortably
Q8_014.9 GB17.4 GB~43 tok/s (est.)Fits comfortably
F1628.0 GB30.5 GB~4 tok/s (est.)Offloads to system RAM (slow)

Want the memory numbers alone, at every quantization level and your own context length? Use the Phi-4 (14B) VRAM calculator.

Buy This HardwareIntel Arc B570 10GB — 10 GB VRAM · 150 W board powerDeploy in the Cloud NowRTX 4090 on RunPod — from $0.34/hr · rate checked 2026-07

or compare on Vast.ai from $0.35/hr (typical low · varies)

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Recommended GPU

The cheapest catalogued GPU that runs Phi-4 (14B) is the Intel Arc B570 (10 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Intel Arc B570 10GB
10 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.
Check price on Amazon

How to Run Phi-4 (14B)

Install Ollama, then run:

ollama run phi4

Weights on Hugging Face: microsoft/phi-4.

Download Phi-4 (14B) — GGUF Quantizations

Pick a quantization and open it in LM Studio, Ollama, or Jan, or download the raw .gguf file directly. Quant list and sizes resolved from Hugging Face.

Phi-4 (14B) — GGUF quants · bartowski/phi-4-GGUF

QuantSizeDownload (.gguf)
Q3_K_M5.97 GB (est.)phi-4-Q3_K_M.gguf
Q4_K_M8.45 GB (est.)phi-4-Q4_K_M.gguf
Q5_K_M9.92 GB (est.)phi-4-Q5_K_M.gguf
Q6_K11.48 GB (est.)phi-4-Q6_K.gguf
Q8_014.88 GB (est.)phi-4-Q8_0.gguf

Download in LM Studio: lms get bartowski/phi-4-GGUF

Want this model on your phone? You can run it on your desktop with LM Studio and chat from your iPhone or iPad over an encrypted link — see Run LM Studio Models on Your Phone (LM Link).

Best for: reasoning, math, stem.

Can I Run Phi-4 (14B) on My GPU?

Phi-4 (14B) — Frequently Asked Questions

How much VRAM does Phi-4 (14B) need?
About 9 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Phi-4 (14B) run on an RTX 4090 (24 GB)?
Yes. Phi-4 (14B) needs about 9 GB at Q4_K_M, inside a 24 GB card, at an estimated 69 tokens/sec.
How do I run Phi-4 (14B) locally?
Install Ollama and run `ollama run phi4`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.

← All Phi-4 Family models | VRAM calculator | Build a PC for this model | Check your own hardware