Llama 3.3 70B Instruct — VRAM, Speed & Local Setup

作者: Jakub Rusinowski · 最后更新: 2024年12月8日

Model library → Llama 3.3 → Llama 3.3 70B Instruct

The current king of open weights. Exceptionally capable at following complex instructions, coding, and creative writing.

Llama 3.3 70B Instruct needs about 43 GB of VRAM at Q2_K_XS (Tight) — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Specifications

Parameters70 Billion
Context window128,000
ArchitectureDense
ProviderMeta
LicenceLlama Community
Specified atQ2_K_XS (Tight)
System RAM64 GB
Record updated2024-12-08

Curated — A hand-written entry from before this catalogue recorded its sources. The figures are long-standing but their provenance is not on file.

Licence

Llama Community — commercial use permitted. Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.

VRAM and Speed by Quantization

Modelled on a reference NVIDIA RTX 4090 (24 GB). Assumes an 8K-token context with an f16 KV cache. A longer window needs more; a quantized KV cache needs less. Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.

QuantBits/weightWeightsVRAM neededEst. speedFit on 24 GB
Q2_K2.6323 GB26.5 GB~4 tok/s (est.)Offloads to system RAM (slow)
Q3_K_M3.4129.8 GB33.3 GB~3 tok/s (est.)Offloads to system RAM (slow)
Q4_K_M4.8342.3 GB45.7 GB~3 tok/s (est.)Offloads to system RAM (slow)
Q5_K_M5.6749.6 GB53.1 GB~2 tok/s (est.)Offloads to system RAM (slow)
Q6_K6.5657.4 GB60.9 GB—Won't fit
Q8_08.5074.4 GB77.9 GB—Won't fit
F1616.00140 GB143.5 GB—Won't fit

Want to set your own context length and KV-cache quantization? Use the interactive VRAM calculator.

购买此硬件 Apple MacBook Pro M5 Pro — 64 GB VRAM · 30 W board power立即云端部署 RunPod 上的 NVIDIA A40 — 低至 $0.44/小时 · 价格核实于 2026-08

或在 Vast.ai 比较

作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。

Recommended GPU

The cheapest catalogued GPU that runs Llama 3.3 70B Instruct is the Apple M5 Pro (64 GB).

联盟营销声明: 本页部分链接为联盟推广链接——如果你通过它们购买,LLM Configurator 可能会获得佣金,而你无需支付任何额外费用。作为亚马逊联盟成员(Amazon Associate),LLM Configurator 会从符合条件的购买中获得收益。
Apple MacBook Pro M5 Pro
64 GB VRAM · 30 W board power
2026年价格波动较大——请以当前商品页价格为准。
在亚马逊查看价格

How to Run Llama 3.3 70B Instruct

Install Ollama, then run:

ollama run llama3.3

Weights on Hugging Face: meta-llama/Llama-3.3-70B-Instruct.

Best for: coding, creative, reasoning.

Can I Run Llama 3.3 70B Instruct on My GPU?

Llama 3.3 70B Instruct — Frequently Asked Questions

How much VRAM does Llama 3.3 70B Instruct need?
About 43 GB at Q2_K_XS (Tight) — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Llama 3.3 70B Instruct run on an RTX 4090 (24 GB)?
No. Llama 3.3 70B Instruct needs about 43 GB at Q2_K_XS (Tight), more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.
How do I run Llama 3.3 70B Instruct locally?
Install Ollama and run `ollama run llama3.3`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.

← All Llama 3.3 models | VRAM calculator | Check your own hardware