Llama 3.3 70B Instruct — VRAM, Speed & Local Setup
作者: Jakub Rusinowski · 最后更新: 2024年12月8日
Model library → Llama 3.3 → Llama 3.3 70B Instruct
The current king of open weights. Exceptionally capable at following complex instructions, coding, and creative writing.
Llama 3.3 70B Instruct needs about 43 GB of VRAM at Q2_K_XS (Tight) — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
Specifications
| Parameters | 70 Billion |
| Context window | 128,000 |
| Architecture | Dense |
| Provider | Meta |
| Licence | Llama Community |
| Specified at | Q2_K_XS (Tight) |
| System RAM | 64 GB |
| Record updated | 2024-12-08 |
Curated — A hand-written entry from before this catalogue recorded its sources. The figures are long-standing but their provenance is not on file.
Licence
Llama Community — commercial use permitted. Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.
VRAM and Speed by Quantization
Modelled on a reference NVIDIA RTX 4090 (24 GB). Assumes an 8K-token context with an f16 KV cache. A longer window needs more; a quantized KV cache needs less. Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.
| Quant | Bits/weight | Weights | VRAM needed | Est. speed | Fit on 24 GB |
|---|---|---|---|---|---|
| Q2_K | 2.63 | 23 GB | 26.5 GB | ~4 tok/s (est.) | Offloads to system RAM (slow) |
| Q3_K_M | 3.41 | 29.8 GB | 33.3 GB | ~3 tok/s (est.) | Offloads to system RAM (slow) |
| Q4_K_M | 4.83 | 42.3 GB | 45.7 GB | ~3 tok/s (est.) | Offloads to system RAM (slow) |
| Q5_K_M | 5.67 | 49.6 GB | 53.1 GB | ~2 tok/s (est.) | Offloads to system RAM (slow) |
| Q6_K | 6.56 | 57.4 GB | 60.9 GB | — | Won't fit |
| Q8_0 | 8.50 | 74.4 GB | 77.9 GB | — | Won't fit |
| F16 | 16.00 | 140 GB | 143.5 GB | — | Won't fit |
Want to set your own context length and KV-cache quantization? Use the interactive VRAM calculator.
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
Recommended GPU
The cheapest catalogued GPU that runs Llama 3.3 70B Instruct is the Apple M5 Pro (64 GB).
How to Run Llama 3.3 70B Instruct
Install Ollama, then run:
ollama run llama3.3
Weights on Hugging Face: meta-llama/Llama-3.3-70B-Instruct.
Best for: coding, creative, reasoning.
Can I Run Llama 3.3 70B Instruct on My GPU?
- Llama 3.3 on AMD Radeon RX 7800 XT
- Llama 3.3 on AMD Radeon RX 7900 XT
- Llama 3.3 on AMD Radeon RX 7900 XTX
- Llama 3.3 on AMD Radeon RX 9060 XT 16GB
- Llama 3.3 on AMD Radeon RX 9070
- Llama 3.3 on AMD Radeon RX 9070 XT
- Llama 3.3 on Apple M1
- Llama 3.3 on Apple M1 Pro
- Llama 3.3 on Apple M2
- Llama 3.3 on Apple M2 Pro
- Llama 3.3 on Apple M3
- Llama 3.3 on Apple M3 Pro
- Llama 3.3 on Apple M4
- Llama 3.3 on Apple M4 Pro
- Llama 3.3 on Apple M5
- Llama 3.3 on Intel Arc B570
- Llama 3.3 on Intel Arc B580
- Llama 3.3 on NVIDIA GeForce RTX 3060 (12GB)
- Llama 3.3 on NVIDIA GeForce RTX 3080 (10GB)
- Llama 3.3 on NVIDIA GeForce RTX 3090
- Llama 3.3 on NVIDIA GeForce RTX 4060 Ti 16GB
- Llama 3.3 on NVIDIA GeForce RTX 4070
- Llama 3.3 on NVIDIA GeForce RTX 4070 Super
- Llama 3.3 on NVIDIA GeForce RTX 4070 Ti
- Llama 3.3 on NVIDIA GeForce RTX 4070 Ti Super
- Llama 3.3 on NVIDIA GeForce RTX 4080
- Llama 3.3 on NVIDIA GeForce RTX 4080 Super
- Llama 3.3 on NVIDIA GeForce RTX 4090
- Llama 3.3 on NVIDIA GeForce RTX 5060 Ti 16GB
- Llama 3.3 on NVIDIA GeForce RTX 5070
- Llama 3.3 on NVIDIA GeForce RTX 5070 Ti
- Llama 3.3 on NVIDIA GeForce RTX 5080
- Llama 3.3 on NVIDIA GeForce RTX 5090
- Llama 3.3 on NVIDIA L40S
- Llama 3.3 on NVIDIA RTX 6000 Ada Generation
Llama 3.3 70B Instruct — Frequently Asked Questions
← All Llama 3.3 models | VRAM calculator | Check your own hardware