Llama 3.2 90B Vision Instruct — 显存、速度与本地部署
作者: Jakub Rusinowski · 最后更新:
The flagship open vision model. Near GPT-4V quality for image understanding. Requires 48GB+ VRAM or multi-GPU setup. Listed in full on the dedicated Llama 3.2 Vision page, which is the canonical entry for this model and carries the install command.
Llama 3.2 90B Vision Instruct 在 Q4_K_M 下约需 54 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。
按量化级别的显存与速度
计算基准:NVIDIA RTX 4090 (24 GB)。包含 8K 上下文的 KV 缓存,因此数值高于上方的主数字。
| 量化 | 显存 | 显存 | 速度(估算) | 适配 |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 33.3 GB | ~3 tok/s | 需卸载 | |
| Q3_K_M 3.41 bpw | 42 GB | ~3 tok/s | 需卸载 | |
| Q4_K_M 4.83 bpw | 57.8 GB | — | 放不下 | |
| Q5_K_M 5.67 bpw | 67.1 GB | — | 放不下 | |
| Q6_K 6.56 bpw | 77 GB | — | 放不下 | |
| Q8_0 8.50 bpw | 98.5 GB | — | 放不下 | |
| F16 16.00 bpw | 181.8 GB | — | 放不下 |
黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Llama 3.2 90B Vision Instruct 显存计算器 →
运行 Llama 3.2 90B Vision Instruct
目录中能运行 Llama 3.2 90B Vision Instruct 的最便宜 GPU 是 Apple M5 Pro (64 GB).
或在 Vast.ai 比较,低至 $0.77/小时 (typical low · varies)
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
如何运行 Llama 3.2 90B Vision Instruct
安装 Ollama,然后运行:
ollama run llama-3-2规格
Curated — A hand-written entry from before this catalogue recorded its sources. The figures are long-standing but their provenance is not on file.
- 参数量
- 90 Billion
- 上下文窗口
- 128,000
- 架构
- Dense + Vision Encoder
- 提供商
- Meta
- 许可证
- Llama Community
- 规格量化
- Q4_K_M
- 系统内存
- 128 GB
- 记录更新于
- 2026-08-15
Weights are downloadable and commercial use is permitted, subject to the licence’s acceptable-use terms.
质量与使用场景
评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。
我的 GPU 能运行 Llama 3.2 90B Vision Instruct 吗?
- Llama 3.2 90B Vision Instruct 在 AMD Radeon RX 9060 XT 8GB 上
- Llama 3.2 90B Vision Instruct 在 AMD Ryzen AI Max+ 395 上
- Llama 3.2 90B Vision Instruct 在 Apple M1 Ultra 上
- Llama 3.2 90B Vision Instruct 在 Apple M2 Max 上
- Llama 3.2 90B Vision Instruct 在 Apple M4 Max 上
- Llama 3.2 90B Vision Instruct 在 Apple M5 Max 上
- Llama 3.2 90B Vision Instruct 在 Intel Arc B570 上
- Llama 3.2 90B Vision Instruct 在 Intel Arc B580 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA A100 80GB (PCIe) 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA DGX Spark 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 3060 (12GB) 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 3070 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 3070 Ti 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 3080 (10GB) 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 4060 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 4070 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 4070 Super 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 4070 Ti 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 5060 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 5060 Ti 8GB 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA GeForce RTX 5070 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA H100 80GB (PCIe) 上
- Llama 3.2 90B Vision Instruct 在 NVIDIA RTX PRO 6000 Blackwell 上
Llama 3.2 Family 的其他尺寸
Llama 3.2 90B Vision Instruct — 常见问题
How much VRAM does Llama 3.2 90B Vision Instruct need?
About 54 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Llama 3.2 90B Vision Instruct run on an RTX 4090 (24 GB)?
No. Llama 3.2 90B Vision Instruct needs about 54 GB at Q4_K_M, more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.
How do I run Llama 3.2 90B Vision Instruct locally?
Install Ollama and run `ollama run llama-3-2`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.
What other sizes does Llama 3.2 Family come in?
Llama 3.2 1B Instruct (2 GB), Llama 3.2 3B Instruct (3 GB), Llama 3.2 11B Vision Instruct (7 GB), Llama 3.2 90B Vision Instruct (54 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.