Granite 4.1 30B — 显存、速度与本地部署
作者: Jakub Rusinowski · 最后更新:
Largest Granite 4.1 tier — a hybrid Mamba2/attention design (interleaved SSM + attention layers), not MoE despite the 4.0 generation using MoE. ~18GB at FP8 (Q4 likely lower, ~12-15GB, unconfirmed). Context extendable to 512K. Apache 2.0.
Granite 4.1 30B 在 FP8 (Q4 est. ~12-15GB, unconfirmed) 下约需 19 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。
按量化级别的显存与速度
计算基准:NVIDIA RTX 4090 (24 GB)。仅含权重与开销:该模型架构未公开,因此未计入 KV 缓存。
| 量化 | 显存 | 显存 | 速度(估算) | 适配 |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 10.7 GB | ~61 tok/s | 可运行 | |
| Q3_K_M 3.41 bpw | 13.6 GB | ~49 tok/s | 可运行 | |
| Q4_K_M 4.83 bpw | 18.9 GB | ~37 tok/s | 可运行 | |
| Q5_K_M 5.67 bpw | 22.1 GB | ~32 tok/s | 勉强 | |
| Q6_K 6.56 bpw | 25.4 GB | ~4 tok/s | 需卸载 | |
| Q8_0 8.50 bpw | 32.7 GB | ~3 tok/s | 需卸载 | |
| F16 16.00 bpw | 60.8 GB | — | 放不下 |
黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Granite 4.1 30B 显存计算器 →
运行 Granite 4.1 30B
目录中能运行 Granite 4.1 30B 的最便宜 GPU 是 AMD Radeon RX 7900 XT (20 GB).
或在 Vast.ai 比较,低至 $0.35/小时 (typical low · varies)
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
如何运行 Granite 4.1 30B
安装 Ollama,然后运行:
ollama run granite4.1:30b规格
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.
- 参数量
- 30 Billion
- 上下文窗口
- 128,000
- 架构
- Hybrid Mamba2 + Attention (SSM)
- 提供商
- IBM
- 许可证
- Apache 2.0
- 规格量化
- FP8 (Q4 est. ~12-15GB, unconfirmed)
- 系统内存
- 32 GB
- 记录更新于
- 2026-04-29
Commercial use permitted. No usage restrictions beyond attribution.
质量与使用场景
评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。
我的 GPU 能运行 Granite 4.1 30B 吗?
- Granite 4.1 30B 在 AMD Radeon RX 9060 XT 8GB 上
- Granite 4.1 30B 在 Apple M1 Max 上
- Granite 4.1 30B 在 Apple M3 Max 上
- Granite 4.1 30B 在 Apple M5 Pro 上
- Granite 4.1 30B 在 Intel Arc B570 上
- Granite 4.1 30B 在 Intel Arc B580 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 3060 (12GB) 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 3070 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 3070 Ti 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 3080 (10GB) 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 4060 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 4070 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 4070 Super 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 4070 Ti 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 5060 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 5060 Ti 8GB 上
- Granite 4.1 30B 在 NVIDIA GeForce RTX 5070 上
- Granite 4.1 30B 在 NVIDIA L40S 上
- Granite 4.1 30B 在 NVIDIA RTX 6000 Ada Generation 上
IBM Granite 4.1 的其他尺寸
Granite 4.1 30B — 常见问题
How much VRAM does Granite 4.1 30B need?
About 19 GB at FP8 (Q4 est. ~12-15GB, unconfirmed) — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Granite 4.1 30B run on an RTX 4090 (24 GB)?
Yes. Granite 4.1 30B needs about 19 GB at FP8 (Q4 est. ~12-15GB, unconfirmed), inside a 24 GB card, at an estimated 37 tokens/sec.
How do I run Granite 4.1 30B locally?
Install Ollama and run `ollama run granite4.1:30b`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.
What other sizes does IBM Granite 4.1 come in?
Granite 4.1 3B (3 GB), Granite 4.1 8B (6 GB), Granite 4.1 30B (19 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.