Qwen 3.5 397B-A17B — 显存、速度与本地部署
作者: Jakub Rusinowski · 最后更新:
Alibaba's Feb 16, 2026 flagship MoE model — 397B total, ~17B active. Even at aggressive sub-Q4 quantization it needs ~125 GB (e.g. a 128GB Mac Studio at ~24 tok/s). Not realistically self-hostable on consumer hardware; Ollama only lists a cloud-hosted `:cloud` tag, not a local download. Best treated as a reference/'what frontier open weights look like' entry rather than a local recommendation.
Qwen 3.5 397B-A17B 在 IQ2_M (sub-Q4) 下约需 240 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。
按量化级别的显存与速度
计算基准:NVIDIA RTX 4090 (24 GB)。仅含权重与开销:该模型架构未公开,因此未计入 KV 缓存。
| 量化 | 显存 | 显存 | 速度(估算) | 适配 |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 131.3 GB | — | 放不下 | |
| Q3_K_M 3.41 bpw | 170 GB | — | 放不下 | |
| Q4_K_M 4.83 bpw | 240.5 GB | — | 放不下 | |
| Q5_K_M 5.67 bpw | 282.2 GB | — | 放不下 | |
| Q6_K 6.56 bpw | 326.3 GB | — | 放不下 | |
| Q8_0 8.50 bpw | 422.6 GB | — | 放不下 | |
| F16 16.00 bpw | 794.8 GB | — | 放不下 |
黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Qwen 3.5 397B-A17B 显存计算器 →
运行 Qwen 3.5 397B-A17B
目录中能运行 Qwen 3.5 397B-A17B 的最便宜 GPU 是 Apple M3 Ultra (512 GB).
作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。
如何运行 Qwen 3.5 397B-A17B
The Ollama tag `qwen3.5:397b-cloud` is cloud-hosted, so there is no local run command for this model.
规格
Curated — A hand-written entry from before this catalogue recorded its sources. The figures are long-standing but their provenance is not on file.
- 参数量
- 397 Billion (~17B active)
- 上下文窗口
- 262,144
- 架构
- Hybrid Gated DeltaNet + MoE
- 提供商
- Alibaba Cloud
- 许可证
- Apache 2.0
- 规格量化
- IQ2_M (sub-Q4)
- 系统内存
- 256 GB
- 记录更新于
- 2026-02-16
Commercial use permitted. No usage restrictions beyond attribution.
质量与使用场景
评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。
我的 GPU 能运行 Qwen 3.5 397B-A17B 吗?
- Qwen 3.5 397B-A17B 在 AMD Radeon RX 7900 XT 上
- Qwen 3.5 397B-A17B 在 AMD Radeon RX 7900 XTX 上
- Qwen 3.5 397B-A17B 在 AMD Radeon RX 9060 XT 8GB 上
- Qwen 3.5 397B-A17B 在 AMD Ryzen AI Max+ 395 上
- Qwen 3.5 397B-A17B 在 Apple M1 Pro 上
- Qwen 3.5 397B-A17B 在 Apple M1 Ultra 上
- Qwen 3.5 397B-A17B 在 Apple M2 上
- Qwen 3.5 397B-A17B 在 Apple M2 Pro 上
- Qwen 3.5 397B-A17B 在 Apple M2 Ultra 上
- Qwen 3.5 397B-A17B 在 Apple M3 上
- Qwen 3.5 397B-A17B 在 Apple M3 Pro 上
- Qwen 3.5 397B-A17B 在 Apple M4 上
- Qwen 3.5 397B-A17B 在 Apple M4 Max 上
- Qwen 3.5 397B-A17B 在 Apple M4 Pro 上
- Qwen 3.5 397B-A17B 在 Apple M5 上
- Qwen 3.5 397B-A17B 在 Apple M5 Max 上
- Qwen 3.5 397B-A17B 在 Intel Arc B570 上
- Qwen 3.5 397B-A17B 在 Intel Arc B580 上
- Qwen 3.5 397B-A17B 在 NVIDIA A100 80GB (PCIe) 上
- Qwen 3.5 397B-A17B 在 NVIDIA DGX Spark 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 3060 (12GB) 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 3070 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 3070 Ti 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 3080 (10GB) 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 3090 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 4060 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 4070 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 4070 Super 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 4070 Ti 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 4090 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 5060 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 5060 Ti 8GB 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 5070 上
- Qwen 3.5 397B-A17B 在 NVIDIA GeForce RTX 5090 上
- Qwen 3.5 397B-A17B 在 NVIDIA H100 80GB (PCIe) 上
- Qwen 3.5 397B-A17B 在 NVIDIA L40S 上
- Qwen 3.5 397B-A17B 在 NVIDIA RTX 6000 Ada Generation 上
- Qwen 3.5 397B-A17B 在 NVIDIA RTX PRO 6000 Blackwell 上
Qwen 3.5 的其他尺寸
Qwen 3.5 397B-A17B — 常见问题
How much VRAM does Qwen 3.5 397B-A17B need?
About 240 GB at IQ2_M (sub-Q4) — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Qwen 3.5 397B-A17B run on an RTX 4090 (24 GB)?
No. Qwen 3.5 397B-A17B needs about 240 GB at IQ2_M (sub-Q4), more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.
How do I run Qwen 3.5 397B-A17B locally?
The Ollama tag `qwen3.5:397b-cloud` is cloud-hosted, so there is no local run command for this model. Running the published tag would send your prompts to a hosted GPU rather than your own machine.
What other sizes does Qwen 3.5 come in?
Qwen 3.5 0.8B (1 GB), Qwen 3.5 2B (2 GB), Qwen 3.5 4B (3 GB), Qwen 3.5 9B (6 GB), Qwen 3.5 27B (17 GB), Qwen 3.5 35B-A3B (22 GB), Qwen 3.5 122B-A10B (74 GB), Qwen 3.5 397B-A17B (240 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.