Alibaba Cloud122B (10B active)Q4_K_M 下约 74 GB 显存

Qwen 3.5 122B-A10B — 显存、速度与本地部署

作者: Jakub Rusinowski · 最后更新:

Large MoE model with 122B total parameters (256 routed + 1 shared expert) and 10B active per token. Q4_K_M is ~74 GB — needs 2x24GB+ GPUs, a single 80GB GPU, or a Mac Studio with 96GB+ unified memory. Delivers near-frontier performance for long-context enterprise tasks.

Qwen 3.5 122B-A10B 在 Q4_K_M 下约需 74 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。

逻辑93
创意91
编程93

按量化级别的显存与速度

计算基准:NVIDIA RTX 4090 (24 GB)。仅含权重与开销:该模型架构未公开,因此未计入 KV 缓存。

量化显存速度(估算)适配
Q2_K
2.63 bpw
40.9 GB~18 tok/s需卸载
Q3_K_M
3.41 bpw
52.8 GB~15 tok/s需卸载
Q4_K_M
4.83 bpw
74.5 GB—放不下
Q5_K_M
5.67 bpw
87.3 GB—放不下
Q6_K
6.56 bpw
100.8 GB—放不下
Q8_0
8.50 bpw
130.4 GB—放不下
F16
16.00 bpw
244.8 GB—放不下

黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Qwen 3.5 122B-A10B 显存计算器 →

运行 Qwen 3.5 122B-A10B

目录中能运行 Qwen 3.5 122B-A10B 的最便宜 GPU 是 AMD Ryzen AI Max+ 395 (128 GB).

联盟营销声明: 本页部分链接为联盟推广链接——如果你通过它们购买,LLM Configurator 可能会获得佣金,而你无需支付任何额外费用。作为亚马逊联盟成员(Amazon Associate),LLM Configurator 会从符合条件的购买中获得收益。
Ryzen AI Max+ 395 Laptop (Strix Halo, up to 128GB)
128 GB VRAM · 120 W board power
2026年价格波动较大——请以当前商品页价格为准。

如何运行 Qwen 3.5 122B-A10B

安装 Ollama,然后运行:

ollama run qwen3.5:122b
Hugging Face 上的权重: Qwen/Qwen3.5-122B-A10B-Instruct ↗

规格

Curated — A hand-written entry from before this catalogue recorded its sources. The figures are long-standing but their provenance is not on file.

参数量
122 Billion (10B active)
上下文窗口
262,144
架构
Hybrid Gated DeltaNet + MoE
提供商
Alibaba Cloud
许可证
Apache 2.0
规格量化
Q4_K_M
系统内存
128 GB
记录更新于
2026-02-24
许可证Apache-2.0允许商业使用

Commercial use permitted. No usage restrictions beyond attribution.

质量与使用场景

评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。

最适合enterpriselong contextresearchmulti gpu

我的 GPU 能运行 Qwen 3.5 122B-A10B 吗?

Qwen 3.5 的其他尺寸

Qwen 3.5 122B-A10B — 常见问题

How much VRAM does Qwen 3.5 122B-A10B need?

About 74 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Qwen 3.5 122B-A10B run on an RTX 4090 (24 GB)?

No. Qwen 3.5 122B-A10B needs about 74 GB at Q4_K_M, more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.

How do I run Qwen 3.5 122B-A10B locally?

Install Ollama and run `ollama run qwen3.5:122b`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.

What other sizes does Qwen 3.5 come in?

Qwen 3.5 0.8B (1 GB), Qwen 3.5 2B (2 GB), Qwen 3.5 4B (3 GB), Qwen 3.5 9B (6 GB), Qwen 3.5 27B (17 GB), Qwen 3.5 35B-A3B (22 GB), Qwen 3.5 122B-A10B (74 GB), Qwen 3.5 397B-A17B (240 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.