Thinking Machines1TNVFP4 (4-bit) 下约 589 GB 显存

Inkling (NVFP4) — 显存、速度与本地部署

作者: Jakub Rusinowski · 最后更新:

The 4-bit deployment build. NVFP4 quantization brings the 975B / 41B-active MoE down to ~600 GB — a single high-memory GPU server rather than a multi-node cluster — while keeping the full 1M-token context and native text / image / audio input. Includes MTP speculative-decoding layers for faster generation and runs under vLLM, SGLang, Transformers 5.14+ and llama.cpp. Blackwell GPUs are recommended for native FP4. Apache 2.0. Thinking Machines reports AIME 2026 97.1%, GPQA Diamond 87.2%, SWE-bench Verified 77.6% and VoiceBench 91.4%. Specs from launch coverage — verify on the Hugging Face model card.

Inkling (NVFP4) 在 NVFP4 (4-bit) 下约需 589 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。

逻辑97
创意93
编程96

按量化级别的显存与速度

计算基准:NVIDIA RTX 4090 (24 GB)。仅含权重与开销:该模型架构未公开,因此未计入 KV 缓存。

量化显存速度(估算)适配
Q2_K
2.63 bpw
321.3 GB—放不下
Q3_K_M
3.41 bpw
416.4 GB—放不下
Q4_K_M
4.83 bpw
589.5 GB—放不下
Q5_K_M
5.67 bpw
691.8 GB—放不下
Q6_K
6.56 bpw
800.3 GB—放不下
Q8_0
8.50 bpw
1036.7 GB—放不下
F16
16.00 bpw
1950.8 GB—放不下

黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Inkling (NVFP4) 显存计算器 →

运行 Inkling (NVFP4)

立即云端部署 RunPod

或在 Vast.ai 比较

作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。

如何运行 Inkling (NVFP4)

安装 Ollama,然后运行:

ollama run inkling
Hugging Face 上的权重: thinkingmachines/Inkling-NVFP4 ↗

规格

Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.

Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.

参数量
975 Billion (41B active)
上下文窗口
1,000,000
架构
Multimodal Mixture-of-Experts — 256 experts, top-6 routed + 2 shared; NVFP4 4-bit
提供商
Thinking Machines
许可证
Apache 2.0
规格量化
NVFP4 (4-bit)
系统内存
1024 GB
记录更新于
2026-07-16
许可证Apache-2.0允许商业使用

Commercial use permitted. No usage restrictions beyond attribution.

质量与使用场景

评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。

最适合multimodalagentic taskslong contextfine tuningenterprise
基准分数来源
AIME 202697.1 %厂商声称 · Thinking Machines (launch)
GPQA Diamond87.2 %厂商声称 · Thinking Machines (launch)
HLE (with tools)46 %厂商声称 · Thinking Machines (launch)
SWE-bench Verified77.6 %厂商声称 · Thinking Machines (launch)
MMMU Pro (Standard 10)73.3 %厂商声称 · Thinking Machines (launch)
VoiceBench91.4 %厂商声称 · Thinking Machines (launch)
Global-MMLU-Lite88.7 %厂商声称 · Thinking Machines (launch)

我的 GPU 能运行 Inkling (NVFP4) 吗?

Inkling 的其他尺寸

Inkling (NVFP4) — 常见问题

How much VRAM does Inkling (NVFP4) need?

About 589 GB at NVFP4 (4-bit) — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Inkling (NVFP4) run on an RTX 4090 (24 GB)?

No. Inkling (NVFP4) needs about 589 GB at NVFP4 (4-bit), more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.

How do I run Inkling (NVFP4) locally?

Install Ollama and run `ollama run inkling`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.

What other sizes does Inkling come in?

Inkling (NVFP4) (589 GB), Inkling (BF16) (589 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.