Inkling — Thinking Machines 的本地 AI 模型

作者: Jakub Rusinowski · 最后更新:

PREVIEW (July 2026). Thinking Machines Lab's first open-weights model — and the first ~1-trillion-parameter open model to natively accept text, image and audio input with a 1M-token context window. Inkling is a sparse multimodal Mixture-of-Experts transformer: 975B total parameters but only ~41B active per token (256 experts, top-6 routed plus 2 always-on shared experts). It was trained on 45 trillion tokens of text, images, audio and video, and reasons natively across modalities with seven selectable effort levels (none → max). Ships in full BF16 (~2 TB VRAM, Hopper+) and a well-calibrated NVFP4 4-bit build (~600 GB, Blackwell), both with MTP speculative-decoding layers for faster inference. Released under Apache 2.0 and pitched as a customization base to fine-tune via the company's Tinker platform — "built for customization, not leaderboard dominance." Datacenter-class hardware required; most users will reach it through a hosted API. Specs from launch coverage and the Hugging Face model card — verify before relying on them.

变体

Inkling 最小的变体在 NVFP4 (4-bit) 下约需 589 GB 显存——量化权重加框架开销,不含 KV 缓存。

模型显存
Inkling (NVFP4) →
1T
~589.5 GB
Inkling (BF16) →
1T
~589.5 GB

显存为 Q4_K_M 下的量化权重加开销,与 GPU 与显存检测器使用同一引擎计算。

如何在本地运行 Inkling

安装 Ollama,然后拉取标签。

ollama run inkling

在上方选择一个尺寸,查看它自己的显存、速度估算和安装命令。

许可证

Apache-2.0允许商业使用

Commercial use permitted. No usage restrictions beyond attribution.

适用于: Inkling (NVFP4), Inkling (BF16)

我的 GPU 能运行 Inkling 吗?

Inkling — 常见问题

How much VRAM does Inkling need?

Inkling needs about 589 GB VRAM at NVFP4 (4-bit) quantization for its smallest variant. Variants: Inkling (NVFP4) (589 GB, NVFP4 (4-bit)); Inkling (BF16) (589 GB, BF16 (full precision)). On Apple Silicon, unified memory counts toward this requirement.

Can I run Inkling on an RTX 4090 (24 GB)?

Inkling's smallest variant needs about 589 GB, which exceeds a single RTX 4090 (24 GB). Use multiple GPUs, a higher-VRAM card, or Apple Silicon with large unified memory.

What quantization should I use for Inkling?

Q4_K_M is the best balance of quality and VRAM for Inkling in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run Inkling with Ollama?

Inkling has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.