Inkling — Thinking Machines 的本地 AI 模型
作者: Jakub Rusinowski · 最后更新:
PREVIEW (July 2026). Thinking Machines Lab's first open-weights model — and the first ~1-trillion-parameter open model to natively accept text, image and audio input with a 1M-token context window. Inkling is a sparse multimodal Mixture-of-Experts transformer: 975B total parameters but only ~41B active per token (256 experts, top-6 routed plus 2 always-on shared experts). It was trained on 45 trillion tokens of text, images, audio and video, and reasons natively across modalities with seven selectable effort levels (none → max). Ships in full BF16 (~2 TB VRAM, Hopper+) and a well-calibrated NVFP4 4-bit build (~600 GB, Blackwell), both with MTP speculative-decoding layers for faster inference. Released under Apache 2.0 and pitched as a customization base to fine-tune via the company's Tinker platform — "built for customization, not leaderboard dominance." Datacenter-class hardware required; most users will reach it through a hosted API. Specs from launch coverage and the Hugging Face model card — verify before relying on them.
变体
Inkling 最小的变体在 NVFP4 (4-bit) 下约需 589 GB 显存——量化权重加框架开销,不含 KV 缓存。
| 模型 | Q4 下显存 | 显存 | 上下文 | 运行 |
|---|---|---|---|---|
| Inkling (NVFP4) → 1T | ~589.5 GB | 1,000,000 | ollama run inkling | |
| Inkling (BF16) → 1T | ~589.5 GB | 1,000,000 | ollama run inkling |
显存为 Q4_K_M 下的量化权重加开销,与 GPU 与显存检测器使用同一引擎计算。
如何在本地运行 Inkling
安装 Ollama,然后拉取标签。
ollama run inkling在上方选择一个尺寸,查看它自己的显存、速度估算和安装命令。
许可证
Commercial use permitted. No usage restrictions beyond attribution.
适用于: Inkling (NVFP4), Inkling (BF16)我的 GPU 能运行 Inkling 吗?
Inkling — 常见问题
How much VRAM does Inkling need?
Inkling needs about 589 GB VRAM at NVFP4 (4-bit) quantization for its smallest variant. Variants: Inkling (NVFP4) (589 GB, NVFP4 (4-bit)); Inkling (BF16) (589 GB, BF16 (full precision)). On Apple Silicon, unified memory counts toward this requirement.
Can I run Inkling on an RTX 4090 (24 GB)?
Inkling's smallest variant needs about 589 GB, which exceeds a single RTX 4090 (24 GB). Use multiple GPUs, a higher-VRAM card, or Apple Silicon with large unified memory.
What quantization should I use for Inkling?
Q4_K_M is the best balance of quality and VRAM for Inkling in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Inkling with Ollama?
Inkling has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.