Inkling (BF16) — VRAM Requirements

Written by Jakub Rusinowski · Last updated July 16, 2026

How much GPU VRAM you need to run Inkling Inkling (BF16) by Thinking Machines locally, a 975B-parameter model. Figures are quantized weights + KV cache + framework overhead, computed from the model's parameter count and published architecture — not a throughput model. See /en/methodology.

Inkling (BF16) needs about 589 GB VRAM at Q4_K_M.

VRAM by Quantization

QuantBits/weightWeightsTotal VRAM
Q2_K2.63320.5 GB321.3 GB
Q3_K_M3.41415.6 GB416.4 GB
Q4_K_M4.83588.7 GB589.5 GB
Q5_K_M5.67691.0 GB691.8 GB
Q6_K6.56799.5 GB800.3 GB
Q8_08.501035.9 GB1036.7 GB
F1616.001950.0 GB1950.8 GB

Switch quantization in the interactive calculator, or see the full Inkling model page.

Deploy in the Cloud NowRunPod

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Add this badge to your model card

Model creators: paste this into your Hugging Face model card README to link readers straight to this VRAM breakdown.

VRAM Requirements

[![VRAM Requirements](https://img.shields.io/badge/Check_VRAM-LLM_Configurator-blue)](https://llmconfigurator.com/en/vram-calculator/inkling-bf16?utm_source=badge&utm_medium=referral&utm_campaign=readme_badge&utm_content=inkling-bf16)

Estimates only — actual VRAM varies with context length, batch size, runtime and KV-cache settings.