DeepSeek V4.1 — local AI model by DeepSeek
Written by Jakub Rusinowski · Last updated
DeepSeek V4.1 — the Flash member shipped 10 September 2026 under MIT with an encoder-decoder design and native image input. A V4.1 Pro was announced without a date and had not shipped as of September 2026.
Variants
The smallest DeepSeek V4.1 variant needs about 172 GB of VRAM at Q4 (experimental) — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| DeepSeek V4.1 Flash → 284B (13B active) | ~172.3 GB | 1,000,000 | ollama run deepseek-v4-1 | |
| DeepSeek V4.1 1.6T (49B active) | ~966.8 GB | 1,000,000 | ollama run deepseek-v4-1 |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run DeepSeek V4.1 locally
Install Ollama, then pull the tag.
ollama run deepseek-v4-1Pick a size above for its own VRAM figure, speed estimate and install command.
Licence
Commercial use permitted. No usage restrictions beyond attribution.
Applies to: DeepSeek V4.1 Flash, DeepSeek V4.1Recommended GPU
The cheapest catalogued GPU that runs DeepSeek V4.1 locally (min 172 GB VRAM) is the Apple M2 Ultra (192 GB).
Can I run DeepSeek V4.1 on my GPU?
- DeepSeek V4.1 on AMD Ryzen AI Max+ 395
- DeepSeek V4.1 on Apple M1 Ultra
- DeepSeek V4.1 on Apple M2 Max
- DeepSeek V4.1 on Apple M2 Ultra
- DeepSeek V4.1 on Apple M4 Max
- DeepSeek V4.1 on Apple M5 Max
- DeepSeek V4.1 on NVIDIA A100 80GB (PCIe)
- DeepSeek V4.1 on NVIDIA DGX Spark
- DeepSeek V4.1 on NVIDIA H100 80GB (PCIe)
- DeepSeek V4.1 on NVIDIA RTX PRO 6000 Blackwell
DeepSeek V4.1 — frequently asked questions
How much VRAM does DeepSeek V4.1 need?
DeepSeek V4.1 needs about 172 GB VRAM at Q4 (experimental) quantization for its smallest variant. Variants: DeepSeek V4.1 Flash (172 GB, Q4 (experimental)); DeepSeek V4.1 (967 GB, Q4 (experimental, datacenter only)). On Apple Silicon, unified memory counts toward this requirement.
Can I run DeepSeek V4.1 on an RTX 4090 (24 GB)?
DeepSeek V4.1's smallest variant needs about 172 GB, which exceeds a single RTX 4090 (24 GB). Use multiple GPUs, a higher-VRAM card, or Apple Silicon with large unified memory.
What quantization should I use for DeepSeek V4.1?
Q4_K_M is the best balance of quality and VRAM for DeepSeek V4.1 in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run DeepSeek V4.1 with Ollama?
DeepSeek V4.1 has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.