Devstral — Local AI Model by Mistral AI
作者: Jakub Rusinowski · 最后更新: 2026年9月11日
Mistral AI's coding specialists, spanning both generations. The original Devstral Small 2505 posts 46.8% on SWE-Bench Verified; its December 2025 successor Devstral Small 2 lifts that to 68.0% at the same 24B, which is the local-agent sweet spot — about 15 GB at Q4_K_M on one consumer card. The 123B flagship reaches 71.6%. Built for agentic software engineering: multi-file editing, repo navigation, and test-driven development. Apache 2.0 throughout.
Licence
| Licence | What it permits | Applies to |
|---|---|---|
Apache-2.0 | Commercial use permitted Commercial use permitted. No usage restrictions beyond attribution. | Devstral-2 123B, Devstral 2 22B, Devstral Small 2505 24B, Devstral Small 2 24B |
Hardware Requirements
| Devstral-2 123B | Min 75 GB VRAM · Q4_K_M · 262,144 ctx · ollama run devstral:123b |
| Devstral 2 22B | Min 14 GB VRAM · Q4_K_M · 128,000 ctx · |
| Devstral Small 2505 24B | Min 15 GB VRAM · Q4_K_M · 128,000 ctx · ollama run devstral:24b |
| Devstral Small 2 24B | Min 15 GB VRAM · Q4_K_M · 262,144 ctx · ollama run devstral-small-2:24b |
Recommended GPU
The cheapest GPU that runs Devstral locally (min 14 GB VRAM) is the AMD Radeon RX 9060 XT 16GB (16 GB).
How to Run Locally
Install Ollama then run: ollama run devstral:123b
Minimum VRAM: 14 GB. For best results use Q4_K_M quantization.
Devstral — Frequently Asked Questions
How much VRAM does Devstral need?
Devstral needs about 14 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Devstral-2 123B (75 GB, Q4_K_M); Devstral 2 22B (14 GB, Q4_K_M); Devstral Small 2505 24B (15 GB, Q4_K_M); Devstral Small 2 24B (15 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Devstral on an RTX 4090 (24 GB)?
Yes — Devstral runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
What quantization should I use for Devstral?
Q4_K_M is the best balance of quality and VRAM for Devstral in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Devstral with Ollama?
Install Ollama, then run: ollama run devstral:123b. This downloads Devstral and starts a local, OpenAI-compatible endpoint — no internet connection is needed after the initial download.