Mistral Large 3 — Local AI Model by Mistral AI
作者: Jakub Rusinowski · 最后更新: 2026年6月26日
Mistral's open-weight flagship, released 2 December 2025: a 675B MoE with 41B active parameters, Apache-2.0, 256K context. Datacentre-scale to self-host; most readers will rent it.
Licence
| Licence | What it permits | Applies to |
|---|---|---|
Apache-2.0 | Commercial use permitted Commercial use permitted. No usage restrictions beyond attribution. | Mistral Large 3 675B-A41B |
Hardware Requirements
| Mistral Large 3 675B-A41B | Min 408 GB VRAM · Q4_K_M · 256,000 ctx · |
Recommended GPU
The cheapest GPU that runs Mistral Large 3 locally (min 408 GB VRAM) is the Apple M3 Ultra (512 GB).
How to Run Locally
Install Ollama then run: ollama run mistral-large-3
Minimum VRAM: 408 GB. For best results use Q4_K_M quantization.
Mistral Large 3 — Frequently Asked Questions
How much VRAM does Mistral Large 3 need?
Mistral Large 3 needs about 408 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Mistral Large 3 675B-A41B (408 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Mistral Large 3 on an RTX 4090 (24 GB)?
Mistral Large 3's smallest variant needs about 408 GB, which exceeds a single RTX 4090 (24 GB). Use multiple GPUs, a higher-VRAM card, or Apple Silicon with large unified memory.
What quantization should I use for Mistral Large 3?
Q4_K_M is the best balance of quality and VRAM for Mistral Large 3 in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Mistral Large 3 with Ollama?
Mistral Large 3 has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.