作者: Jakub Rusinowski · 最后更新: 2024年10月16日
Mistral AI针对资源受限环境的轻量级模型系列。Ministral 8B的推理任务表现远超其规格,在6 GB VRAM GPU上即可运行。Ministral 3B专为手机和边缘设备设计。
| Licence | What it permits | Applies to |
|---|---|---|
Mistral Research | Research / non-commercial only Research / non-commercial only — this licence does NOT permit shipping a commercial product. | Ministral 3B, Ministral 8B |
| Ministral 3B | Min 3 GB VRAM · Q4_K_M · 32,768 ctx · ollama run ministral:3b |
| Ministral 8B | Min 6 GB VRAM · Q4_K_M · 32,768 ctx · ollama run ministral:8b |
The cheapest GPU that runs Ministral locally (min 3 GB VRAM) is the Intel Arc B570 (10 GB).
Install Ollama then run: ollama run ministral:3b
Minimum VRAM: 3 GB. For best results use Q4_K_M quantization.
Ministral needs about 3 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Ministral 3B (3 GB, Q4_K_M); Ministral 8B (6 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Yes — Ministral runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
Q4_K_M is the best balance of quality and VRAM for Ministral in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
Install Ollama, then run: ollama run ministral:3b. This downloads Ministral and starts a local, OpenAI-compatible endpoint — no internet connection is needed after the initial download.