Mistral AI24BQ4_K_M 下约 15 GB 显存

Devstral Small 2505 24B — 显存、速度与本地部署

作者: Jakub Rusinowski · 最后更新:

The first-generation Devstral Small (release 2505), and still the local coding model with the clearest published number attached to it: 46.8% on SWE-Bench Verified, which beat the prior open-source state of the art by 6 points at release. Finetuned from Mistral Small 3.1, so it inherits the 128K context. At 24B dense it is light enough for a single RTX 4090 or a 32GB Mac (~14 GB at Q4_K_M). Apache 2.0.

Devstral Small 2505 24B 在 Q4_K_M 下约需 15 GB 显存——量化权重加框架开销,不含 KV 缓存。在 Apple Silicon 上,这部分来自统一内存。

逻辑79
创意72
编程85

按量化级别的显存与速度

计算基准:NVIDIA RTX 4090 (24 GB)。包含 8K 上下文的 KV 缓存,因此数值高于上方的主数字。

量化显存速度(估算)适配
Q2_K
2.63 bpw
10 GB~74 tok/s可运行
Q3_K_M
3.41 bpw
12.4 GB~60 tok/s可运行
Q4_K_M
4.83 bpw
16.6 GB~45 tok/s可运行
Q5_K_M
5.67 bpw
19.2 GB~39 tok/s可运行
Q6_K
6.56 bpw
21.8 GB~34 tok/s勉强
Q8_0
8.50 bpw
27.6 GB~4 tok/s需卸载
F16
16.00 bpw
50.1 GB~2 tok/s需卸载

黑色标记 = NVIDIA RTX 4090 (24 GB) 上的可用显存。 估算来自内存带宽屋顶线模型,详见 方法说明页. Devstral Small 2505 24B 显存计算器 →

运行 Devstral Small 2505 24B

目录中能运行 Devstral Small 2505 24B 的最便宜 GPU 是 AMD Radeon RX 9060 XT 16GB (16 GB).

购买此硬件 AMD Radeon RX 9060 XT 16GB — 16 GB VRAM · 160 W board power立即云端部署 RunPod 上的 RTX 4090 — 低至 $0.34/小时 · 价格核实于 2026-07

或在 Vast.ai 比较,低至 $0.35/小时 (typical low · varies)

作为亚马逊联盟成员,我们从符合条件的购买中获得收入。云 GPU 链接为推荐链接——我们可能获得佣金,您无需额外付费。

联盟营销声明: 本页部分链接为联盟推广链接——如果你通过它们购买,LLM Configurator 可能会获得佣金,而你无需支付任何额外费用。作为亚马逊联盟成员(Amazon Associate),LLM Configurator 会从符合条件的购买中获得收益。
AMD Radeon RX 9060 XT 16GB
16 GB VRAM · 160 W board power
2026年价格波动较大——请以当前商品页价格为准。

如何运行 Devstral Small 2505 24B

安装 Ollama,然后运行:

ollama run devstral:24b
Hugging Face 上的权重: mistralai/Devstral-Small-2505 ↗

规格

Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.

Preview — The model is released, but these specs are thin or rest on a single source. Individual fields may be wrong.

参数量
24 Billion
上下文窗口
128,000
架构
Dense Transformer
提供商
Mistral AI
许可证
Apache 2.0
规格量化
Q4_K_M
系统内存
32 GB
记录更新于
2026-08-15
许可证Apache-2.0允许商业使用

Commercial use permitted. No usage restrictions beyond attribution.

质量与使用场景

评分由模型作者或独立评测方发布——衡量质量而非吞吐量,并非我们实测。

最适合software engineeringagentic codingconsumer gpulocal inference
基准分数来源
SWE-Bench Verified46.8 / 100 %已发布 · https://mistral.ai/news/devstral/

我的 GPU 能运行 Devstral Small 2505 24B 吗?

Devstral 的其他尺寸

Devstral Small 2505 24B — 常见问题

How much VRAM does Devstral Small 2505 24B need?

About 15 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Devstral Small 2505 24B run on an RTX 4090 (24 GB)?

Yes. Devstral Small 2505 24B needs about 15 GB at Q4_K_M, inside a 24 GB card, at an estimated 45 tokens/sec.

How do I run Devstral Small 2505 24B locally?

Install Ollama and run `ollama run devstral:24b`. That pulls the weights and starts a local OpenAI-compatible endpoint; after the download nothing leaves the machine.

What other sizes does Devstral come in?

Devstral-2 123B (75 GB), Devstral 2 22B (14 GB), Devstral Small 2505 24B (15 GB), Devstral Small 2 24B (15 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.