Arcee AI26B (3B active)~17 GB VRAM at Q4_K_M
Licence: OpenMDW-1.1Local route: Official localHardware: Consumer localMaturity: CurrentEvidence: Moderate

Evidence: benchmark results vendor-reported

Trinity Mini — VRAM, speed & local setup

Written by Jakub Rusinowski · Last updated

Arcee AI's 26B-total, about 3B-active mixture-of-experts model with 128 routed experts (8 active) and one shared expert. It is tuned for reasoning, supports a 128K context, and Arcee publishes an official GGUF set that works with llama.cpp (b7061 or newer), plus vLLM, Transformers and LM Studio recipes. Licensed under OpenMDW-1.1. Benchmark charts on the card are vendor-reported.

Trinity Mini needs about 17 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

VRAM and speed by quantization

Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.

QuantVRAMSpeed (est.)Fit
Q2_K
2.63 bpw
9.5 GB~275 tok/sFits
Q3_K_M
3.41 bpw
13 GB~248 tok/sFits
Q4_K_M
4.83 bpw
16.9 GB~211 tok/sFits
Q5_K_M
5.67 bpw
19.6 GB~194 tok/sFits
Q6_K
6.56 bpw
22.5 GB~179 tok/sTight
Q8_0
8.50 bpw
28.7 GB~23 tok/sOffload
F16
16.00 bpw
53.2 GB~15 tok/sOffload

Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. Trinity Mini VRAM calculator →

Get Trinity Mini running

The cheapest catalogued GPU that runs Trinity Mini is the AMD Radeon RX 7900 XT (20 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 7900 XT 20GB
20 GB VRAM · 315 W board power
2026 prices are volatile — check the current listing.

How to run Trinity Mini

No verified Ollama library tag. Arcee publishes official GGUF files that work with llama.cpp (b7061 or newer) and LM Studio.

Weights on Hugging Face: arcee-ai/Trinity-Mini ↗

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.

Parameters
26.1 Billion (3B active)
Context window
131,072
Architecture
Mixture-of-Experts (128 routed + 1 shared, interleaved local and global attention)
Provider
Arcee AI
Licence
OpenMDW-1.1
Specified at
Q4_K_M
System RAM
32 GB
Record updated
2026-10-06
LicenceOpenMDW-1.1Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Limits and caveats

  • Benchmark claims are vendor-reported and only available as a chart image.
  • No verified Ollama library tag.
  • Card recommends unusually low sampling temperature (0.15); other runtimes may need these settings explicitly.

Quality and use cases

Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.

Best forreasoningagentictool usechat

Other Trinity sizes

Trinity Mini — frequently asked questions

How much VRAM does Trinity Mini need?

About 17 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Trinity Mini run on an RTX 4090 (24 GB)?

Yes. Trinity Mini needs about 17 GB at Q4_K_M, inside a 24 GB card, at an estimated 211 tokens/sec.

How do I run Trinity Mini locally?

No verified Ollama library tag. Arcee publishes official GGUF files that work with llama.cpp (b7061 or newer) and LM Studio. Running the published tag would send your prompts to a hosted GPU rather than your own machine.

What other sizes does Trinity come in?

Trinity Mini (17 GB), Trinity-Large-Thinking (241 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.