Trinity — local AI model by Arcee AI
Written by Jakub Rusinowski · Last updated
Trinity is Arcee AI's open-weight sparse mixture-of-experts family. Trinity Mini (26B total, about 3B active) is the locally practical member with an official GGUF; Trinity-Large-Thinking (about 398B total, about 13B active) is a reasoning and agent model for multi-GPU servers. The whole family moved from Apache-2.0 to OpenMDW-1.1 on 29 May 2026, so older Apache wording is stale.
Variants
The smallest Trinity variant needs about 17 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| Trinity Mini → 26B (3B active) | ~16.7 GB | 131,072 | cloud-hosted tag | |
| Trinity-Large-Thinking → 398B (13B active) | ~241.5 GB | 262,144 | cloud-hosted tag |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run Trinity locally
Install Ollama, then pull the tag.
No verified Ollama library tag. Arcee publishes official GGUF files that work with llama.cpp (b7061 or newer) and LM Studio.
Pick a size above for its own VRAM figure, speed estimate and install command.
Licence
Commercial use permitted. No usage restrictions beyond attribution.
Applies to: Trinity Mini, Trinity-Large-ThinkingRecommended GPU
The cheapest catalogued GPU that runs Trinity locally (min 17 GB VRAM) is the AMD Radeon RX 7900 XT (20 GB).
Trinity — frequently asked questions
How much VRAM does Trinity need?
Trinity needs about 17 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Trinity Mini (17 GB, Q4_K_M); Trinity-Large-Thinking (241 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Trinity on an RTX 4090 (24 GB)?
Yes — Trinity runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
What quantization should I use for Trinity?
Q4_K_M is the best balance of quality and VRAM for Trinity in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Trinity with Ollama?
Trinity has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.