Evidence: benchmark results vendor-reported
Trinity-Large-Thinking — VRAM, speed & local setup
Written by Jakub Rusinowski · Last updated
The reasoning and agent post-trained checkpoint of Arcee's Trinity-Large: about 398B parameters in total, about 13B active per token, 256 experts with 4 active. Thinking tokens inside <think> blocks must be kept in the conversation history for agent loops to work. The repository is about 797 GB in BF16, so this is a multi-GPU datacenter model. The older Trinity-Large-Preview is a different checkpoint and is deliberately not listed.
Trinity-Large-Thinking needs about 241 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
VRAM and speed by quantization
Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.
| Quant | Memory | VRAM | Speed (est.) | Fit |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 132.4 GB | — | Too big | |
| Q3_K_M 3.41 bpw | 171.2 GB | — | Too big | |
| Q4_K_M 4.83 bpw | 242 GB | — | Too big | |
| Q5_K_M 5.67 bpw | 283.8 GB | — | Too big | |
| Q6_K 6.56 bpw | 328.2 GB | — | Too big | |
| Q8_0 8.50 bpw | 424.9 GB | — | Too big | |
| F16 16.00 bpw | 798.6 GB | — | Too big |
Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. Trinity-Large-Thinking VRAM calculator →
Get Trinity-Large-Thinking running
The cheapest catalogued GPU that runs Trinity-Large-Thinking is the Apple M3 Ultra (512 GB).
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
How to run Trinity-Large-Thinking
No GGUF or Ollama route found. Official serving is vLLM 0.11.1 or newer on a multi-GPU node.
Specifications
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: languages.
- Parameters
- 398.6 Billion (13B active)
- Context window
- 262,144
- Architecture
- Mixture-of-Experts (256 routed + 1 shared, interleaved local and global attention)
- Provider
- Arcee AI
- Licence
- OpenMDW-1.1
- Specified at
- Q4_K_M
- System RAM
- 256 GB
- Record updated
- 2026-10-06
Commercial use permitted. No usage restrictions beyond attribution.
Limits and caveats
- Datacenter-class hardware only.
- Thinking tokens must be preserved in agent loops or tool calls can malform.
- The model card also advertises 512K context; the config and the tech report give 262,144, which is what this page uses.
- The 'Trinity-Large-Preview' checkpoint is intentionally not listed as current.
Quality and use cases
Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.
Other Trinity sizes
Trinity-Large-Thinking — frequently asked questions
How much VRAM does Trinity-Large-Thinking need?
About 241 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Trinity-Large-Thinking run on an RTX 4090 (24 GB)?
No. Trinity-Large-Thinking needs about 241 GB at Q4_K_M, more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.
How do I run Trinity-Large-Thinking locally?
No GGUF or Ollama route found. Official serving is vLLM 0.11.1 or newer on a multi-GPU node. Running the published tag would send your prompts to a hosted GPU rather than your own machine.
What other sizes does Trinity come in?
Trinity Mini (17 GB), Trinity-Large-Thinking (241 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.