Arcee AI398B (13B active)~241 GB VRAM at Q4_K_M
Licence: OpenMDW-1.1Local route: Official localHardware: Datacenter-only in practiceMaturity: CurrentEvidence: Moderate

Evidence: benchmark results vendor-reported

Trinity-Large-Thinking — VRAM, speed & local setup

Written by Jakub Rusinowski · Last updated

The reasoning and agent post-trained checkpoint of Arcee's Trinity-Large: about 398B parameters in total, about 13B active per token, 256 experts with 4 active. Thinking tokens inside <think> blocks must be kept in the conversation history for agent loops to work. The repository is about 797 GB in BF16, so this is a multi-GPU datacenter model. The older Trinity-Large-Preview is a different checkpoint and is deliberately not listed.

Trinity-Large-Thinking needs about 241 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

VRAM and speed by quantization

Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.

QuantVRAMSpeed (est.)Fit
Q2_K
2.63 bpw
132.4 GB—Too big
Q3_K_M
3.41 bpw
171.2 GB—Too big
Q4_K_M
4.83 bpw
242 GB—Too big
Q5_K_M
5.67 bpw
283.8 GB—Too big
Q6_K
6.56 bpw
328.2 GB—Too big
Q8_0
8.50 bpw
424.9 GB—Too big
F16
16.00 bpw
798.6 GB—Too big

Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. Trinity-Large-Thinking VRAM calculator →

Get Trinity-Large-Thinking running

The cheapest catalogued GPU that runs Trinity-Large-Thinking is the Apple M3 Ultra (512 GB).

Buy this hardware Apple Mac Studio M3 Ultra — 512 GB VRAM · 60 W board powerDeploy in the cloud now RunPod

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Apple Mac Studio M3 Ultra
512 GB VRAM · 60 W board power
2026 prices are volatile — check the current listing.

How to run Trinity-Large-Thinking

No GGUF or Ollama route found. Official serving is vLLM 0.11.1 or newer on a multi-GPU node.

Weights on Hugging Face: arcee-ai/Trinity-Large-Thinking ↗

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source. Still unconfirmed: languages.

Parameters
398.6 Billion (13B active)
Context window
262,144
Architecture
Mixture-of-Experts (256 routed + 1 shared, interleaved local and global attention)
Provider
Arcee AI
Licence
OpenMDW-1.1
Specified at
Q4_K_M
System RAM
256 GB
Record updated
2026-10-06
LicenceOpenMDW-1.1Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Limits and caveats

  • Datacenter-class hardware only.
  • Thinking tokens must be preserved in agent loops or tool calls can malform.
  • The model card also advertises 512K context; the config and the tech report give 262,144, which is what this page uses.
  • The 'Trinity-Large-Preview' checkpoint is intentionally not listed as current.

Quality and use cases

Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.

Best forreasoningagentictool uselong context

Other Trinity sizes

Trinity-Large-Thinking — frequently asked questions

How much VRAM does Trinity-Large-Thinking need?

About 241 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Trinity-Large-Thinking run on an RTX 4090 (24 GB)?

No. Trinity-Large-Thinking needs about 241 GB at Q4_K_M, more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.

How do I run Trinity-Large-Thinking locally?

No GGUF or Ollama route found. Official serving is vLLM 0.11.1 or newer on a multi-GPU node. Running the published tag would send your prompts to a hosted GPU rather than your own machine.

What other sizes does Trinity come in?

Trinity Mini (17 GB), Trinity-Large-Thinking (241 GB). Every size shares the family's training and licence; the larger ones score higher and need proportionally more memory.