MoE (Mixture of Experts) architecture is behind some of the most exciting models of 2026: Llama 4, DeepSeek V3, Qwen 3 MoE, and Gemma 4. Understanding MoE helps you choose the right model and manage your VRAM more effectively.
What Is Mixture of Experts?
In a traditional "dense" LLM, every parameter is active for every token. In a Mixture of Experts model, the architecture is split into many specialized sub-networks (the "experts"), and only a small fraction are activated per token:
Dense model (e.g., Llama 3.1 8B):
- All 8 billion parameters active per token
- Every layer fully participates in every computation
MoE model (e.g., Qwen 3.5 35B-A3B):
- 35 billion total parameters stored in memory
- Only ~3 billion parameters active per token ("A3B" = Active 3B)
- A router network selects which 2–4 experts handle each token
Why MoE Models Are Exciting
The key insight: you store more knowledge (more total parameters) but only pay the compute cost of a smaller model.
| Property | Dense 7B | MoE 35B-A3B |
|---|---|---|
| Total parameters | 7B | 35B |
| Active parameters/token | 7B | ~3B |
| VRAM needed (Q4) | ~5 GB | ~20 GB |
| Inference speed (tokens/sec) | 60 t/s (RTX 4090) | 25–35 t/s (RTX 4090) |
| Knowledge/capability | 7B level | ~25–30B equivalent |
| Reasoning quality | Good | Significantly better |
You get the knowledge of a 35B model at roughly the inference cost of a 7B model — at the expense of more VRAM.
Major MoE Models in 2026
Llama 4 Scout (17B active / 109B total)
- Active params: 17B per token
- Total params: 109B across 16 experts
- VRAM (Q4): ~10.5 GB
- Best for: Coding, reasoning tasks where you want more than 8B quality without 70B hardware
- Context: 128k tokens
ollama run llama4:scoutLlama 4 Maverick (17B active / 400B total)
- Active params: 17B per token
- Total params: 400B across many experts
- VRAM (Q4): ~242 GB — every one of the 400B parameters must be resident, even though only 17B activate per token
- Best for: Multi-GPU servers or 256 GB+ unified-memory systems
- Notable: Strong multimodal capabilities
ollama run llama4:maverickQwen 3.5 35B-A3B
- Active params: 3B per token
- Total params: 35B
- VRAM (Q4): ~20 GB
- Speed: ~30 t/s on RTX 4090 (fast despite large total size)
- Best for: Users with 24GB GPU wanting best quality/speed balance
DeepSeek V3.2 (685B total, 37B active)
- The flagship open-weight MoE model
- VRAM: ~370 GB total — only feasible via cloud or multi-GPU cluster
- Speed: Runs efficiently on large clusters due to low active parameters
- Available via DeepSeek API or cloud GPU rental
Gemma 4 26B-A4B
- Google's MoE entry in the Gemma 4 family
- Active: 4B params per token out of 26B total
- VRAM (Q4): ~16 GB (fits RTX 4060 Ti 16GB)
- Excellent multimodal support
Mixtral 8x7B (Classic)
The original widely-adopted MoE model:
- 8 experts, 2 active per token
- ~26B total params, ~13B active VRAM
- VRAM: ~14GB at Q4 (fits dual-GPU setups or 16GB cards)
- Still competitive for many tasks
VRAM vs Compute: The MoE Tradeoff
MoE models have an unusual VRAM profile:
Dense Llama 3.1 70B Q4:
VRAM: 40 GB (must fit all weights)
Compute/token: 70B ops
Typical speed: 5-8 t/s on RTX 4090
MoE Llama 4 Scout Q4:
VRAM: 66.6 GB (109B total params - ALL experts resident)
Bytes streamed/token: 10.3 GB (17B active params)
Runs on: H100 80GB, or 96 GB+ unified memory
MoE Llama 4 Maverick Q4:
VRAM: 242 GB (400B total params - ALL experts resident)
Bytes streamed/token: 10.3 GB (17B active params)
Runs on: multi-GPU server, or 256 GB+ unified memoryCritical insight: MoE reduces compute (speed) but not VRAM — all experts must be loaded into memory even if only 2–4 are used per token. A 685B MoE model still needs ~370GB VRAM, even though only 37B parameters run per token.
When to Choose MoE vs Dense
Choose MoE when:
- You have enough VRAM to load the full model
- You want maximum capability from your hardware
- You're running batch inference where throughput matters less than quality
- You need stronger reasoning than same-VRAM dense models provide
Choose Dense when:
- VRAM is borderline — dense models are more predictable
- You need maximum tokens/second (interactive chat)
- You're doing fine-tuning (MoE is harder to fine-tune)
- You want smaller total model storage on disk
Running MoE Models Efficiently
Ollama handles MoE models automatically. Some tips:
# Check available VRAM before loading large MoE
ollama list
# For Llama 4 Scout (10.5GB) on an 8GB card - it will CPU offload
# Force full GPU offload if you have enough VRAM
OLLAMA_NUM_GPU=99 ollama run llama4:scout
# For Qwen 3.5 35B-A3B on RTX 4090 24GB:
ollama run qwen3.5:35b-a3b-q4_K_MMoE and Context Length
MoE models often have very large context windows (because the architecture scales well):
| Model | Context Window |
|---|---|
| Llama 4 Scout | 128k tokens |
| Llama 4 Maverick | 128k tokens |
| DeepSeek V3.2 | 128k tokens |
| Qwen 3.5 35B-A3B | 32k tokens |
| Gemma 4 26B-A4B | 128k tokens |
Large context + MoE efficiency = excellent for RAG, document analysis, and long-form tasks.
Frequently Asked Questions
Does Ollama support MoE models automatically? Yes — Ollama handles MoE models transparently via llama.cpp. You don't need any special configuration.
Is MoE better than dense for my use case? For most interactive chat tasks, a dense 8B model is faster. For complex reasoning where you can wait 2–3 seconds per response, an MoE model with the same VRAM budget will give better answers.
Can I fine-tune MoE models? Technically yes with frameworks like LLaMA-Factory, but it's more complex. LoRA fine-tuning works but requires careful expert routing. Most users fine-tune dense models instead.
Why does a 685B model need 370GB VRAM if only 37B are active? All expert weight matrices must be in VRAM for the router to select from them. You can't know in advance which experts will be selected, so all must be loaded. This is the fundamental MoE memory constraint.