Mixture of Experts (MoE) Models Explained: Local AI 2026

By Jakub RusinowskiLast updated 7 min readAdvanced

MoE (Mixture of Experts) architecture is behind some of the most exciting models of 2026: Llama 4, DeepSeek V3, Qwen 3 MoE, and Gemma 4. Understanding MoE helps you choose the right model and manage your VRAM more effectively.

What Is Mixture of Experts?

In a traditional "dense" LLM, every parameter is active for every token. In a Mixture of Experts model, the architecture is split into many specialized sub-networks (the "experts"), and only a small fraction are activated per token:

Dense model (e.g., Llama 3.1 8B):

  • All 8 billion parameters active per token
  • Every layer fully participates in every computation

MoE model (e.g., Qwen 3.5 35B-A3B):

  • 35 billion total parameters stored in memory
  • Only ~3 billion parameters active per token ("A3B" = Active 3B)
  • A router network selects which 2–4 experts handle each token

Why MoE Models Are Exciting

The key insight: you store more knowledge (more total parameters) but only pay the compute cost of a smaller model.

PropertyDense 7BMoE 35B-A3B
Total parameters7B35B
Active parameters/token7B~3B
VRAM needed (Q4)~5 GB~20 GB
Inference speed (tokens/sec)60 t/s (RTX 4090)25–35 t/s (RTX 4090)
Knowledge/capability7B level~25–30B equivalent
Reasoning qualityGoodSignificantly better

You get the knowledge of a 35B model at roughly the inference cost of a 7B model — at the expense of more VRAM.

Major MoE Models in 2026

Llama 4 Scout (17B active / 109B total)

  • Active params: 17B per token
  • Total params: 109B across 16 experts
  • VRAM (Q4): ~10.5 GB
  • Best for: Coding, reasoning tasks where you want more than 8B quality without 70B hardware
  • Context: 128k tokens
bash
ollama run llama4:scout

Llama 4 Maverick (17B active / 400B total)

  • Active params: 17B per token
  • Total params: 400B across many experts
  • VRAM (Q4): ~242 GB — every one of the 400B parameters must be resident, even though only 17B activate per token
  • Best for: Multi-GPU servers or 256 GB+ unified-memory systems
  • Notable: Strong multimodal capabilities
bash
ollama run llama4:maverick

Qwen 3.5 35B-A3B

  • Active params: 3B per token
  • Total params: 35B
  • VRAM (Q4): ~20 GB
  • Speed: ~30 t/s on RTX 4090 (fast despite large total size)
  • Best for: Users with 24GB GPU wanting best quality/speed balance

DeepSeek V3.2 (685B total, 37B active)

  • The flagship open-weight MoE model
  • VRAM: ~370 GB total — only feasible via cloud or multi-GPU cluster
  • Speed: Runs efficiently on large clusters due to low active parameters
  • Available via DeepSeek API or cloud GPU rental

Gemma 4 26B-A4B

  • Google's MoE entry in the Gemma 4 family
  • Active: 4B params per token out of 26B total
  • VRAM (Q4): ~16 GB (fits RTX 4060 Ti 16GB)
  • Excellent multimodal support

Mixtral 8x7B (Classic)

The original widely-adopted MoE model:

  • 8 experts, 2 active per token
  • ~26B total params, ~13B active VRAM
  • VRAM: ~14GB at Q4 (fits dual-GPU setups or 16GB cards)
  • Still competitive for many tasks

VRAM vs Compute: The MoE Tradeoff

MoE models have an unusual VRAM profile:

shell
Dense Llama 3.1 70B Q4:
  VRAM: 40 GB (must fit all weights)
  Compute/token: 70B ops
  Typical speed: 5-8 t/s on RTX 4090

MoE Llama 4 Scout Q4:
  VRAM: 66.6 GB (109B total params - ALL experts resident)
  Bytes streamed/token: 10.3 GB (17B active params)
  Runs on: H100 80GB, or 96 GB+ unified memory

MoE Llama 4 Maverick Q4:
  VRAM: 242 GB (400B total params - ALL experts resident)
  Bytes streamed/token: 10.3 GB (17B active params)
  Runs on: multi-GPU server, or 256 GB+ unified memory

Critical insight: MoE reduces compute (speed) but not VRAM — all experts must be loaded into memory even if only 2–4 are used per token. A 685B MoE model still needs ~370GB VRAM, even though only 37B parameters run per token.

When to Choose MoE vs Dense

Choose MoE when:

  • You have enough VRAM to load the full model
  • You want maximum capability from your hardware
  • You're running batch inference where throughput matters less than quality
  • You need stronger reasoning than same-VRAM dense models provide

Choose Dense when:

  • VRAM is borderline — dense models are more predictable
  • You need maximum tokens/second (interactive chat)
  • You're doing fine-tuning (MoE is harder to fine-tune)
  • You want smaller total model storage on disk

Running MoE Models Efficiently

Ollama handles MoE models automatically. Some tips:

bash
# Check available VRAM before loading large MoE
ollama list

# For Llama 4 Scout (10.5GB) on an 8GB card - it will CPU offload
# Force full GPU offload if you have enough VRAM
OLLAMA_NUM_GPU=99 ollama run llama4:scout

# For Qwen 3.5 35B-A3B on RTX 4090 24GB:
ollama run qwen3.5:35b-a3b-q4_K_M

MoE and Context Length

MoE models often have very large context windows (because the architecture scales well):

ModelContext Window
Llama 4 Scout128k tokens
Llama 4 Maverick128k tokens
DeepSeek V3.2128k tokens
Qwen 3.5 35B-A3B32k tokens
Gemma 4 26B-A4B128k tokens

Large context + MoE efficiency = excellent for RAG, document analysis, and long-form tasks.

Frequently Asked Questions

Does Ollama support MoE models automatically? Yes — Ollama handles MoE models transparently via llama.cpp. You don't need any special configuration.

Is MoE better than dense for my use case? For most interactive chat tasks, a dense 8B model is faster. For complex reasoning where you can wait 2–3 seconds per response, an MoE model with the same VRAM budget will give better answers.

Can I fine-tune MoE models? Technically yes with frameworks like LLaMA-Factory, but it's more complex. LoRA fine-tuning works but requires careful expert routing. Most users fine-tune dense models instead.

Why does a 685B model need 370GB VRAM if only 37B are active? All expert weight matrices must be in VRAM for the router to select from them. You can't know in advance which experts will be selected, so all must be loaded. This is the fundamental MoE memory constraint.