Aleph Alpha78B (3.5B active)~48 GB VRAM at Q4_K_M
Licence: Apache-2.0Local route: Official local, custom pluginHardware: ServerMaturity: Early releaseEvidence: Limited

Evidence: benchmark results vendor-reported

Kolibri 1 — VRAM, speed & local setup

Written by Jakub Rusinowski · Last updated

Aleph Alpha's German and English reasoning mixture-of-experts model: 78B parameters in total, about 3.5B active per token, 384 routed experts plus one shared expert per layer. Reasoning effort (none, low, medium, high) and tool calling are built in. Officially it runs through Aleph Alpha's pinned vLLM plugin; stock llama.cpp, Ollama and LM Studio do not load it yet. Benchmark results come from Aleph Alpha's own model card.

Kolibri 1 needs about 48 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

VRAM and speed by quantization

Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.

QuantVRAMSpeed (est.)Fit
Q2_K
2.63 bpw
26.6 GB~39 tok/sOffload
Q3_K_M
3.41 bpw
34.3 GB~35 tok/sOffload
Q4_K_M
4.83 bpw
48.4 GB~29 tok/sOffload
Q5_K_M
5.67 bpw
56.3 GB—Too big
Q6_K
6.56 bpw
65 GB—Too big
Q8_0
8.50 bpw
84.1 GB—Too big
F16
16.00 bpw
157.2 GB—Too big

Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. Kolibri 1 VRAM calculator →

Get Kolibri 1 running

The cheapest catalogued GPU that runs Kolibri 1 is the NVIDIA DGX Spark (128 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
NVIDIA DGX Spark (128GB)
128 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.

How to run Kolibri 1

No Ollama or LM Studio route today: stock llama.cpp does not support the architecture. Official route: vLLM with Aleph Alpha's aleph-alpha-inference plugin, which pins one vLLM minor version. Community routes: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX conversions that need about 64 GB of unified memory.

Weights on Hugging Face: Aleph-Alpha/Kolibri-1 ↗

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.

Parameters
78.1 Billion (3.46B active)
Context window
262,144
Architecture
Mixture-of-Experts (384 routed + 1 shared, 4:1 sliding-window to full attention)
Provider
Aleph Alpha
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
64 GB
Record updated
2026-10-06
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Limits and caveats

  • Released on 3 October 2026; maturity is 'Early release'.
  • No stock Ollama, llama.cpp or LM Studio path; the official route needs Aleph Alpha's version-pinned vLLM plugin, and MLX builds are community conversions.
  • German and English only (deliberate depth over breadth); text only.
  • Weights-and-config licence grant only; code, architecture and training method stay with Aleph Alpha.
  • The community GGUF is experimental and unverified on GPU backends and for contexts above 8,192 tokens.

Quality and use cases

Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.

Best forreasoningagentictool usemultilinguallong context

Kolibri 1 — frequently asked questions

How much VRAM does Kolibri 1 need?

About 48 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Kolibri 1 run on an RTX 4090 (24 GB)?

No. Kolibri 1 needs about 48 GB at Q4_K_M, more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.

How do I run Kolibri 1 locally?

No Ollama or LM Studio route today: stock llama.cpp does not support the architecture. Official route: vLLM with Aleph Alpha's aleph-alpha-inference plugin, which pins one vLLM minor version. Community routes: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX conversions that need about 64 GB of unified memory. Running the published tag would send your prompts to a hosted GPU rather than your own machine.