Evidence: benchmark results vendor-reported
Kolibri 1 — VRAM, speed & local setup
Written by Jakub Rusinowski · Last updated
Aleph Alpha's German and English reasoning mixture-of-experts model: 78B parameters in total, about 3.5B active per token, 384 routed experts plus one shared expert per layer. Reasoning effort (none, low, medium, high) and tool calling are built in. Officially it runs through Aleph Alpha's pinned vLLM plugin; stock llama.cpp, Ollama and LM Studio do not load it yet. Benchmark results come from Aleph Alpha's own model card.
Kolibri 1 needs about 48 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
VRAM and speed by quantization
Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.
| Quant | Memory | VRAM | Speed (est.) | Fit |
|---|---|---|---|---|
| Q2_K 2.63 bpw | 26.6 GB | ~39 tok/s | Offload | |
| Q3_K_M 3.41 bpw | 34.3 GB | ~35 tok/s | Offload | |
| Q4_K_M 4.83 bpw | 48.4 GB | ~29 tok/s | Offload | |
| Q5_K_M 5.67 bpw | 56.3 GB | — | Too big | |
| Q6_K 6.56 bpw | 65 GB | — | Too big | |
| Q8_0 8.50 bpw | 84.1 GB | — | Too big | |
| F16 16.00 bpw | 157.2 GB | — | Too big |
Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. Kolibri 1 VRAM calculator →
Get Kolibri 1 running
The cheapest catalogued GPU that runs Kolibri 1 is the NVIDIA DGX Spark (128 GB).
or compare on Vast.ai from $0.77/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
How to run Kolibri 1
No Ollama or LM Studio route today: stock llama.cpp does not support the architecture. Official route: vLLM with Aleph Alpha's aleph-alpha-inference plugin, which pins one vLLM minor version. Community routes: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX conversions that need about 64 GB of unified memory.
Specifications
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.
- Parameters
- 78.1 Billion (3.46B active)
- Context window
- 262,144
- Architecture
- Mixture-of-Experts (384 routed + 1 shared, 4:1 sliding-window to full attention)
- Provider
- Aleph Alpha
- Licence
- Apache-2.0
- Specified at
- Q4_K_M
- System RAM
- 64 GB
- Record updated
- 2026-10-06
Commercial use permitted. No usage restrictions beyond attribution.
Limits and caveats
- Released on 3 October 2026; maturity is 'Early release'.
- No stock Ollama, llama.cpp or LM Studio path; the official route needs Aleph Alpha's version-pinned vLLM plugin, and MLX builds are community conversions.
- German and English only (deliberate depth over breadth); text only.
- Weights-and-config licence grant only; code, architecture and training method stay with Aleph Alpha.
- The community GGUF is experimental and unverified on GPU backends and for contexts above 8,192 tokens.
Quality and use cases
Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.
Kolibri 1 — frequently asked questions
How much VRAM does Kolibri 1 need?
About 48 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.
Does Kolibri 1 run on an RTX 4090 (24 GB)?
No. Kolibri 1 needs about 48 GB at Q4_K_M, more than a single RTX 4090's 24 GB. It needs a larger card, several GPUs, or Apple Silicon with enough unified memory — or it runs with part of the weights offloaded to system RAM, which is much slower.
How do I run Kolibri 1 locally?
No Ollama or LM Studio route today: stock llama.cpp does not support the architecture. Official route: vLLM with Aleph Alpha's aleph-alpha-inference plugin, which pins one vLLM minor version. Community routes: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX conversions that need about 64 GB of unified memory. Running the published tag would send your prompts to a hosted GPU rather than your own machine.