Kolibri 1 and the European take on open-weight models: what Aleph Alpha shipped, and what it takes to run it

Kolibri 1 is Aleph Alpha's Apache-2.0 German-English 78B MoE. What the licence covers, how to run it, what the benchmarks show, and what 'European' means here.

October 6, 202618 min readJakub Rusinowski
Kolibri 1 is a German-English reasoning model from Aleph Alpha, published on 3 October 2026 with Apache-2.0 weights. It is a mixture-of-experts model: 78.1 billion parameters in total, about 3.46 billion active per token, trained from scratch. It is a real European open-weight release, but it is not a laptop model. The weights alone are about 79 GB in FP8, officially it runs on a version-pinned vLLM plugin, and the routes onto smaller hardware are community conversions only: an experimental GGUF that needs a patched llama.cpp, and unofficial MLX builds that need a Mac with about 64 GB of unified memory. This post covers what shipped, what the licence does and doesn't grant, what the vendor's own benchmarks say, and what "European" actually means for a model like this.

If you want to know whether your machine can run it, check the Kolibri 1 page or the analyzer.


What shipped on 3 October

Kolibri 1
DeveloperAleph Alpha Research GmbH (provider: Aleph Alpha GmbH)
Released3 October 2026 (the Hugging Face repositories were created on 2 October)
ArchitectureMixture-of-experts, 50 layers, trained from scratch (no base-model dependency)
Parameters78,103,074,560 total · 3,457,573,120 active per token
Experts384 routed experts per layer, 6 chosen per token, plus 1 shared expert
LanguagesGerman and English (a deliberate choice of depth over breadth)
Context262,144 tokens native; validated by the vendor up to 1,048,576; recommended ≤ 262,144
FeaturesReasoning effort none / low / medium / high, tool calling, structured extraction
LicenceApache-2.0 for the published weights and configuration files
ArtifactsFP8 (about 79 GB) and BF16 (about 156 GB), both public and ungated
Official runtimevLLM, through Aleph Alpha's aleph-alpha-inference plugin
Every number above comes from the model cards and the Hugging Face API. The benchmark results later in this post are the vendor's own.

Why this is a European story

Aleph Alpha spent 2021 to 2024 as the company people pointed to when asked for "Germany's answer to OpenAI". In September 2024 it announced a pivot away from building its own models, and Reuters described it in April 2026 as having since abandoned development of large language models in favour of specialised business applications. Its last open-weight line before this was Pharia-1-LLM-7B, and that one shipped under the Open Aleph License 1.0, which allows non-commercial and non-administrative use only. That is why I haven't added Pharia to the library as a normal recommendation.

Kolibri changes the picture, and the timing is odd:

  • 24 April 2026: Reuters reports that Canadian AI company Cohere has agreed to buy Aleph Alpha at an undisclosed price, with Germany's Schwarz Group committing $600 million to Cohere's next funding round.
  • 16 September 2026: the two companies sign a definitive agreement. The combined company will operate as Cohere, with dual headquarters in Berlin and Toronto and Aleph Alpha's Heidelberg office kept as a research centre. Schwarz Group's commitment, reported in April as $600 million, is given as €500 million in the September coverage, and the companies say the AI services will run on STACKIT, the group's sovereign cloud. The deal is still subject to regulatory approval and is expected to close later in 2026.
  • 3 October 2026: Aleph Alpha publishes a from-scratch 78B model with Apache-2.0 weights.

So a company reported to have left model-building behind in April released a large model six months later, in the middle of being absorbed. I couldn't find any statement on what the combined company plans for the Kolibri line, and I'd treat that as unknown. What you can rely on is narrower: the weights already published are under Apache-2.0, a licence whose text describes the grant as perpetual and irrevocable. Future releases are another matter. (That's a reading of the licence text, not legal advice.)

What "European" does and doesn't mean here

"European model" can mean five different things. For Kolibri, from the model card:

LayerWhat the card saysWhat it means for you
CompanyGerman developer, pending combination with Canada's CohereThe ownership picture is changing; the weights are already out
Training data20T tokens, bilingual: about 62.5% English, 23.9% German, 13.6% code; web data from Common Crawl, filtered, with PII redaction and decontaminationA German share this large is unusual for a model this size
Synthetic dataGenerated with permissively licensed open models: Gemma-4-26B-A4B (English rephrases), Mistral-NeMo-12B (German rephrases), Qwen3-32B (quality annotations)The pipeline leaned on American, French and Chinese open models
Compute768 NVIDIA B200 GPUs for 21 days (pre-training), about 6.4e23 FLOPs in total; the card doesn't name the data centreEuropean model, American silicon; location not stated
Compliance postureSignatory of the EU's General-Purpose AI Code of Practice (published by the Commission on 10 July 2025); a published copyright-rightsholder contact; an estimated 950 MWh for pre-training, mid-training and long-context trainingMore transparency than most open releases offer

That last row is the real European angle. The model card reads like it was written with an AI Act documentation checklist next to it: training-data categories, a rightsholder contact, an energy estimate, and an explicit statement that the model is built for systems where a person reviews its output rather than for unsupervised operation. None of that makes the model better. It makes it easier to evaluate for a regulated company, which is exactly the customer Aleph Alpha and Cohere say they are targeting. Whether it satisfies any given legal obligation is a question for your counsel, not for a model card.

"Sovereign" is a property of where you run a model, not of where it was made. Apache-2.0 weights can sit in a German data centre, on a STACKIT instance or in a cupboard. That is the whole point of open weights, and it is the reason I care about Kolibri's licence more than its origin.

What the licence actually grants

The card's licence section is short and unusually explicit. The weights are published under Apache-2.0, and the rights granted apply only to the weights and configuration files in the repository. It then states that the licence "especially does not extend to underlying code, model architecture, parameter settings or any training method", all of which Aleph Alpha keeps.

That is why this site says open-weight, not open source. In practice:

  • You can download, run, fine-tune, host and build products on the weights, commercially.
  • You do not get the training code, the recipe or a licence to reimplement the architecture as your own product. (You also need Aleph Alpha's plugin, which is published under Apache-2.0 on PyPI, to serve it with vLLM.)
  • Compare this with Arcee's Trinity, which moved to OpenMDW-1.1 in May 2026, or NVIDIA's Nemotron, which uses a custom licence. Apache-2.0, MIT, OpenMDW-1.1 and custom licences are different things, and the library now labels each one separately.

Inside the model

A mixture-of-experts model reads all 78.1B parameters into memory but computes with only 3.46B per token: 384 routed experts per layer, 6 selected, 1 shared
Figure 1. Kolibri 1 holds 78.1B parameters in memory and computes with 3.46B per token.
Mixture-of-experts, and why memory is the catch. Each of the 50 layers has 384 small routed experts and one shared expert. A router picks 6 routed experts per token, so only about 3.46B parameters do work for any one token. That is why a 78B model can be cheap to compute. It does not make it cheap to hold: every expert has to be loaded, because the next token may need any of them. Active parameters set speed; total parameters set the memory floor. The card says so itself.
Fifty layers drawn as a strip: four sliding-window layers (512-token window, positional encoding) then one full-attention layer (no positional encoding), repeated ten times
Figure 2. Only 10 of 50 layers use full attention, so only those hold a cache that grows with context.
Mostly local attention, which keeps long context affordable. Four out of five layers only look at the previous 512 tokens. Every fifth layer attends across the whole context and uses no positional encoding. That has two consequences. First, the KV cache, which normally grows with context and eats VRAM, only grows in 10 of the 50 layers. Second, because positional encoding lives only in the sliding-window layers, the vendor says the context can be extended past its trained 262,144 tokens without position scaling; it validated serving up to 1,048,576.

You can estimate the cost from the published config. With 4 KV heads, a head size of 128 and an FP8 KV cache (the setting Aleph Alpha uses in its evaluations), the 10 full-attention layers need 2 × 10 × 4 × 128 = 10,240 bytes per token, which is 10 KiB. At 262,144 tokens that is about 2.5 GiB, and at 1,048,576 tokens about 10 GiB. A BF16 cache doubles those figures. That's my arithmetic from the config, not a vendor figure, and the small bounded window caches in the other 40 layers come on top. The analyzer is the authority for fit verdicts.

A tokenizer built for German. Kolibri has a 128,000-token vocabulary trained with a new method Aleph Alpha calls UniBPE, which places more token boundaries on word-part boundaries. The technical report measures 4.90 bytes per token on German, the best of the twelve tokenizers it compares (GPT-5's is 4.35), and 4.58 on English, a little behind GPT-5's 4.67. Those are vendor measurements, but they are easy to check on your own documents, and fewer tokens per page means cheaper long German contexts. Reasoning you can dial. Set reasoning_effort to low, medium or high, or turn thinking off with none. The model was trained to reason in German when asked in German, with a language-consistency reward during reinforcement learning. Tool calling works together with reasoning. The recommended sampling settings are temperature 1.0, top_p 0.97 and top_k 128.

What Aleph Alpha's benchmarks say, and what they don't

Grouped bars of five vendor-reported benchmark rows for Kolibri 1 and six peers, labelled 'vendor-reported, single harness'
Figure 3. Vendor-reported results from Aleph Alpha's own model card. A dash is a missing score, not a zero.

Everything below comes from one table on Aleph Alpha's own model card. The card compares Kolibri with thirteen other models; I show six of them, chosen because they are the ones you are likely to be weighing it against. Aleph Alpha ran all models with the same evaluation setup, Kolibri at reasoning effort high. These are vendor-reported results, not independent rankings.

ModelOverall ENOverall DEGPQA Diamond (EN)SWE-Bench VerifiedTerminalBench 2.1
Kolibri 1 (3.5B active)75.570.884.366.427.7
Qwen3.5 35B-A3B74.769.883.871.639.7
Gemma 4 26B-A4B IT71.966.381.157.8–
GPT-OSS 120B72.370.276.4–29.2
Nemotron 3 Super 120B-A12B73.067.978.060.239.7
Mistral Small 4 119B-A6B63.161.474.760.821.0
Qwen3.8 27B (dense)80.279.989.272.676.8

A dash means the model produced no score (it can't call tools, or the run exceeded its window). "Overall" is an unweighted mean of category averages, as defined on the card.

Reading it fairly:

  • Among the eleven other sparse (mixture-of-experts) models in the card's table, Kolibri has the best overall English and German averages. That is the headline, and it matches the "German and English at low serving cost" pitch. The German margin over GPT-OSS 120B is small (70.8 against 70.2).
  • It doesn't win on agentic coding. On SWE-Bench Verified, Qwen3.5 35B-A3B (71.6) and Qwen3.6 35B-A3B (73.8, not shown above) are ahead of Kolibri's 66.4. On TerminalBench 2.1, Qwen3.5 and Nemotron 3 Super both score 39.7 against Kolibri's 27.7.
  • The dense 27B model is clearly ahead overall. The card lists Qwen3.8 27B as a dense baseline, with about 27B parameters active per token against Kolibri's roughly 3B. If memory isn't your constraint and compute is, that comparison matters more than the sparse one.
  • The FP8 and BF16 cards differ slightly. The BF16 card reports overall 75.0 (EN) and 71.0 (DE) against 75.5 and 70.8 on the FP8 card. Differences of that size are in the noise, and a reminder that these are single runs.

My rule on this site, as with the decision-model benchmarks, is to show vendor numbers labelled as vendor numbers and to say that the only benchmark that decides anything for you is a hundred of your own examples. For German in particular, run your own documents through it before you believe any table.

Can you run it?

Decision flow: do you have 2x80GB-class GPUs? yes: official vLLM plugin route. no: unofficial community routes (patched llama.cpp GGUF, or MLX on a 64 GB Mac), or rent a GPU
Figure 4. The official route needs 80 GB-class GPUs; every other route is community-made.
RouteMade byWhat you needStatus
vLLM + aleph-alpha-inference pluginAleph AlphaFP8: 2× A100 80 GB, 2× H100 SXM5, or 1× H200 / B200 / B300. BF16: 4× A100 80 GB, 4× H100, 2× H200, or 1× B200 / B300Official. The plugin pins one vLLM minor version (vLLM 0.29 at launch)
Community GGUF on patched llama.cppHob-forge (independent)The Q4_K_M file is 47.5 GB and Q8_0 is 83.1 GB, plus a llama.cpp built from commit 836d571 with the author's patch appliedExperimental, not endorsed by Aleph Alpha or llama.cpp
Community MLX conversionsSeveral independent authors on Hugging FaceA Mac with about 64 GB of unified memory (the 4-bit builds are about 45 GB)Unofficial, not endorsed by Aleph Alpha. I haven't run them
Ollama, LM Studio, stock llama.cppnonen/aNot supported yet: the architecture isn't in upstream llama.cpp
The official route is two commands once you have the GPUs:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice

For the BF16 weights, serve Aleph-Alpha/Kolibri-1-BF16 and drop the FP8 KV flag. There is also a container image at ghcr.io/aleph-alpha/aleph-alpha-inference.

The community routes are more interesting for anyone with a big-RAM box. The author of the Hob-forge GGUF reports that on a CPU-only desktop (Ryzen 7 7800X3D, 128 GB RAM) the Q4_K_M ran at roughly 12 to 15 tokens per second on short prompts and handled a German explanation, an arithmetic problem and a tool call correctly. They are equally clear about what they did not test: CUDA, Vulkan and Metal backends, contexts above 8,192 tokens, multiple tool calls in one turn, and quality against the original on benchmarks. Treat it as a proof that the architecture can be ported, not as a supported setup. If you try it, expect to rebuild when upstream support lands.

On a Mac, several people have published MLX conversions of the FP8 weights within days of the release. Their own cards put the 4-bit builds at about 45 GB, which is why they need 64 GB of unified memory. I haven't run any of them, and none comes from Aleph Alpha or from the main MLX community account, so check the card of the one you pick for the exact mlx-lm version it needs.

If you don't own the hardware, an hour on a rented H100- or H200-class GPU is a cheaper way to evaluate Kolibri on your own German documents than buying anything. RunPod and Vast.ai both rent that class of GPU by the hour. Rent GPUs on RunPod → Rent GPUs on Vast.ai → This page contains affiliate links.

My take

Kolibri is the first European open-weight release in a while that I think deserves a place next to the American and Chinese ones, and it is just as clearly not for everyone. Three things I find convincing: the licence is a real commercial one, the German-first tokenizer and data mix are a genuine design decision, and the documentation is the most audit-friendly I've seen on a model this size. Three things hold me back: it's three days old, the only official serving path is a pinned plugin, and every quality claim so far comes from the vendor.

Who it fits: a team with access to 2× H100-class hardware (owned, rented or on a European cloud) that needs German and English document work, retrieval-augmented generation or agentic tool use, and values being able to show a regulator where the model came from. Who it doesn't: anyone who wants a model on a 24 GB card or an ordinary laptop this month. For that, the library already has smaller options.

What I'll watch: whether llama.cpp merges the architecture upstream, which would put Kolibri in Ollama and LM Studio for anyone with 64 to 128 GB of RAM or unified memory; independent German-language evaluations; and what Cohere says about the Kolibri line once the merger closes. I'm also checking the other European and Polish open-weight models (Mistral, Apertus, EuroLLM, Bielik, PLLuM) and will add them to the library only as I verify them, so expect a follow-up.

Get started

Frequently asked questions

What is Kolibri 1? Aleph Alpha's mixture-of-experts reasoning model for German and English, released on 3 October 2026. It has 78.1 billion parameters in total, about 3.46 billion active per token, adjustable reasoning effort and tool calling, and Apache-2.0 weights. Is Kolibri 1 open source? It is open-weight. The Apache-2.0 licence covers the published weights and configuration files only; Aleph Alpha explicitly excludes code, architecture, parameter settings and training methods. Can I run Kolibri 1 on Ollama or LM Studio? Not today. The architecture isn't in upstream llama.cpp, which Ollama and LM Studio build on. The official route is vLLM with Aleph Alpha's plugin; an experimental community GGUF needs a patched llama.cpp, and there are unofficial MLX builds for Macs with about 64 GB. How much memory does Kolibri 1 need? The FP8 weights are about 79 GB and the BF16 weights about 156 GB, and the KV cache comes on top. Aleph Alpha's minimum for FP8 is two 80 GB A100s or H100s, or a single H200, B200 or B300. Sparse compute doesn't reduce memory, because every expert must be loaded. Is Kolibri 1 better than Qwen, Gemma or GPT-OSS? On Aleph Alpha's own table it has the best overall English and German averages among the sparse models it compared, but it trails Qwen3.5 35B-A3B on agentic coding and a dense Qwen3.8 27B overall. These are vendor-run results, not independent rankings. Is Aleph Alpha still independent? Cohere and Aleph Alpha signed a definitive combination agreement on 16 September 2026, subject to regulatory approval and expected to close later in 2026. The combined company will operate as Cohere. Does Kolibri 1 comply with the EU AI Act? I can't answer that and neither can a model card. Aleph Alpha's card says it is a signatory of the EU General-Purpose AI Code of Practice and publishes a copyright-rightsholder contact, training-data categories and an energy estimate. Whether that satisfies your obligations is a question for your legal advisor.

Sources

Primary: Kolibri-1 (FP8) model card and Kolibri-1-BF16 model card plus Hugging Face API metadata (read 6 October 2026) · aleph-alpha-inference on PyPI · Hob-forge/Kolibri-1-GGUF (independent) · Pharia-1-LLM-7B-control LICENSE (Open Aleph License 1.0) · European Commission: General-Purpose AI Code of Practice · Arcee: "Trinity is moving to OpenMDW-1.1" (29 May 2026) · Aleph Alpha launch post (3 October 2026) · Aleph Alpha technical report.

Company news (press coverage of the Cohere and Aleph Alpha agreements): tech.eu, 24 April 2026 · Yahoo Finance · Trending Topics · AI Weekly · Unite.AI · European Open Source AI Index on Aleph Alpha's 2024 pivot.

Benchmark numbers are vendor-reported. The training figures (768 B200 GPUs, 20 trillion tokens, the 512-token window, 10 full-attention layers, tokenizer compression) come from Aleph Alpha's launch post and technical report, which I read on 6 October 2026; the 21-day pre-training time and the 950 MWh estimate come from the model card. Where a date or figure in the company-news paragraph comes from Reuters reporting as carried by these outlets, the outlet is linked above.