Aleph Alpha’s Sovereign A.I. Model Kolibri Is No Match for the Open-Weight Leaders
Trending Topics Jakob Steinschaden
Aleph Alpha released Kolibri, a German AI model built to run in customers’ own data centers. It looks decent on older tests, but newer open models have already moved on.
Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Aleph Alpha has put Kolibri out in the open, with full weights on Hugging Face and an Apache 2.0 license, and it picked Germany’s Unity Day to do it. The pitch is classic Aleph Alpha: a sovereign model for places that do not want their data sent to cloud systems in the U.S. or China, and that prefer to keep the whole stack inside their own walls.
Kolibri is a mixture-of-experts model with 78 billion parameters, but only about 3 billion are active at a time. It has a context window of up to one million tokens, was trained on 20 trillion tokens of pre-training and nearly 24 trillion tokens in total, and ran on 768 Nvidia B200 GPUs in Germany and Finland. German is a big part of the mix too: 21.3 percent of the training data is German, with very little machine translation, and Aleph Alpha says its tokenizer handles German more efficiently than GPT-5, Gemini or Qwen.
On paper, the model does well against the rivals Aleph Alpha chose to compare it with. It tops the AIME 2025 and 2026 math tests, GPQA Diamond and LiveCodeBench in those head-to-heads, and the company says it can adjust how much it “thinks” to balance cost and quality. It was also trained to answer “I don’t know” when the documents do not contain the answer, which is a refreshingly unheroic trait for an AI model.
But the timing hurts. The three models Aleph Alpha picked for comparison all date from the spring, while Chinese labs have since shipped stronger open-weight systems such as Qwen3.8, GLM-5.3, Kimi K3 and Xiaomi’s MiMo-V2.6-Pro. Kolibri is not measured against any of those, and on current Artificial Analysis data the likely result would be somewhere around 15 to 20 points, which would leave it behind at least 20 open-weight models. Even in its own weight class, others are now ahead.
That does not make Kolibri meaningless. It may still fit the customers Aleph Alpha cares about, especially in public administration, industry and aerospace, where sovereignty matters more than leaderboard bragging rights. But as a showcase for Europe catching up in open-weight AI, it arrives looking a bit like a very polished postcard from a race that has already sprinted on.
My take — AI-written commentary, not fact-checked reporting
This is what happens when a company sells sovereignty and then measures itself against yesterday’s models. The open-weight world is moving fast, and a respectable German model built for controlled deployment is still not the same thing as a leader. The uncomfortable truth is that “made in Europe” is not a benchmark.
Read more about this at: Trending Topics