Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
MarkTechPost Asif Razzaq ● Covered by 2 sources
Aleph Alpha dropped Kolibri, an open-weight English-German AI model with 78.1B total parameters. It only uses 3.46B per token and can stretch to 1,048,576 tokens of context.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Aleph Alpha has put out Kolibri, a bilingual English-German mixture-of-experts model that is open-weight, Apache 2.0 licensed, and aimed at deployments where control matters more than splashy demos. Public administration, industry, and aerospace are the intended home turf. Not exactly the kind of product that wants to live on a hype reel.
The model is big on paper and selective in practice. Kolibri has 78.1B total parameters, but only 3.46B are active for each token. That works out to 4.4% of the model being used at a time. Aleph Alpha also says the FP8 checkpoint is about 78GB and can run on a single B200, B300, or H200, or on two H100 SXM5 GPUs through vLLM. It ships with dedicated reasoning and tool-call parsers, and users can set reasoning effort per request.
The architecture is built around sparse routing and a hybrid attention setup. Kolibri uses 50 transformer blocks with a width of 2,560. Every MoE layer scores 384 routed experts, sends each token to the top 6, and always includes one shared expert. The attention stack mixes grouped-query attention with full attention every fifth block, while the other layers use sliding-window attention over the previous 512 tokens. That lets only 10 layers grow with context length, which is how Aleph Alpha says the design can support sequences four times longer than a full-attention model at matched compute.
Training and post-training were handled end to end by Aleph Alpha teams in Germany, with training infrastructure in Germany and Finland. The pre-training run covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens at 65,536 sequence length and then a 201B-token long-context stage on 262,144-token sequences. The company added more than 2T German tokens from the web or synthetic generation, then followed with supervised fine-tuning and reinforcement learning on more than 1.2M internal tasks. The Merlin-Arthur protocol is meant to make the model abstain when retrieved context doesn’t support an answer.
On the benchmark sheet, Aleph Alpha says Kolibri leads its English tests for GPQA Diamond, AIME 2025, and AIME 2026, and tops its German and English overall scores among the 12 MoE models it compared. It still trails some competitors in specific areas, including BFCL v4. But the headline is clear: this is a large model that’s trying hard to be practical, not just impressive.
My take — AI-written commentary, not fact-checked reporting
Kolibri looks like the sort of model that actually matches the European AI policy mood: controlled, licensed, and built for deployment instead of demo theater. That’s refreshing, because the industry has spent enough time pretending “open” means “a bigger scoreboard and a louder launch thread.”
Read more about this at: MarkTechPost