Aleph Alpha's Kolibri: A Sovereign Open-Weight Model for Regulated Industries
Aleph Alpha's Kolibri: A Sovereign Open-Weight Model for Regulated Industries
On October 3rd, 2026, Aleph Alpha released Kolibri, an open-weight bilingual English-German Mixture-of-Experts (MoE) Transformer. The model, with 78.1 billion total parameters and 3.46 billion active per token, supports a context length of up to 1 million tokens. It's available for download with full weights on Hugging Face under the permissive Apache 2.0 license. This release is strategically timed to Germany's Unity Day, underscoring its sovereign AI ambitions.
Kolibri is engineered for mission-critical work in regulated sectors like public administration, industrials, and aerospace. Its design prioritizes sovereignty, grounding, and efficiency. Aleph Alpha emphasizes full supply-chain integrity—from data curation and training on German and Finnish infrastructure to final evaluations—all under European law. This approach aims to provide customers with freedom of deployment and IP safety, making compliance an inherited property of the model.
Technical Deep Dive: Architecture and Training
Kolibri's architecture is a study in efficiency. It uses 384 experts with only 6 active per token. To manage inference costs, 40 of its 50 layers employ a 512-token sliding window attention, while the remaining 10 layers process full context. This allows the model to handle long contexts without exploding compute costs. The model was trained on 768 B200 GPUs in three stages: 20 trillion tokens of pre-training, 3.44 trillion tokens of mid-training, and 200 billion tokens of long-context adaptation, totaling nearly 24 trillion tokens.
The training pipeline, dubbed the 'Model Factory,' is a key differentiator. It's implemented as code, enabling hundreds of ablation experiments and automated recovery from hardware failures. Between Kolibri Origin (a 30B model released internally) and Kolibri, the team scaled from 7.5T to 20T training tokens and from a 65k to a 1M token context window in just three months. This velocity is a testament to the maturity of their infrastructure. The model also introduces exact quantile balancing for expert routing, improving load balance and model quality.
Bilingual by Design: German and English Data Strategy
Aleph Alpha made a significant effort to make Kolibri genuinely bilingual. German constitutes 21.3% of pre-training tokens, sourced from a 2.4T-token unique pool. This pool was built from curated Common Crawl data (retuned for German's linguistic quirks), rephrased German documents, and only a small amount of translation. They argue that translated text carries the cultural fingerprint of its source language, so they prioritized organic German data. This strategy yields a model that reasons directly in German, avoiding the 'translationese' problem.
The company also developed a new tokenizer training method called UniBPE, which combines BPE with a Unigram objective. This approach better respects German morphology, leading to superior compression rates on German text compared to other state-of-the-art models. This translates to more efficient inference and lower costs for German-language tasks.
Grounding and Abstention: Reducing Hallucinations
A core focus for Kolibri is reducing hallucinations. The model is trained to abstain from answering when the context doesn't support an answer. This is achieved through dedicated abstention data and Aleph Alpha's in-house Merlin-Arthur protocol. This adversarial training setup uses two auxiliary models: 'Merlin' generates contexts that encourage correct answers, while 'Morgana' strips out evidence to lure the model into hallucinating. The result is a model that knows when to say 'I don't know.'
The numbers support this focus. On the AA-Omniscience benchmark, Kolibri abstains instead of answering incorrectly on 44% of items, versus 14.8% for Kolibri Origin. Its M/A grounding score, which certifies how much of an answer provably came from the document, is 0.23, compared to 0.00 for Origin and many other models. This is a critical feature for regulated industries where a wrong answer is unacceptable.
Benchmark Performance and Market Position
Kolibri's performance is competitive, especially for its size. On average English benchmarks, it scores 75.5%, and on German, 70.8%. It matches or exceeds models with up to four times its active parameter count on math, coding, and long-context tasks. For instance, it scores 96.9% on AIME 2025 (EN), better than Nemotron 3 Super's 91.7%. On agentic benchmarks like τ²-bench telecom, it scores 94.7%, outperforming much larger models.
However, critics like Trending Topics point out that on some general knowledge benchmarks, Kolibri trails the very latest open-weight leaders. For example, on Humanity's Last Exam (EN), Kolibri scores 21.5% versus Qwen3.8's 35.6%. Aleph Alpha counters that public benchmarks don't capture the specialized needs of their customers. They've built internal evaluation suites for verticals like aerospace and public administration, where Kolibri shows dramatic improvements over its predecessor, jumping from 0.14 to 0.59 in aerospace. The company argues that contextualized performance and sovereignty are more important than leaderboard supremacy.
Why It Matters: Sovereignty and Compliance
Kolibri is not just another open-weight model; it's a strategic statement. It's built with the EU AI Act, GDPR, and copyright law in mind. The model is trained and built in Germany and Finland, with no foreign control. This is a direct response to European concerns about data sovereignty and dependence on US or Chinese AI providers. For EU public buyers, this offers a path to compliant AI deployment on-premise, without sending sensitive data to third-party inference services.
The model is open-weight, not fully open-source, as training code remains proprietary. This distinction matters for some, but for Aleph Alpha's target market, the full transparency of weights and data curation decisions is sufficient. The release also comes amid leadership changes and a strategic partnership with Cohere, raising questions about the company's future direction. But for now, Kolibri stands as a credible, sovereign alternative for organizations where control and compliance are non-negotiable.
Getting Started with Kolibri
You can download Kolibri from Hugging Face (Aleph-Alpha/Kolibri-1). It requires the aleph-alpha-inference package, which provides a Kolibri-specific vLLM plugin. Serving it with reasoning and tool-calling enabled is straightforward, and settings are provided to extend context beyond 262k tokens to the full 1M. Recommended sampling parameters are temperature=1.0, top_p=0.97, and top_k=128.
Related News

Extra Big Ass Intelligence: A Satirical Take on AI Hype Goes Viral

AI Painters Create Art in Simulated Oil Studio, No Images Needed

Cloudflare Clef: Open-Weight Decision Models & RL Fine-Tuning Platform

OpenAI & Synopsys Unveil GPT-Synopsys to Transform Chip Design

OpenDLSS Brings NVIDIA DLSS 5 to Vulkan with Bit-Exact Precision

