Aleph Alpha Releases Kolibri: A 78B Sovereign Open-Weight Model for Regulated Industries
Key point
Aleph Alpha has released Kolibri, a 78B-parameter Mixture-of-Experts model with 3.46B active parameters, designed for sovereign, regulated industries with strong German language support and agentic capabilities.
Details
Aleph Alpha has released Kolibri, a sovereign open-weight model licensed under Apache 2.0 and available on Hugging Face. Designed for regulated sectors such as public administration, industry, and aerospace, Kolibri emphasizes German language proficiency, reasoning, and agent capabilities. The model features a Mixture-of-Experts (MoE) architecture with 78.1B total parameters and 3.46B active parameters, supporting a context window of up to 1M tokens.
Performance and Efficiency
Kolibri positions itself on the Pareto frontier of cost-efficiency, delivering quality comparable to models with significantly larger active parameter counts. In benchmark comparisons against Nemotron 3 Super (120B-A12B), Kolibri demonstrated superior performance in several key areas:
- Mathematics: Kolibri scored 96.9 on AIME 2025 (vs. 91.7 for Nemotron 3 Super) and 88.8 on Math Average (DE).
- Agentic Tasks: It achieved 94.7 on τ²-bench Telecom (vs. 68.1) and 38.1 on τ³-bench Banking (vs. 15.5).
- Coding: Kolibri scored 85.9 on LiveCodeBench v6 (vs. 82.0).
- Hallucination Control: The model is trained to abstain from answering when context does not support a response, using the Merlin-Arthur protocol. On the AA-Omniscience Index, Kolibri scored -32.8.
Sovereignty and Architecture
The model is developed by a German team using infrastructure in Germany and Finland, ensuring compliance with the EU AI Act, GDPR, and the General-Purpose AI Code of Practice. Key architectural innovations include:
- Attention Mechanism: A hybrid pattern using sliding windows (512 tokens) for 40 layers and full attention for 10 layers to optimize memory and compute.
- Tokenizer: A new UniBPE tokenizer that improves German compression rates by respecting morphological structures.
- Training Data: Trained on 20T tokens, with 21.3% being German data curated through specialized pipelines to avoid translation artifacts.
- Grounding: Uses the Merlin-Arthur protocol to train the model to abstain from answering when context does not support a response.
Deployment
Kolibri is served via vLLM. It supports four reasoning effort levels (none, low, medium, high) to allow users to trade off between speed, cost, and quality. For long-context tasks exceeding 262,144 tokens, specific vLLM flags are required to enable the full 1M token context window.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.