AlephAlpha Releases Kolibri: Open-Source MoE Model for Sovereign English-German AI
Key point
AlephAlpha has released Kolibri, an open-weight Mixture-of-Experts model with 78.1B total parameters and 3.46B active parameters, designed for sovereign and regulated domains with a focus on German language performance.
Details
Model Architecture and Efficiency
AlephAlpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) Transformer licensed under Apache 2.0. The model targets sovereign and regulated domains, offering 78.1B total parameters with only 3.46B active parameters per token (4.4%). It utilizes a Decoder-only Causal Transformer with Sandwich Normalization and a Hybrid Attention structure: 40 blocks use Sliding-Window Attention (SWA) with a window size of 512, while 10 blocks use Full Attention. The MoE layer selects the Top-6 experts out of 384 routed experts using Sigmoid routing, combined with one shared expert.
Training Data and Long-Context Capabilities
Kolibri was pre-trained on 24T tokens, with German comprising more than 20% of the data mix. The training pipeline included pre-training, mid-training, and a Long-context extension phase to support up to 1M tokens context. The tokenizer, UniBPE with a 128,000 vocabulary, shows improved efficiency for German, achieving 4.90 bytes/token. In terms of morphological alignment, Kolibri creates splits at morpheme boundaries with a precision of 65% in German, compared to 49% for the best of the external tokenisers evaluated.
Performance Benchmarks
In evaluations, Kolibri claims a Pareto frontier position for quality and serving cost against open models:
- Base Model: Compared to Nemotron Nano/Super and Qwen3.5 Base, it achieves 1.6x throughput and scores +22.7 points in English and +20.4 points in German.
- Post-trained Model: Compared to Qwen3.6/3.8 and Gemma 4 IT, it achieves 2.7x throughput and scores +21.4 points in English and +24.4 points in German.
- Specific Scores: MMLU 81.0, GSM8K 89.8, HumanEval 86.7.
- Long Context: At 1M tokens, Kolibri scores 63.2 on RULER. The source notes it lies on the Pareto frontier among the open models evaluated in the report.
Post-Training and Compliance
The model underwent SFT and RL using 1.2M+ curated tasks focused on reasoning, tool use, and instruction following. Training was conducted on 768 NVIDIA B200 GPUs. To comply with EU AI Act and GDPR, the model includes specific filters for illegal content, PII masking, and abstention training for ungrounded answers. German language consistency was enforced during RL by translating environments and applying consistency rewards.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.