Aleph Alpha Releases Kolibri: 78B Open-Weight German-English MoE, How to Choose Sovereign AI?
German AI company Aleph Alpha released Kolibri on October 3 (Day of German Reunification): a 78.1B total / 3.46B active German-English MoE with 1M context, Merlin-Arthur anti-hallucination protocol, exact quantile balancing routing, Apache 2.0 license for commercial use, trained on 768 B200 GPUs for 21 days over 20T tokens.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
TL;DR
On October 3, 2026 — German Reunification Day — Aleph Alpha released Kolibri: a 78.1B total / 3.46B active German-English MoE under Apache 2.0, with 1M token context, a novel Merlin-Arthur hallucination reduction protocol, and exact quantile balancing for expert routing. Trained on 768 B200 GPUs for 21 days over 20T tokens, it punches well above its 3B active parameter weight class.
Architecture & Parameters
Kolibri's core specifications:
| Metric | Value |
|---|---|
| Total parameters | 78.1B |
| Active parameters per token | 3.46B |
| Experts | 384 total / 6 active + 1 shared |
| Layers | 50 (all MoE) |
| Attention heads | 48 query / 4 KV |
| Model dimension | 2,560 |
| Tokenizer vocabulary | 128,000 |
| Max context | 1,000,000 (1M tokens) |
| Pre-training tokens | 20T (from 200T+ raw) |
| Knowledge cutoff | EN: Sep 2024, DE: Aug 2025, Combined: Jun 2026 |
Three key architectural design choices stand out:
384 small experts over fewer wide ones. Aleph Alpha's ablation testing showed that more, narrower experts outperformed fewer, wider ones at the same total parameter budget. This diverges from Qwen3.8's 128-expert approach, pushing sparsity further.
Exact quantile balancing. An improvement over Kimi K3's estimated approach — computed exactly at fixed cost independent of batch size, improving both load balance and model quality.
Sliding window + periodic full attention. 40 layers use 512-token sliding window attention; 10 layers (every 5th) use full attention. This hybrid balances efficiency with information aggregation for long contexts.
Training Details
Hardware and timeline:
- GPUs: 768 NVIDIA B200
- Pre-training duration: 21 days
- Stage 1: 20T tokens at 16K sequence length (21 days)
- Stage 2: 3.44T tokens mid-training at 64K
- Stage 3: 200B tokens long-context adaptation at 256K
- Total training: ~24T tokens
- Interruptions: 38 unplanned over 21 days (~1 per 10,000 GPU-hours), all auto-recovered
Post-training:
- SFT: 174B tokens of synthetic data generated, combined with open-source datasets into 268B tokens of training mix
- RL: 1.2 million curated tasks across code/math reasoning, agentic tasks, instruction following, Q&A, tool calling
- Method: Asynchronous training (generation parallel with training)
Pre-training data composition: English ~62%, German ~21.3% (~4.3T tokens), Code ~14%. Only 6% of German data is translated — the rest is organic or rephrased native German.
Merlin-Arthur Hallucination Reduction Protocol
Kolibri's most innovative feature is the Merlin-Arthur three-player training game for hallucination reduction:
- Arthur: The model being trained and shipped
- Merlin: Generates contexts that increase Arthur's correct-answer probability
- Morgana: Strips evidence from contexts to lure Arthur into hallucination
The training objective: Arthur must learn to answer Merlin's well-supported contexts while abstaining on Morgana's redacted contexts — any guess, even a lucky one, is scored as incorrect. This forces genuine reading comprehension rather than pattern-matched guessing.
Benchmarks validate the design. On the AA-Omniscience Non-Hallucination metric, Kolibri scores 44.0 — far ahead of Kolibri Origin's 14.8 and Nemotron 3 Super's 13.9. On the M/A Grounding Score (composite hallucination control), Kolibri scores 0.23, the highest among all compared models.
Benchmark Performance
Comparison with peer models (selected key benchmarks):
| Benchmark | Kolibri | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B | Mistral Small 4 119B-A6B |
|---|---|---|---|---|
| AIME 2025 | 96.9 | 84.6 | 91.7 | 79.8 |
| AIME 2025 (DE) | 87.5 | 82.9 | 85.6 | 72.3 |
| GPQA Diamond | 84.3 | 83.4 | 78.0 | 74.7 |
| LiveCodeBench v6 | 85.9 | 82.5 | 82.0 | 71.2 |
| HumanEval+ | 92.7 | 92.8 | 94.7 | 92.8 |
| BFCL v4 | 61.4 | 67.2 | 61.0 | 58.0 |
| τ²-bench telecom | 94.7 | 99.1 | 68.1 | 41.5 |
| τ³-bench banking | 38.1 | 10.6 | 15.5 | 5.7 |
Key finding: Kolibri (3B active) is competitive with or close to Nemotron 3 Super (12B active, 120B total) — a model with ~4x the active parameters. However, Nemotron 3 Super still leads on Overall (EN) 75.5 vs 80.2 and Overall (DE) 70.8 vs 79.9.
On industry-specific internal benchmarks, Kolibri excels: Automotive supplier 99.0, German public sector 75.0, Semiconductors 80.4, Aerospace 58.9.
Bilingual Tokenizer
Kolibri uses a custom UniBPE tokenizer combining BPE's bottom-up approach with Unigram's training objective. German text compression: 4.90 bytes/token — slightly behind DeepSeek V4's 3.72 and Gemini's 4.13, but superior in morphological splitting.
For example, "Bundessozialgerichtes" (Federal Social Court's) is split as "Bundes|sozial|gericht|es" by Kolibri, versus competitors' typical "Bund|ess|oz|ial|gericht|es" — the latter being semantically incorrect.
Four-Level Controllable Reasoning
Kolibri supports four reasoning effort levels: none / low / medium / high. Users can trade cost and latency against answer quality:
- none: No reasoning chain, for simple Q&A
- low: Lightweight reasoning, for standard tasks
- medium: Moderate reasoning, for complex analysis
- high: Deep reasoning, for math proofs and complex code
This allows deployment teams to flexibly control cost and latency per task on a single model.
Deployment
Kolibri can be served via vLLM:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choiceExtended to 1M context:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice \
--max-model-len 1048576 \
--hf-overrides '{"max_position_embeddings": 1048576}'Recommended sampling parameters: temperature 1.0, top_p 0.97, top_k 128.
Efficiency data: The 78B model handles 18 concurrent 256K long-context queries on two H100s (vs. only 3 for a 123B model), decoding 28% faster.
Sovereign AI Significance
Kolibri's "sovereignty" operates on two levels:
Construction: Trained on infrastructure in Germany and Finland under European and German law with no foreign control. The full pipeline — data curation, pre-training, post-training, optimization — is owned by Aleph Alpha. Compliance with EU AI Act, General-Purpose AI Code of Practice, GDPR, and copyright law was built in from the design phase.
Delivery: The model is small enough (78B) for on-premise deployment without sending data to third-party inference services. IP safety is guaranteed. Compliance comes as an inherited property of the model.
What This Means for Developers
For teams deploying AI in the European market, Kolibri offers an option that doesn't depend on US or Chinese providers. The Apache 2.0 license permits commercial use, and weights are downloadable from Hugging Face.
However, Kolibri's English overall capability (Overall EN 75.5) still trails Nemotron 3 Super (80.2) — it's not the best choice for English-only scenarios. Its core value lies in the combination of German-English bilingual ability, hallucination abstention, and European compliance — a combination that is unique in the market.
Model weights: https://huggingface.co/Aleph-Alpha/Kolibri-1
Tech report: https://aleph-alpha.com/downloads/tech-report.pdf
Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
n8n + OpenAI affiliate site
Automate content and affiliate monetization
Claude + n8n automation agency
Charge monthly for agent workflow builds
Related tutorials
Related news
- llama.cpp Adds Native Decision Model Support: 5 Open GGUFs, How to Cut Agent Routing Cost?
- Xiaomi MiMo-V2.6-Pro Hits Agent Arena: #5 Open-Source, #2 in Confirmed Success — How to Choose Agent Models?
- LangChain Open-Sources Model Router: 64% Cost Cut in Agent Coding with No Quality Loss — How to Build Yours
- Karpathy Proposes 4-Rung LLM Output Ladder: STE100, Diagrams, HTML, Video — How Humans Understand Autonomous Agent Work?