WayToClawEarn
Medium impactAleph Alpha 官方博客

Aleph Alpha Releases Kolibri: 78B Open-Weight German-English MoE, How to Choose Sovereign AI?

German AI company Aleph Alpha released Kolibri on October 3 (Day of German Reunification): a 78.1B total / 3.46B active German-English MoE with 1M context, Merlin-Arthur anti-hallucination protocol, exact quantile balancing routing, Apache 2.0 license for commercial use, trained on 768 B200 GPUs for 21 days over 20T tokens.

Edisen Lu · WayToClawEarnVia Aleph Alpha 官方博客Published Oct 4, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · Aleph Alpha 官方博客

TL;DR

On October 3, 2026 — German Reunification Day — Aleph Alpha released Kolibri: a 78.1B total / 3.46B active German-English MoE under Apache 2.0, with 1M token context, a novel Merlin-Arthur hallucination reduction protocol, and exact quantile balancing for expert routing. Trained on 768 B200 GPUs for 21 days over 20T tokens, it punches well above its 3B active parameter weight class.

Architecture & Parameters

Kolibri's core specifications:

MetricValue
Total parameters78.1B
Active parameters per token3.46B
Experts384 total / 6 active + 1 shared
Layers50 (all MoE)
Attention heads48 query / 4 KV
Model dimension2,560
Tokenizer vocabulary128,000
Max context1,000,000 (1M tokens)
Pre-training tokens20T (from 200T+ raw)
Knowledge cutoffEN: Sep 2024, DE: Aug 2025, Combined: Jun 2026

Three key architectural design choices stand out:

384 small experts over fewer wide ones. Aleph Alpha's ablation testing showed that more, narrower experts outperformed fewer, wider ones at the same total parameter budget. This diverges from Qwen3.8's 128-expert approach, pushing sparsity further.

Exact quantile balancing. An improvement over Kimi K3's estimated approach — computed exactly at fixed cost independent of batch size, improving both load balance and model quality.

Sliding window + periodic full attention. 40 layers use 512-token sliding window attention; 10 layers (every 5th) use full attention. This hybrid balances efficiency with information aggregation for long contexts.

Training Details

Hardware and timeline:

  • GPUs: 768 NVIDIA B200
  • Pre-training duration: 21 days
  • Stage 1: 20T tokens at 16K sequence length (21 days)
  • Stage 2: 3.44T tokens mid-training at 64K
  • Stage 3: 200B tokens long-context adaptation at 256K
  • Total training: ~24T tokens
  • Interruptions: 38 unplanned over 21 days (~1 per 10,000 GPU-hours), all auto-recovered

Post-training:

  • SFT: 174B tokens of synthetic data generated, combined with open-source datasets into 268B tokens of training mix
  • RL: 1.2 million curated tasks across code/math reasoning, agentic tasks, instruction following, Q&A, tool calling
  • Method: Asynchronous training (generation parallel with training)

Pre-training data composition: English ~62%, German ~21.3% (~4.3T tokens), Code ~14%. Only 6% of German data is translated — the rest is organic or rephrased native German.

Merlin-Arthur Hallucination Reduction Protocol

Kolibri's most innovative feature is the Merlin-Arthur three-player training game for hallucination reduction:

  • Arthur: The model being trained and shipped
  • Merlin: Generates contexts that increase Arthur's correct-answer probability
  • Morgana: Strips evidence from contexts to lure Arthur into hallucination

The training objective: Arthur must learn to answer Merlin's well-supported contexts while abstaining on Morgana's redacted contexts — any guess, even a lucky one, is scored as incorrect. This forces genuine reading comprehension rather than pattern-matched guessing.

Benchmarks validate the design. On the AA-Omniscience Non-Hallucination metric, Kolibri scores 44.0 — far ahead of Kolibri Origin's 14.8 and Nemotron 3 Super's 13.9. On the M/A Grounding Score (composite hallucination control), Kolibri scores 0.23, the highest among all compared models.

Benchmark Performance

Comparison with peer models (selected key benchmarks):

BenchmarkKolibriQwen3.6-35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6B
AIME 202596.984.691.779.8
AIME 2025 (DE)87.582.985.672.3
GPQA Diamond84.383.478.074.7
LiveCodeBench v685.982.582.071.2
HumanEval+92.792.894.792.8
BFCL v461.467.261.058.0
τ²-bench telecom94.799.168.141.5
τ³-bench banking38.110.615.55.7

Key finding: Kolibri (3B active) is competitive with or close to Nemotron 3 Super (12B active, 120B total) — a model with ~4x the active parameters. However, Nemotron 3 Super still leads on Overall (EN) 75.5 vs 80.2 and Overall (DE) 70.8 vs 79.9.

On industry-specific internal benchmarks, Kolibri excels: Automotive supplier 99.0, German public sector 75.0, Semiconductors 80.4, Aerospace 58.9.

Bilingual Tokenizer

Kolibri uses a custom UniBPE tokenizer combining BPE's bottom-up approach with Unigram's training objective. German text compression: 4.90 bytes/token — slightly behind DeepSeek V4's 3.72 and Gemini's 4.13, but superior in morphological splitting.

For example, "Bundessozialgerichtes" (Federal Social Court's) is split as "Bundes|sozial|gericht|es" by Kolibri, versus competitors' typical "Bund|ess|oz|ial|gericht|es" — the latter being semantically incorrect.

Four-Level Controllable Reasoning

Kolibri supports four reasoning effort levels: none / low / medium / high. Users can trade cost and latency against answer quality:

  • none: No reasoning chain, for simple Q&A
  • low: Lightweight reasoning, for standard tasks
  • medium: Moderate reasoning, for complex analysis
  • high: Deep reasoning, for math proofs and complex code

This allows deployment teams to flexibly control cost and latency per task on a single model.

Deployment

Kolibri can be served via vLLM:

terminal
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

Extended to 1M context:

terminal
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice \
  --max-model-len 1048576 \
  --hf-overrides '{"max_position_embeddings": 1048576}'

Recommended sampling parameters: temperature 1.0, top_p 0.97, top_k 128.

Efficiency data: The 78B model handles 18 concurrent 256K long-context queries on two H100s (vs. only 3 for a 123B model), decoding 28% faster.

Sovereign AI Significance

Kolibri's "sovereignty" operates on two levels:

Construction: Trained on infrastructure in Germany and Finland under European and German law with no foreign control. The full pipeline — data curation, pre-training, post-training, optimization — is owned by Aleph Alpha. Compliance with EU AI Act, General-Purpose AI Code of Practice, GDPR, and copyright law was built in from the design phase.

Delivery: The model is small enough (78B) for on-premise deployment without sending data to third-party inference services. IP safety is guaranteed. Compliance comes as an inherited property of the model.

What This Means for Developers

For teams deploying AI in the European market, Kolibri offers an option that doesn't depend on US or Chinese providers. The Apache 2.0 license permits commercial use, and weights are downloadable from Hugging Face.

However, Kolibri's English overall capability (Overall EN 75.5) still trails Nemotron 3 Super (80.2) — it's not the best choice for English-only scenarios. Its core value lies in the combination of German-English bilingual ability, hallucination abstention, and European compliance — a combination that is unique in the market.

Model weights: https://huggingface.co/Aleph-Alpha/Kolibri-1

Tech report: https://aleph-alpha.com/downloads/tech-report.pdf

Aleph AlphaKolibri开源模型MoE主权AI德英双语Apache-2.0Merlin-Arthur

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.