Article Summary (Model: gpt-5.6-sol)
Subject: Sovereign Bilingual MoE
The Gist:
Aleph Alpha released Kolibri, an Apache 2.0 open-weight English-German mixture-of-experts model aimed at regulated, on-premise enterprise and government work. It has 78B total parameters but activates roughly 3B per token, supports contexts up to 1M tokens, and offers four reasoning-effort levels. The company emphasizes European control of the training stack, efficient local deployment, German-language specialization, agentic capabilities, and document-grounded abstention.
Key Claims/Facts:
- High-velocity pipeline: Kolibri was trained on 768 B200 GPUs using nearly 24T tokens; an automated, versioned pipeline handled evaluation and 38 training interruptions without manual intervention.
- German by design: German made up 21.3% of pre-training tokens, drawing heavily on curated German web text and synthetic rephrasing; the custom UniBPE tokenizer targets German morphology and compression.
- Grounding and efficiency: Merlin-Arthur training teaches the model to abstain when evidence is absent, while sparse experts and mostly sliding-window attention reduce serving costs.
Discussion Summary (Model: gpt-5.6-sol)
Consensus: Cautiously Optimistic—the thread strongly praises the unusually detailed technical report and open weights, while questioning benchmark competitiveness, hallucination claims, and the meaning of “sovereign.”
Top Critiques & Pushback:
Better Alternatives / Prior Art:
Expert Context: