Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built specifically for German and English. Kolibri carries 78.1B total parameters but activates only 3.46B—roughly 4.4%—per token. It supports a context window of up to 1,048,576 tokens, allows users to configure reasoning effort on a per-request basis, and ships under the Apache 2.0 license on Hugging Face. The model targets sovereign deployment in regulated sectors such as public administration, industry, and aerospace.
Is It Deployable?
Yes. The FP8 checkpoint is approximately 78GB and runs on a single B200, B300, or H200 GPU, or on two H100 SXM5 GPUs. It is served through vLLM with dedicated Kolibri reasoning and tool-call parsers.
What Is Kolibri?
Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Europe—positioning it as a sovereign alternative for organizations that require data residency and regulatory compliance without relying on US- or China-based model providers. A sparse MoE architecture keeps inference costs low despite the large total parameter count, making a 78.1B-parameter model practical on a single high-end GPU.
Key Specifications
- Architecture: Mixture-of-Experts transformer (sparse activation)
- Total parameters: 78.1B
- Active parameters per token: 3.46B (~4.4%)
- Languages: English and German
- Context length: up to 1,048,576 tokens
- Configurable reasoning effort: per-request
- License: Apache 2.0
- Checkpoint format: FP8 (~78GB)
- Hardware requirements: 1× B200 / B300 / H200, or 2× H100 SXM5
- Serving stack: vLLM, with dedicated reasoning and tool-call parsers
Why It Matters in 2026
As of 2026, the AI industry has shifted decisively toward sparse MoE architectures—following the trend set by Mixtral, DeepSeek, and Qwen—because they deliver large-model quality at small-model inference cost. Kolibri pushes this efficiency further: its ~4.4% activation ratio is among the most aggressive in the open-weight space, while its million-token context window places it in the same tier as frontier proprietary models.
Equally significant is the sovereignty angle. With EU AI Act enforcement now in full effect, European public-sector and defense-adjacent organizations face strict requirements around data residency, auditability, and vendor independence. An Apache 2.0 model that can be self-hosted on a single GPU, trained end to end in Europe, and deployed without US cloud dependency directly addresses that demand.
Availability
Kolibri-1 is available now on Hugging Face under the Apache 2.0 license. The vLLM integration—including reasoning and tool-call parsers—makes it straightforward to deploy in existing inference pipelines.
via MarkTechPost
