Aleph Alpha Releases Kolibri for Self-Hosted German and English AI
The downloadable model targets regulated work, with strong company-reported math results. Its low active parameter count does not mean small hardware requirements.
Loading page…
The downloadable model targets regulated work, with strong company-reported math results. Its low active parameter count does not mean small hardware requirements.
Listen to this story
Kolibri 1 gives organizations a German- and English-focused model they can run on their own infrastructure, but its sparse computation does not make deployment lightweight: all 78.1 billion weights must remain in memory. Aleph Alpha reports strong results on a German math benchmark, while Qwen leads on tool use and long-context tests; these comparisons are company-run, not independent validation. The release offers a sovereignty-oriented option with open weights, but its training code and methods remain proprietary.
Kolibri activates 3.46 billion parameters per token, but its weights alone require about 78 GB in eight-bit format.
The weights and configuration files use Apache 2.0; Aleph Alpha retains rights to its training code and methods.
Aleph Alpha reports 87.5 on German AIME 2025 versus Qwen3.6-35B-A3B’s 82.9, while Qwen leads BFCL v4 and LongBench Pro.
Government agencies and industrial companies have a new downloadable option for German and English AI work. Aleph Alpha released Kolibri 1 on October 3, 2026, with weights available on Hugging Face under Apache 2.0. It is designed to run on customer-controlled infrastructure, without sending internal documents to an outside inference service.
The release targets public administration, industry and aerospace rather than treating every language and workload equally. Aleph Alpha says it built the model in Germany and trained it in Germany and Finland. Its sovereignty pitch combines control over development with customers’ freedom to choose where the model runs.
Open weights do not mean every part of development is open. The Apache 2.0 license covers weights and configuration files; Aleph Alpha retains rights to its training code and methods.
Kolibri contains 78.1 billion parameters—the learned values that shape its responses—but activates only 3.46 billion for each token, or chunk of text. Its mixture-of-experts design routes each token through six of 384 specialist components, plus a shared expert. That limits the portion doing work at any moment.
The storage requirement is different. All the weights must remain in memory, even when most are inactive. AI engineer Tejas Kumar’s technical walkthrough puts the weights alone at about 78 GB in eight-bit floating-point format. Running the model also requires Aleph Alpha’s plugin for vLLM, the software used to serve responses.
The full model contains 78.1 billion parameters.
Only 3.46 billion parameters are active for each token.
German accounts for 21.3% of Kolibri’s pretraining tokens. Aleph Alpha reports 20 trillion pretraining tokens overall and says translation supplied only 6% of the data. A bilingual tokenizer—the system that divides text into chunks—also supports its emphasis on German rather than treating the language as an afterthought.
Kumar tested that tokenizer on Germany’s Basic Law. Kolibri represented the German text in 35,190 tokens, versus 41,482 for the tokenizer used by GPT-5. The English translation produced nearly equal counts. This was a tokenizer experiment, not a model-quality test: fewer chunks let more of the same German text fit into a fixed context window.
In Aleph Alpha’s launch benchmarks, Kolibri scored 87.5 on the German version of the AIME 2025 math test, versus 82.9 for Qwen3.6-35B-A3B. But Qwen led on the overall BFCL v4 tool-calling test, 67.2 to 61.4, and LongBench Pro, 70.8 to 64.5. These are company-run comparisons, not independent validation.
Length also needs qualification. Kolibri’s longest trained context was 262,144 tokens; Aleph Alpha validated it up to 1,048,576, according to Kumar. The larger figure describes tested reach, not the length used throughout training.
Its attention design reduces the work needed for long inputs. Forty of its 50 layers look only at the preceding 512 tokens. Every fifth layer looks across the full preceding text, combining local processing with periodic access to the wider document.
Aleph Alpha trained Kolibri to decline answers when supplied documents lack the evidence. Its Merlin-Arthur method alternates examples where supporting evidence remains visible with examples where that evidence is hidden. The model learns either to answer from the document or acknowledge that it cannot.
Users can also choose none, low, medium or high reasoning effort for each request. Aleph Alpha presents those settings as a way to trade response time and cost against answer quality on the same model.
Loading discussion...
Join the conversation
Explain when outside knowledge would help rather than undermine trust.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.