
Aleph Alpha released its Kolibri language model on October 3, 2026, with full weights on Hugging Face under the permissive Apache 2.0 license. The Heidelberg-based company trained the model from scratch on hardware in Germany and Finland, treating German as an equal language alongside English. With it, a German company is once again offering an open model that keeps up with current rivals in its class and that government agencies and companies can run on their own hardware.
Key takeaways
- Kolibri has 78.1 billion parameters, but only 3.46 billion are active for each token. This design keeps compute costs low but requires a lot of memory.
- A little over a fifth of its 20 trillion training tokens is German text, and a custom tokenizer splits German compound words more efficiently. According to Aleph Alpha, the context window reaches just over one million tokens.
- In Aleph Alpha’s own comparisons, Kolibri edges out similarly efficient open models such as Qwen3.5 35B-A3B but trails the denser Qwen3.8 27B by a clear margin. Independent measurements are not yet available.
- Running it takes about 78 gigabytes of GPU memory, officially at least two A100 or H100 cards or one H200. There is no official version for Ollama or LM Studio yet.
- Kolibri arrives as Aleph Alpha merges with Canadian AI company Cohere; regulators still have to sign off.
What is inside the model
Kolibri is a so-called mixture-of-experts model. Instead of running every word through the entire network, a router in each of its 50 layers picks six of 384 specialized subnetworks, plus one shared expert. That way, the knowledge of a 78-billion-parameter model sits in memory while the compute load is closer to that of a model with three to four billion parameters. Aleph Alpha deliberately chose the size based on operating costs: a 123-billion-parameter variant it tested could handle only three very long-context requests at once on two H100 cards, while Kolibri handles 18 and generates text 28 percent faster.
Attention, the mechanism a model uses to take earlier parts of a text into account, is also tuned for efficiency. Only 10 of the 50 layers look across the full context; the other 40 see only the last 512 tokens. The model was trained with a context of 262,144 tokens, but Aleph Alpha has validated quality up to 1,048,576 tokens, roughly 700,000 English words. For latency-sensitive applications, the company itself recommends staying below the smaller limit. Users can set how much reasoning effort the model puts into an answer at four levels, from “none” to “high.”
German as a training language, not a translation
The most interesting part of the release concerns the data. Based on experiments with small models, Aleph Alpha concluded that roughly 20 percent German text yields the best results. With 20 trillion training tokens, that meant about four trillion German tokens. After cleanup, open German datasets supplied only 390 billion. The team closed the gap in two ways: with its own processing of the Common Crawl web archive, which produced 1.3 trillion tokens of original German text, and with about one trillion tokens in which a language model rewrote German documents in new forms, such as encyclopedia entries or question-and-answer dialogues.
One detail from the blog post shows why this was necessary. Common filters discard texts with too many long words as low quality. German administrative prose routinely exceeds the threshold set for English, so the default setting would have removed exactly the kind of language government agencies write in. The team used translations from English only sparingly, because a translated corpus imports the world of the English-language web: president instead of chancellor. On top of that comes a custom tokenizer that breaks German compound words into pieces more effectively and, according to the model card, processes 4.7 bytes of German text per token. The fewer tokens a text needs, the cheaper and faster each request becomes.
How Kolibri compares
Aleph Alpha mainly compares Kolibri with models that activate a similarly small number of parameters per token. In that class, it leads on the overall average: 75.5 points in English and 70.8 in German, versus 74.7 and 69.8 for Qwen3.5 35B-A3B and 73.0 and 67.9 for the much larger Nemotron 3 Super 120B-A12B. On math tasks such as AIME 2025, Kolibri scores 96.9 percent in English. To its credit, the model card also includes the dense Qwen3.8 27B, which uses about eight times as many parameters per token and leads clearly with 80.2 and 79.9 points, especially in German. All figures come from Aleph Alpha’s own test setup; independent leaderboards do not list the model yet.
One focus stands out that benchmarks rarely reward: Kolibri was trained to say “I don’t know” when an answer is not in the document at hand. On a test that measures exactly this kind of abstention, it scores 85.6 percent. For agencies and companies that want to point the model at their own files, that matters more than an extra percentage point on a math test. Aleph Alpha’s internal industry tests, for example for the German public sector, show clear gains over Kolibri Origin, a predecessor that was never released, but they cannot be verified from the outside.
Who can use Kolibri
The Apache 2.0 license allows commercial use with no additional conditions. The hurdle is hardware. In its standard version with 8-bit weights, Kolibri takes up about 78 gigabytes; Aleph Alpha lists the official minimum as two A100 or H100 cards with 80 gigabytes each, or a single H200 or B200. It runs on the vLLM serving software with a dedicated add-on package that exposes an API compatible with OpenAI’s. Existing applications can often be switched over just by changing an address. The model supports tool calling, which also makes it suitable for agents that trigger searches or programs.
There is movement for hobbyists with powerful machines, too. Within a day, unofficial user-converted versions appeared on Hugging Face, including heavily compressed variants for Apple’s MLX and in the GGUF format for llama.cpp; a 4-bit GGUF version weighs in at 47.5 gigabytes. They do not come from Aleph Alpha. According to their description, the GGUF files require a self-patched build of llama.cpp, so Ollama and LM Studio cannot load them yet. A user upload in full 16-bit precision has recently appeared in Ollama’s library, but at 173 gigabytes it is beyond ordinary computers. Anyone who wants to try Kolibri locally should wait for official support in those apps. As Aleph Alpha argued in its study on the political slant of Chinese models, published yesterday, a model’s origin increasingly matters for open models. Here, Kolibri has a real argument.
Our take: sovereignty with footnotes
Kolibri is not a frontier model, and Aleph Alpha does not claim it is. It is a tool for a clearly defined job: processing German and English documents in government and industry without sending data to outside cloud services. For that purpose, the combination of an open license, traceable data provenance and moderate operating costs makes sense. Even Kolibri does not do without outside components, though: the model card lists Google’s Gemma 4, Mistral NeMo and Qwen3 as helpers in generating and rating synthetic training data. And who will own the model going forward depends on the merger. Aleph Alpha and Cohere signed their agreement on September 16; the combined company is to be called Cohere and be headquartered in Berlin and Toronto, with Heidelberg remaining a research site. Kolibri’s weights, however, are already openly available under Apache 2.0, and that cannot be undone. For users in Germany, that is the real value of this release.
Sources
- Aleph Alpha: Kolibri Has Landed – A Sovereign Open-Weight Model
- Hugging Face: Modellkarte Aleph-Alpha/Kolibri-1
- The Decoder: Aleph Alpha veröffentlicht deutsch-englisches Sprachmodell Kolibri
- SiliconANGLE: Cohere and Aleph Alpha agree to merge
- Hugging Face: Inoffizielle GGUF-Umwandlung von Kolibri-1
- Startup Fortune: Aleph Alpha launches Kolibri

