Aleph Alpha has published the full weights of a new language model called Kolibri on Hugging Face, under the Apache 2.0 license. The company announced the release in a post dated October 3, and TestingCatalog reported it. The pitch is aimed at governments and regulated industries that want to process their own documents on hardware they control, not hand them to an outside company’s servers.
Start with what “open-weight” buys a buyer like that. The weights are the trained model itself, a very large file of numbers. Once it is downloaded, a ministry, hospital, or bank can run it on its own machines, so internal documents never travel to someone else’s data center. With a closed model reached through an API, every prompt leaves the building. Apache 2.0 adds a second benefit: a downloaded copy keeps working whatever the lab decides to do next.
The model is a Mixture-of-Experts design, meaning it holds many specialist sub-networks and routes each token through only a few of them. Per TestingCatalog, Kolibri has 78.1 billion parameters in total, 384 experts, and uses six of them at a time, which works out to 3.46 billion parameters active per token. That keeps the cost of each answer low. It does not shrink the memory bill, because the whole 78 billion still has to sit on the hardware.
The headline context window is one million tokens. Read the detail, though. The lab’s long-context training stopped at 256,000 tokens, and Aleph Alpha supplies serving settings that stretch the model to 1,048,576. A buyer planning to feed it entire case archives should test that upper range before trusting it.
Benchmarks exist, and all of them are the vendor’s. Aleph Alpha reports 96.9 on AIME 2025, a math competition test, 85.9 on LiveCodeBench v6 for coding, and 61.4 on BFCL v4, which scores tool calling. It also gives an overall score of 75.5 in English and 70.8 in German. These were run with the lab’s own harnesses at the highest reasoning setting. The company claims Kolibri holds its own against models that activate as much as four times more parameters per token, across a spread of tasks from math and code to very long documents. The report cites no independent evaluation, and the figures come without scores for rival models to set beside them.
The grounding claim is the most interesting one for regulated work. Kolibri was trained to withhold an answer when the supplied documents do not support one. On the AA-Omniscience test it avoided giving a wrong answer on 44 percent of items, against 14.8 percent for its predecessor, Kolibri Origin. The comparison is against the lab’s own earlier model, not a competitor, so it shows progress without showing rank.
Then the sovereignty argument. Aleph Alpha says it controlled data curation, training, evaluation, the weights, and deployment, and that this is meant to satisfy European compliance and sovereignty requirements. The model was built and trained in Germany and Finland, and German makes up 21.3 percent of its pre-training tokens. Those are statements of intent. TestingCatalog cites no certification or audit, and “mission-critical” describes the target market rather than demonstrated reliability. The verifiable part is simpler: the weights are public, so anyone can inspect and host them.
A procurement team can settle in a week what the benchmark table cannot: load the weights on its own hardware and run its own German-language documents through it, including the questions the documents cannot answer.
Originally reported by TestingCatalog on October 3, 2026.