Aleph Alpha has rolled out Kolibri, a German-English language model boasting 78 billion total parameters, according to the official technical documentation. The model runs on a mixture-of-experts architecture that fires up roughly three billion parameters per token, striking a calculated balance between operational overhead and output quality across both supported languages.
German constitutes 21.3 percent of the training corpus, backed by a dedicated data pipeline assembled specifically for this project. Curiously, the pipeline also leaned on synthetic training data generated by Chinese models, a detail that adds a sharp irony to the company's grand narrative of unadulterated European technological independence.
Aleph Alpha has positioned Kolibri squarely at public administrations, aviation, and heavy industry—sectors where regulatory compliance under the EU AI Act is non-negotiable. The model accommodates context windows stretching up to one million tokens and was trained across 768 B200 GPUs deployed physically in Germany and Finland. For all the talk of proprietary borders, the weights sit openly on Hugging Face under an Apache 2.0 license.
They promised an airtight, independent European AI stack forged strictly under local regulatory frameworks. Then they shipped a model trained in part on synthetic outputs cooked up by Chinese systems and slapped the label of digital sovereignty on it. That is one way to define independence.