Summarize this article with:
Aleph Alpha's Kolibri is a 78.1B-parameter Mixture-of-Experts model built for German and English, released with open weights under Apache 2.0. This article covers its architecture, 1M-token context, published benchmarks, hardware requirements and what it means for sovereign AI in Europe, including how it fits into a multi-provider architecture with EdenAI.
As AI moves deeper into government, industry and other regulated environments, the question is no longer only how capable a model is. It is also where it runs, who controls it, what happens to the data, and how freely an organization can change its infrastructure or providers.
That is the problem Aleph Alpha is targeting with Kolibri (German for "hummingbird"), an English-German Mixture-of-Experts model released on October 3, 2026, the Day of German Unity.**
What makes Kolibri significant is the combination: open weights, European development, German and English specialization, efficient inference, long context, and the option to deploy it on an organization's own infrastructure and under its own governance model.
That makes Kolibri a useful case study in what European AI sovereignty looks like when it becomes an engineering problem rather than only a policy objective.
What Is Kolibri?
Kolibri is a reasoning model developed by Aleph Alpha with a specific focus on German and English. It was trained from scratch and contains 78.1 billion parameters in total, of which only about 3.46 billion (4.4%) are active for each token.
It is built for long-context and agentic workloads. Aleph Alpha lists multi-step reasoning, retrieval-augmented generation (RAG), agentic tool calling, coding and German- and English-language assistants among its primary use cases.
The model is published on Hugging Face as Kolibri-1, in an FP8 version and a BF16 version, under the Apache 2.0 license. It is the first language model Aleph Alpha has released under Apache 2.0. Its earlier open models used a non-commercial license.
"Open-weight" is the more precise description than "open source". The Apache 2.0 license covers the published weights and configuration files, which gives organizations broad freedom to deploy, adapt and commercialize the model. Aleph Alpha explicitly keeps the rights to its training code, training methods and other artifacts outside the repository.
Inside Kolibri: MoE, Reasoning and Long Context
Kolibri's architecture is built on a simple trade-off: a large model capacity does not mean the full model has to run for every token.
The model has 50 transformer layers. In each Mixture-of-Experts layer, a router scores 384 specialized "experts" and sends each token to the 6 most relevant ones, while 1 shared expert always runs. Only that small subset participates in the computation.
This matters because serving a model involves more than counting parameters. Sparse activation keeps the computation per token low, but the full set of weights still has to sit in memory. "3.46B active parameters" therefore does not mean Kolibri can be served like a conventional 3.46B-parameter model.
Kolibri also keeps long contexts affordable through a hybrid attention design. Four out of every five attention layers only look at the most recent 512 tokens, while the remaining layers attend to the full context.
Controllable reasoning and tool calling
Kolibri exposes adjustable reasoning effort: none, low, medium or high, set per request. Applications can trade response speed and compute against more deliberate reasoning only when a task needs it.
Tool calling is supported and can be combined with reasoning. That makes the model suitable for agentic workflows where it has to call APIs, run searches or execute code, rather than only generate text.
The 1M-token context comes with an important detail
Kolibri is advertised with a context length of up to 1,048,576 tokens, roughly one million. That number needs context.
The model was trained on sequences of up to 262,144 tokens, its native context length. Because of how positional information is handled, it can be extended beyond that without retraining, and Aleph Alpha has validated quality and serving efficiency up to 1M tokens. Serving more than 262K tokens requires an explicit configuration change, and Aleph Alpha recommends staying at 262,144 tokens or less for latency-sensitive deployments and complex tasks.
For production systems, the distinction matters. A million-token context is valuable for large document collections, long technical specifications and regulatory material. But latency, memory consumption, throughput and the cost of processing very large prompts still apply.
The real value of long context is not simply "more tokens". It is the ability to design applications around larger working sets of information when the workload actually benefits from them.
A Model Built Around German and English
Kolibri's language focus is one of its most distinctive characteristics.
Instead of adding German on top of a primarily English model, Aleph Alpha designed Kolibri around two languages by choice, favoring depth over breadth. German accounts for more than one-fifth of the pre-training data, about 4.3 trillion tokens, alongside roughly 62% English and 14% code. Aleph Alpha also emphasized organic German data rather than relying mainly on translations from English.
The tokenizer was built for the same purpose. It is designed around German word structure, including compound words, so German text is split into fewer tokens without making English less efficient. Fewer tokens per document means lower cost and more usable context for the same German text.
The focus continues into post-training: Kolibri was trained to reason in German when working in German, so users can follow its reasoning in their own language.
Language quality is not only about grammatically correct output. Tokenization, training data, terminology, domain knowledge and the language mix throughout training all affect how well a model handles real-world documents. For organizations that work with German administrative, industrial, legal or technical material, this is a central part of Kolibri's positioning.
How Does Kolibri Perform?
Aleph Alpha evaluates Kolibri across knowledge, math, coding, agentic tasks, grounding, long context and industry-specific RAG, in both English and German.
Headline results include 96.9 on AIME 2025, 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6 and 64.5 on LongBench Pro. Across all categories, Kolibri reaches an overall score of 75.5 in English and 70.8 in German, the highest among the Mixture-of-Experts models in Aleph Alpha's comparison.
It does not lead everywhere. On LongBench Pro, for example, Qwen3.6 35B-A3B scores higher (70.8), and several competitors do better on some knowledge and grounding benchmarks. These are also Aleph Alpha's own evaluations, run with its open-source evaluation framework, rather than an independent benchmark study.
The more interesting point is the relationship between quality and serving cost. Aleph Alpha positions Kolibri on a quality-throughput Pareto frontier: strong output quality for the amount of compute needed to serve it.
For production AI, that is often what decides. A model does not run in isolation from its infrastructure, and a slightly higher benchmark score can matter less than the combination of quality, throughput, hardware requirements, latency and operational control.
From Kolibri Origin to Kolibri
Kolibri was not a single jump from an empty training run to a 78B model.
Aleph Alpha first built Kolibri Origin, a smaller model with 30B total and 3B active parameters and a 65K-token context window. It served to validate and automate the training pipeline before the approach was scaled up.
That pipeline covers data ingestion and curation, experiments and ablations, pre-training, post-training and evaluation. Aleph Alpha highlights automation that kept training running through hardware failures or interrupted data connections without manual intervention.
The full model was then trained in three stages: about 20 trillion tokens of pre-training, 3.44 trillion tokens of mid-training and a long-context phase. Pre-training alone ran on 768 NVIDIA B200 GPUs for about three weeks.
The result is less a single model release than a repeatable European capability for training and improving models. That matters as organizations start treating AI infrastructure as a long-term capability rather than a one-time model purchase.
What Does “Sovereign AI” Actually Mean?
The word sovereign is used often in AI discussions, and it can mean very different things. Running a model somewhere in Europe does not automatically make an AI system sovereign.
Sovereignty is better understood as a set of control points across four layers: the model, the infrastructure it runs on, the data it processes, and the operations that let an organization change course when its requirements change.
Kolibri addresses several of these layers. Its weights are public under Apache 2.0, it can run on infrastructure the organization controls or on a partner it chooses, and Aleph Alpha says it was developed on infrastructure in Germany and Finland. Aleph Alpha is also a signatory of the EU General-Purpose AI Code of Practice and publishes a training-data summary in the European Commission's template.
But open weights are not a complete compliance solution. Organizations still need security controls, access management, infrastructure governance, retention policies, monitoring and documented data processing. GDPR compliance depends on how the whole application and data environment is designed, not on which model it uses.
That is why sovereign AI is ultimately an architecture question.
From Open Weights to Production in Regulated Industries
Public administrations, industrial companies and aerospace organizations often handle information they cannot send to arbitrary external services. A self-hostable open-weight model changes that equation.
A public administration could build an internal assistant on its own document repositories. An industrial company could connect the model to technical documentation and internal knowledge bases. An aerospace organization could run long-context workflows on engineering or operational material while keeping control of the infrastructure. Aleph Alpha's own industry evaluations reflect these scenarios, with RAG tests on German public-sector, aerospace, automotive and semiconductor material.
Aleph Alpha positions Kolibri for human-in-the-loop systems, where a person reviews the output before it is acted on, rather than fully autonomous decision-making. The model does not remove the need for governance. It gives organizations more choice over where and how that governance is implemented, which becomes critical as AI moves from isolated chat interfaces into RAG pipelines, internal agents, document processing and coding workflows that touch private data.
What it takes to run Kolibri
The practical question is not whether an organization can download Kolibri. It is whether it can operate it reliably.
The FP8 model needs about 78 GB of memory for its weights. Aleph Alpha lists these configurations:
- Minimum: 2× A100 80 GB, 2× H100 SXM5, or a single H200, B200 or B300.
- Recommended: 2× H100 SXM5, 2× H200, or a single B200 or B300.
Serving runs on vLLM through Aleph Alpha's aleph-alpha-inference package, available as a container image, and exposes an OpenAI-compatible API. Beyond GPUs, teams also need serving infrastructure, monitoring, networking, model lifecycle management and enough capacity for their expected load.
Long context adds further considerations. A workload that occasionally sends hundreds of thousands of tokens has very different requirements from one that sends short prompts continuously at high throughput.
The decision should therefore be based on the complete workload:
model capability + context requirements + throughput + latency + infrastructure + governance + cost.
Self-hosting provides a high degree of control, but it also moves more operational responsibility to the organization. That is where a multi-provider architecture becomes useful.
Where EdenAI Fits
Sovereignty does not have to mean committing an entire AI stack to one model.
An organization might choose Kolibri where German-language performance, open weights, deployment control or sovereign infrastructure matter most. Other workloads have different requirements, and access to additional models and providers is useful there.
The challenge is that integrating several providers directly quickly becomes its own project. Different APIs, authentication systems, model identifiers, request formats, capabilities, routing logic, monitoring and fallback mechanisms all have to be maintained.
EdenAI provides a provider-agnostic layer that connects applications to multiple AI providers through one unified API. Its European endpoint is designed for workloads that require European data residency: prompts, files, requests and outputs are processed and routed within Europe, with zero data retention and access to EU-compatible providers and models.
This gives two complementary options:
- Self-hosting Kolibri maximizes control over the model and the inference infrastructure.
- A European multi-provider gateway reduces integration complexity while keeping provider choice and European routing requirements.
The two can coexist. A company can run a sovereign model internally and use a gateway for other approved providers when a workload needs different capabilities. The principle is choice without unnecessary integration complexity.
The Bigger Picture
Kolibri is more than another large language model release. It reflects a broader shift in how European organizations think about AI infrastructure.
For years, the conversation focused on access to ever more capable models. For regulated and mission-critical workloads, the questions are now broader: Who controls the model? Where does it run? Where does the data go? How replaceable is the provider? And what happens when the AI landscape changes?
Kolibri answers part of that by combining open weights, a German-English focus, long-context and agentic capabilities, and a design aimed at deployments where control matters. The rest depends on infrastructure, data governance, provider strategy and application architecture.
That is why Kolibri is worth watching: not as an argument for one particular model, but as an example of European AI moving toward systems where organizations control the technology beneath their applications. As that ecosystem grows, the ability to connect these models without giving up provider choice or operational flexibility becomes just as important as the models themselves.



