AI NEWS
Generative AI
8 min reading

Aleph Alpha Kolibri: Open-Weight German-English LLM for Sovereign AI

Summarize this article with:

summary

Aleph Alpha's Kolibri is a 78.1B-parameter Mixture-of-Experts model built for German and English, released with open weights under Apache 2.0. This article covers its architecture, 1M-token context, published benchmarks, hardware requirements and what it means for sovereign AI in Europe, including how it fits into a multi-provider architecture with EdenAI.

As AI moves deeper into government, industry and other regulated environments, the question is no longer only how capable a model is. It is also where it runs, who controls it, what happens to the data, and how freely an organization can change its infrastructure or providers.

That is the problem Aleph Alpha is targeting with Kolibri (German for "hummingbird"), an English-German Mixture-of-Experts model released on October 3, 2026, the Day of German Unity.**

What makes Kolibri significant is the combination: open weights, European development, German and English specialization, efficient inference, long context, and the option to deploy it on an organization's own infrastructure and under its own governance model.

That makes Kolibri a useful case study in what European AI sovereignty looks like when it becomes an engineering problem rather than only a policy objective.

What Is Kolibri?

Kolibri is a reasoning model developed by Aleph Alpha with a specific focus on German and English. It was trained from scratch and contains 78.1 billion parameters in total, of which only about 3.46 billion (4.4%) are active for each token.

It is built for long-context and agentic workloads. Aleph Alpha lists multi-step reasoning, retrieval-augmented generation (RAG), agentic tool calling, coding and German- and English-language assistants among its primary use cases.

The model is published on Hugging Face as Kolibri-1, in an FP8 version and a BF16 version, under the Apache 2.0 license. It is the first language model Aleph Alpha has released under Apache 2.0. Its earlier open models used a non-commercial license.

"Open-weight" is the more precise description than "open source". The Apache 2.0 license covers the published weights and configuration files, which gives organizations broad freedom to deploy, adapt and commercialize the model. Aleph Alpha explicitly keeps the rights to its training code, training methods and other artifacts outside the repository.

MODEL AT A GLANCE

Kolibri

A sovereign-focused English-German Mixture-of-Experts model from Aleph Alpha, released with open weights on October 3, 2026.

Total parameters 78.1B
Active / token 3.46B About 4.4% of the model
Experts per layer 384 + 1 shared 6 routed experts per token
Languages German · English
Context 1M tokens max 262K native and recommended
Reasoning & tools None to high Adjustable effort, tool calling
License Apache 2.0 Open weights
FP8 footprint ~78 GB BF16 version also available

Inside Kolibri: MoE, Reasoning and Long Context

Kolibri's architecture is built on a simple trade-off: a large model capacity does not mean the full model has to run for every token.

The model has 50 transformer layers. In each Mixture-of-Experts layer, a router scores 384 specialized "experts" and sends each token to the 6 most relevant ones, while 1 shared expert always runs. Only that small subset participates in the computation.

How Kolibri's sparse Mixture-of-Experts architecture works

Input token Each token enters the MoE layer
→
Router Picks the 6 most relevant of 384 experts
→
Sparse expert activation
Shared expert (always on)
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
→
Output Selected expert outputs are combined

Illustrative view of one layer, repeated across Kolibri's 50 layers. Of 78.1B total parameters, only around 3.46B are active for each token.

This matters because serving a model involves more than counting parameters. Sparse activation keeps the computation per token low, but the full set of weights still has to sit in memory. "3.46B active parameters" therefore does not mean Kolibri can be served like a conventional 3.46B-parameter model.

Kolibri also keeps long contexts affordable through a hybrid attention design. Four out of every five attention layers only look at the most recent 512 tokens, while the remaining layers attend to the full context.

Controllable reasoning and tool calling

Kolibri exposes adjustable reasoning effort: none, low, medium or high, set per request. Applications can trade response speed and compute against more deliberate reasoning only when a task needs it.

Tool calling is supported and can be combined with reasoning. That makes the model suitable for agentic workflows where it has to call APIs, run searches or execute code, rather than only generate text.

The 1M-token context comes with an important detail

Kolibri is advertised with a context length of up to 1,048,576 tokens, roughly one million. That number needs context.

The model was trained on sequences of up to 262,144 tokens, its native context length. Because of how positional information is handled, it can be extended beyond that without retraining, and Aleph Alpha has validated quality and serving efficiency up to 1M tokens. Serving more than 262K tokens requires an explicit configuration change, and Aleph Alpha recommends staying at 262,144 tokens or less for latency-sensitive deployments and complex tasks.

For production systems, the distinction matters. A million-token context is valuable for large document collections, long technical specifications and regulatory material. But latency, memory consumption, throughput and the cost of processing very large prompts still apply.

The real value of long context is not simply "more tokens". It is the ability to design applications around larger working sets of information when the workload actually benefits from them.

LONG CONTEXT

1M tokens, with a practical 262K recommendation

262,144 tokens Native trained context, recommended for efficient serving and complex tasks
1,048,576 tokens Maximum validated context, enabled through an extra serving setting

The maximum context window and the recommended production context are not the same thing. Choose the context size that matches your actual information and latency requirements.

A Model Built Around German and English

Kolibri's language focus is one of its most distinctive characteristics.

Instead of adding German on top of a primarily English model, Aleph Alpha designed Kolibri around two languages by choice, favoring depth over breadth. German accounts for more than one-fifth of the pre-training data, about 4.3 trillion tokens, alongside roughly 62% English and 14% code. Aleph Alpha also emphasized organic German data rather than relying mainly on translations from English.

The tokenizer was built for the same purpose. It is designed around German word structure, including compound words, so German text is split into fewer tokens without making English less efficient. Fewer tokens per document means lower cost and more usable context for the same German text.

The focus continues into post-training: Kolibri was trained to reason in German when working in German, so users can follow its reasoning in their own language.

Language quality is not only about grammatically correct output. Tokenization, training data, terminology, domain knowledge and the language mix throughout training all affect how well a model handles real-world documents. For organizations that work with German administrative, industrial, legal or technical material, this is a central part of Kolibri's positioning.

How Does Kolibri Perform?

Aleph Alpha evaluates Kolibri across knowledge, math, coding, agentic tasks, grounding, long context and industry-specific RAG, in both English and German.

Headline results include 96.9 on AIME 2025, 84.3 on GPQA Diamond, 85.9 on LiveCodeBench v6 and 64.5 on LongBench Pro. Across all categories, Kolibri reaches an overall score of 75.5 in English and 70.8 in German, the highest among the Mixture-of-Experts models in Aleph Alpha's comparison.

It does not lead everywhere. On LongBench Pro, for example, Qwen3.6 35B-A3B scores higher (70.8), and several competitors do better on some knowledge and grounding benchmarks. These are also Aleph Alpha's own evaluations, run with its open-source evaluation framework, rather than an independent benchmark study.

BENCHMARKS

Kolibri vs. comparable sparse models

Scores out of 100, as published by Aleph Alpha. Higher is better.

Kolibri (3.46B active)
Qwen3.6 35B-A3B (3B active)
Mistral Small 4 119B-A6B (6B active)
OverallEnglish
75.5
71.4
63.1
OverallGerman
70.8
67.3
61.4
AIME 2025Math, English
96.9
84.6
79.8
GPQA DiamondScience, English
84.3
83.4
74.7
LiveCodeBench v6Coding
85.9
82.5
71.2
LongBench ProLong context
64.5
70.8
56.4

Source: Kolibri-1 model card (Aleph Alpha, October 2026). Kolibri evaluated at reasoning effort "high". Overall scores are unweighted averages across Aleph Alpha's benchmark categories.

The more interesting point is the relationship between quality and serving cost. Aleph Alpha positions Kolibri on a quality-throughput Pareto frontier: strong output quality for the amount of compute needed to serve it.

For production AI, that is often what decides. A model does not run in isolation from its infrastructure, and a slightly higher benchmark score can matter less than the combination of quality, throughput, hardware requirements, latency and operational control.

From Kolibri Origin to Kolibri

Kolibri was not a single jump from an empty training run to a 78B model.

Aleph Alpha first built Kolibri Origin, a smaller model with 30B total and 3B active parameters and a 65K-token context window. It served to validate and automate the training pipeline before the approach was scaled up.

That pipeline covers data ingestion and curation, experiments and ablations, pre-training, post-training and evaluation. Aleph Alpha highlights automation that kept training running through hardware failures or interrupted data connections without manual intervention.

The full model was then trained in three stages: about 20 trillion tokens of pre-training, 3.44 trillion tokens of mid-training and a long-context phase. Pre-training alone ran on 768 NVIDIA B200 GPUs for about three weeks.

The result is less a single model release than a repeatable European capability for training and improving models. That matters as organizations start treating AI infrastructure as a long-term capability rather than a one-time model purchase.

What Does “Sovereign AI” Actually Mean?

The word sovereign is used often in AI discussions, and it can mean very different things. Running a model somewhere in Europe does not automatically make an AI system sovereign.

Sovereignty is better understood as a set of control points across four layers: the model, the infrastructure it runs on, the data it processes, and the operations that let an organization change course when its requirements change.

AI SOVEREIGNTY

Sovereignty is an architecture, not a model checkbox

01

Model

Who controls the model, weights, lineage, and ability to deploy or adapt it?

02

Infrastructure

Where does inference run, and who controls the underlying compute environment?

03

Data

Where are prompts, documents, outputs, logs, and other sensitive information processed?

04

Operations

Can the organization change providers, infrastructure, or deployment strategies when requirements change?

Kolibri addresses several of these layers. Its weights are public under Apache 2.0, it can run on infrastructure the organization controls or on a partner it chooses, and Aleph Alpha says it was developed on infrastructure in Germany and Finland. Aleph Alpha is also a signatory of the EU General-Purpose AI Code of Practice and publishes a training-data summary in the European Commission's template.

But open weights are not a complete compliance solution. Organizations still need security controls, access management, infrastructure governance, retention policies, monitoring and documented data processing. GDPR compliance depends on how the whole application and data environment is designed, not on which model it uses.

That is why sovereign AI is ultimately an architecture question.

From Open Weights to Production in Regulated Industries

Public administrations, industrial companies and aerospace organizations often handle information they cannot send to arbitrary external services. A self-hostable open-weight model changes that equation.

A public administration could build an internal assistant on its own document repositories. An industrial company could connect the model to technical documentation and internal knowledge bases. An aerospace organization could run long-context workflows on engineering or operational material while keeping control of the infrastructure. Aleph Alpha's own industry evaluations reflect these scenarios, with RAG tests on German public-sector, aerospace, automotive and semiconductor material.

Aleph Alpha positions Kolibri for human-in-the-loop systems, where a person reviews the output before it is acted on, rather than fully autonomous decision-making. The model does not remove the need for governance. It gives organizations more choice over where and how that governance is implemented, which becomes critical as AI moves from isolated chat interfaces into RAG pipelines, internal agents, document processing and coding workflows that touch private data.

What it takes to run Kolibri

The practical question is not whether an organization can download Kolibri. It is whether it can operate it reliably.

The FP8 model needs about 78 GB of memory for its weights. Aleph Alpha lists these configurations:

  • Minimum: 2× A100 80 GB, 2× H100 SXM5, or a single H200, B200 or B300.
  • Recommended: 2× H100 SXM5, 2× H200, or a single B200 or B300.

Serving runs on vLLM through Aleph Alpha's aleph-alpha-inference package, available as a container image, and exposes an OpenAI-compatible API. Beyond GPUs, teams also need serving infrastructure, monitoring, networking, model lifecycle management and enough capacity for their expected load.

Long context adds further considerations. A workload that occasionally sends hundreds of thousands of tokens has very different requirements from one that sends short prompts continuously at high throughput.

The decision should therefore be based on the complete workload:

model capability + context requirements + throughput + latency + infrastructure + governance + cost.

Self-hosting provides a high degree of control, but it also moves more operational responsibility to the organization. That is where a multi-provider architecture becomes useful.

Where EdenAI Fits

Sovereignty does not have to mean committing an entire AI stack to one model.

An organization might choose Kolibri where German-language performance, open weights, deployment control or sovereign infrastructure matter most. Other workloads have different requirements, and access to additional models and providers is useful there.

The challenge is that integrating several providers directly quickly becomes its own project. Different APIs, authentication systems, model identifiers, request formats, capabilities, routing logic, monitoring and fallback mechanisms all have to be maintained.

EdenAI provides a provider-agnostic layer that connects applications to multiple AI providers through one unified API. Its European endpoint is designed for workloads that require European data residency: prompts, files, requests and outputs are processed and routed within Europe, with zero data retention and access to EU-compatible providers and models.

MULTI-PROVIDER ARCHITECTURE

One application, multiple approved AI providers

Your application One integration
↓
EdenAI Unified API · routing · provider abstraction
↓
Kolibri Self-hosted or partner-hosted for sovereign workloads
European managed providers Approved alternative models
Other EU-routed models Additional capabilities and fallback

This gives two complementary options:

  • Self-hosting Kolibri maximizes control over the model and the inference infrastructure.
  • A European multi-provider gateway reduces integration complexity while keeping provider choice and European routing requirements.

The two can coexist. A company can run a sovereign model internally and use a gateway for other approved providers when a workload needs different capabilities. The principle is choice without unnecessary integration complexity.

Approach Infrastructure control Provider choice Operational effort Best suited for
Self-hosted Kolibri Very high Limited to models you host High Strict control and private workloads
European managed provider Medium–high Provider-specific Low–medium Teams reducing infrastructure responsibility
Multi-provider gateway Depends on deployment High Medium Applications requiring model and provider flexibility

The Bigger Picture

Kolibri is more than another large language model release. It reflects a broader shift in how European organizations think about AI infrastructure.

For years, the conversation focused on access to ever more capable models. For regulated and mission-critical workloads, the questions are now broader: Who controls the model? Where does it run? Where does the data go? How replaceable is the provider? And what happens when the AI landscape changes?

Kolibri answers part of that by combining open weights, a German-English focus, long-context and agentic capabilities, and a design aimed at deployments where control matters. The rest depends on infrastructure, data governance, provider strategy and application architecture.

That is why Kolibri is worth watching: not as an argument for one particular model, but as an example of European AI moving toward systems where organizations control the technology beneath their applications. As that ecosystem grows, the ability to connect these models without giving up provider choice or operational flexibility becomes just as important as the models themselves.

FAQ

Kolibri is an English-German Mixture-of-Experts model developed by Aleph Alpha and released on October 3, 2026. It has 78.1 billion total parameters, with 3.46 billion active per token, and supports adjustable reasoning, tool calling and long-context workloads.
Kolibri is an open-weight model. Its weights and configuration files are published on Hugging Face under the Apache 2.0 license, which allows commercial use and adaptation. Aleph Alpha keeps the rights to its training code and methods, so open-weight is the more precise description than open source.
Yes. Kolibri supports up to 1,048,576 tokens. Its native trained context is 262,144 tokens, and Aleph Alpha recommends staying at or below that length for serving efficiency and complex tasks.
German makes up more than one-fifth of Kolibri's pre-training data, about 4.3 trillion tokens. The model also uses a tokenizer designed for German word structure and was trained to reason in German.
Yes. The weights can be downloaded from Hugging Face and served with vLLM through Aleph Alpha's aleph-alpha-inference package, which exposes an OpenAI-compatible API.
The FP8 weights take about 78 GB of memory. The minimum is two A100 80 GB or two H100 SXM5 GPUs, or a single H200, B200 or B300. Aleph Alpha recommends two H100 SXM5, two H200, or a single B200 or B300.
Kolibri is an open-weight European model that organizations can deploy on infrastructure they control or choose. This increases control over the model, infrastructure and data processing, although sovereignty and compliance still depend on the complete system architecture.
EdenAI provides a unified API for multiple AI providers. Its EU endpoint is designed for European data residency, with zero data retention and EU-compatible providers, so organizations can combine provider choice with European routing requirements.

Similar articles

AI NEWS
Generative AI
Google Gemini 4 Argon: Cost and Performance Benchmarks for Production AI Systems
10/1/2026
·
Written byClément Moreau
AI NEWS
Generative AI
When Your LLM Router Turns Against You: The Hidden Security Risk in AI Agents
9/15/2026
·
Written byClément Moreau
AI NEWS
Generative AI
Quasar 438B is on Eden AI: Benchmarks, Pricing, and API Access
9/4/2026
·
Written byClément Moreau
let’s start

Start building with Eden AI

A single interface to integrate the best AI technologies into your products.