Résumez cet article avec :
Three European LLMs launched within five weeks of each other. All three offer million-token context windows, but each is built on a very different bet. We compare Multiverse Computing's Quasar 438B, Mistral Large 4 and Aleph Alpha's Kolibri on capability, benchmarks, real cost per task and sovereign deployment. We also show when to use each one, and why many teams may end up routing between all three.
Within five weeks in autumn 2026, three European AI labs shipped their flagship models. Multiverse Computing released Quasar 438B on September 2. Aleph Alpha followed with Kolibri on October 3. Three days later, Mistral Large 4 entered public preview.
They come from three different countries: Spain, Germany and France. All three can handle around one million tokens of context, and all three target demanding enterprise workloads. Yet they make very different trade-offs. Quasar bets on reasoning and agentic coding, Mistral Large 4 on breadth and multimodality, and Kolibri on German-English specialization and sovereign deployment.
That makes "Which is the best European AI model?" the wrong question. The more useful one is: which model fits a given workload, and what does it actually cost to complete that workload?
We have covered each model in depth separately. This article puts them side by side.
Three European models, three different bets
Quasar 438B comes from Multiverse Computing, a San Sebastián company that spent six years compressing other companies' models before building its own. Quasar pairs 438 billion parameters with a 1 million-token context window and always-on reasoning. It scores 43 on the Artificial Analysis Intelligence Index, the highest of any European model and 13th out of 178 models worldwide. Two constraints matter: it supports English and Spanish only, and it is API-only, since the weights have not been released. Its natural territory is complex, multi-step work where planning, tool use and sustained context matter.
Mistral Large 4, nicknamed "Le Chonk", takes the broadest approach. Its granular Mixture-of-Experts architecture holds 1.05 trillion parameters but activates only about 49 billion per token. A 1.6 billion-parameter vision encoder lets it read images as well as text, although it only generates text. Its training data covers more than 160 languages, including every official EU language. Mistral positions it for coding agents, document processing, cybersecurity, finance and industrial visual tasks. During the preview it is available through the API only, and its weights are scheduled for October 27, 2026 under a custom Mistral license. That is a notable shift from Mistral Large 3, which was released under Apache 2.0.
Kolibri is deliberately specialized. It contains 78.1 billion parameters, but only around 3.46 billion (about 4.4%) are active for each token. It was designed around German and English rather than adding German on top of an English model:
- German makes up more than a fifth of its pre-training data, roughly 4.3 trillion tokens.
- Its tokenizer is built for German compound words.
- It reasons in German when it works in German.
It supports tool calling and adjustable reasoning effort, from none to high. Its weights are on Hugging Face under Apache 2.0, which makes it the only one of the three you can download and run on your own infrastructure today.
The important difference is not size. It is what each team chose to optimize for.
Benchmarks: strong results, but no single leaderboard
Only one score is directly comparable between two of the three models: the Artificial Analysis Intelligence Index, an independent composite of reasoning, knowledge and agentic evaluations.
- Quasar scores 43, which means it remains the highest-scoring European model even after Mistral Large 4's release.
- Mistral Large 4 scores 38, which still makes it the strongest open-weight model from outside China on that index.
- Kolibri has not been listed yet.
Beyond that shared yardstick, each vendor emphasizes different evaluations, and the choice itself tells you what each model was built for:
- Quasar 438B: 69.3 on Terminal-Bench v2.1, which measures real tasks in terminal environments rather than isolated coding questions. It also scores 75.0 on AA-LCR, which tests reasoning across long documents rather than simply accepting a long input. That combination points to coding agents, repository analysis and multi-step automation.
- Mistral Large 4: 62% on DeepSWE v1.1 for long-horizon software engineering, 67% on the Finch finance benchmark, and 73% on DIOR-RSVG visual grounding. Most of these figures are self-reported and come from a preview checkpoint. On the live DeepSWE leaderboard, the best published configurations reach around 69%, so Large 4's result is strong but not a lead.
- Kolibri: 96.9 on AIME 2025, 84.3 on GPQA Diamond and 85.9 on LiveCodeBench v6, for an overall score of 75.5 in English and 70.8 in German. These are Aleph Alpha's own evaluations, and they compare Kolibri against sparse models in its own weight class, such as Qwen3.6 35B-A3B and Mistral Small 4, rather than against 400B+ models. The meaningful reading is quality per unit of compute, not raw rank.
These results were produced with different evaluations, methodologies and reporting conventions, so they should not be merged into one leaderboard. A benchmark can tell you where a model is strong; it cannot tell you whether that strength matters for your application.
For production teams, the useful question is not "which model has the highest score?" It is which model produces the most useful result for the work you actually need done?
The 1M-token context: what fits vs. what you should send
All three models can reach roughly one million tokens of context. That makes it possible to pass an entire repository, a large document collection or a long-running agent history without aggressively splitting it into chunks.
But a maximum context window is not a recommendation to fill it on every request. Every additional token adds input cost and latency, and more context also means more irrelevant information competing with the evidence that actually matters.
Kolibri illustrates the distinction well. It supports up to 1,048,576 tokens, but its native trained context is 262,144 tokens. Aleph Alpha recommends staying at or below that length for latency-sensitive deployments and complex tasks, and that is the figure to use for production planning. The same principle applies to Quasar and Mistral Large 4: long-context recall should be tested on your own data, not assumed from the window size.
A large context window changes what an application can do. It does not remove the need to decide what the model should actually see.
That decision becomes even more important once cost enters the picture.
Beyond token prices: what does a completed task cost?
Model comparisons usually stop at the price per million input and output tokens. That number is useful, but it can mislead for agentic and long-running applications.
Consider an AI coding agent. The first request includes a large slice of the repository, and the model proposes a change. A tool call runs the tests, which fail. The agent receives the error output, reasons again and retries, perhaps several times, before the task is done.
The application did not buy one output. It bought a workflow. A more useful production metric is:
Cost per successful task = total workflow spend ÷ successfully completed tasks
Total workflow spend includes:
- input tokens, including repeated and uncached context
- output and reasoning tokens
- tool calls
- retries and failed attempts
This changes how each model's pricing should be read.
Quasar costs $0.60 per million input tokens and $1.80 per million output tokens, and its reasoning is always on. That means output tokens include reasoning, not just the visible answer. In a measured call through Eden AI, a one-sentence question used 22 input tokens and 919 output tokens, for about $0.0017. That is cheap for a hard reasoning step and wasteful for simple extraction, so Quasar is best reserved for the steps that genuinely need it.
Mistral Large 4 lists $0.68 per million input tokens, $2.09 per million output tokens and just $0.07 per million for cached input. For agents that repeatedly send the same instructions, documentation or repository context, caching changes the economics. A 500,000-token prompt costs about $0.34 at the standard rate and about $0.035 from cache. Mistral's pricing page also notes that batch processing halves the price for non-urgent work.
Kolibri raises a different economics question, because it has no token price when you host it yourself. The relevant costs become GPU infrastructure, concurrency, power, engineering time and maintenance. Its sparse design is what makes self-hosting realistic:
- Kolibri: the FP8 weights need about 78 GB of memory, which fits on two H100s or a single B200.
- Mistral Large 4: the weights alone need roughly 1 TB at 8-bit, which means multi-GPU, often multi-node, deployments.
This is why comparing only the cheapest token rate can lead to the wrong conclusion. A more expensive model can be cheaper per task if it needs fewer retries, produces fewer invalid outputs or completes multi-step work more reliably. Conversely, an excellent model used on trivial requests makes an application more expensive than it needs to be.
The right objective is not minimum token price. It is minimum cost for an acceptable successful outcome.
Sovereignty is an architecture, not a label
"European AI" is an increasingly loose category. A model built by a European company does not make every deployment that uses it sovereign. Sovereignty depends on four things:
- the model itself
- where inference runs
- where data is processed
- whether an organization can change or move the system without depending on a single provider
We explored this gap in The Missing Ring in Europe's AI Sovereignty Chain.
The three models sit at very different points on this map.
Kolibri covers the most layers. It was developed in Germany and trained on infrastructure in Germany and Finland. Its weights are public under Apache 2.0, so an organization can deploy it on hardware it controls or with a hosting partner it chooses. Aleph Alpha is also a signatory of the EU General-Purpose AI Code of Practice.
Quasar offers a different kind of value. It gives access to Europe's most capable reasoning model without requiring you to run any infrastructure. Because the weights are not released, however, control over the model and inference stays with the provider.
Mistral Large 4 is in transition. Today it is an API model served from Mistral's European infrastructure, with zero-data-retention options. From October 27, the weights become available under a custom license. That creates a self-hosting path, but both the license terms and the hardware requirements need to be evaluated before relying on it.
Even open weights are not a full compliance answer. GDPR compliance depends on how the whole application handles data, including access controls, retention, logging and monitoring, not on which model sits at its center.
Which model should you choose?
In practice, the choice follows the workload:
- A software company building an autonomous coding agent will likely prioritize Quasar's reasoning and terminal performance.
- A multinational building a multilingual assistant across European markets will value Mistral Large 4's language coverage and image understanding.
- A public administration processing sensitive German documents may value deployment control above benchmark scores, which points to Kolibri.
Notice that each model wins several rows. That is the real takeaway: these choices are not mutually exclusive, and a production system does not need to commit to one model for every task.
Why the strongest architecture may use all three
Imagine an enterprise application that handles several kinds of requests:
- Simple classification and extraction go to a small, inexpensive model.
- A multilingual document containing images goes to Mistral Large 4.
- A difficult reasoning or coding step goes to Quasar.
- A German-language workflow involving sensitive internal data runs on a self-hosted Kolibri deployment.
This is where a model gateway becomes useful. Instead of embedding provider-specific integrations throughout an application, teams put a routing layer between the application and the models. That layer can then weigh capability, language, cost, latency, availability and data requirements on every request. Our guide to LLM routing strategies covers the main approaches.
The conceptual shift matters: model selection becomes a runtime decision rather than a permanent architectural commitment. An organization could start with Mistral Large 4 as its general-purpose route, send its hardest reasoning tasks to Quasar, and reserve Kolibri for workloads where German specialization or self-hosting matters. The same architecture provides fallback: if a model becomes unavailable, too slow or unsuitable for a request, another route takes over.
With Eden AI, Quasar is already callable as compactifai/quasar-438b through the same OpenAI-compatible endpoint as hundreds of other models. You can check the current availability of other models in the Eden AI model catalog. For workloads that must stay in Europe, the EU endpoint keeps prompts, files and outputs routed within Europe with zero data retention. That route can run alongside a self-hosted Kolibri deployment.
This is also why benchmarking should happen at the workflow level. Run the same real tasks through several routes and measure:
- task success rate
- accepted-output rate
- number of retries
- total tokens consumed, including reasoning
- latency
- tool-call reliability
- cost per successful task
- data and deployment requirements
The result is a model comparison much closer to what production actually looks like.
What this says about European AI in 2026
The most interesting conclusion is not that one of these three models is Europe's winner. It is that European AI is no longer converging on a single strategy:
- Mistral is pursuing scale, breadth and multimodality.
- Multiverse Computing has shown that a compression specialist can reach the top of Europe's capability rankings.
- Aleph Alpha is building around specialization, deployment control and sovereignty.
The map has broadened, too. European AI was largely a French and German story for two years; its highest-scoring model is now Spanish. For a fuller view of the ecosystem, see our overview of the best European AI model providers.
For teams building AI products, that diversity is good news. The best model is not necessarily the one with the highest benchmark score or the lowest token price. It is the one that delivers the right combination of capability, reliability, economics and control for each task.
As applications become more agentic, the definition of cost will keep shifting from price per token to price per successful outcome. That may be the most important metric to watch this year.




