AI NEWS
IA Générative
8 min de lecture

Quasar 438B vs Mistral Large 4 vs Kolibri: Best European LLM in 2026

Résumez cet article avec :

Résumé

Three European LLMs launched within five weeks of each other. All three offer million-token context windows, but each is built on a very different bet. We compare Multiverse Computing's Quasar 438B, Mistral Large 4 and Aleph Alpha's Kolibri on capability, benchmarks, real cost per task and sovereign deployment. We also show when to use each one, and why many teams may end up routing between all three.

Within five weeks in autumn 2026, three European AI labs shipped their flagship models. Multiverse Computing released Quasar 438B on September 2. Aleph Alpha followed with Kolibri on October 3. Three days later, Mistral Large 4 entered public preview.

They come from three different countries: Spain, Germany and France. All three can handle around one million tokens of context, and all three target demanding enterprise workloads. Yet they make very different trade-offs. Quasar bets on reasoning and agentic coding, Mistral Large 4 on breadth and multimodality, and Kolibri on German-English specialization and sovereign deployment.

That makes "Which is the best European AI model?" the wrong question. The more useful one is: which model fits a given workload, and what does it actually cost to complete that workload?

We have covered each model in depth separately. This article puts them side by side.

Three European models, three different bets

Quasar 438B comes from Multiverse Computing, a San Sebastián company that spent six years compressing other companies' models before building its own. Quasar pairs 438 billion parameters with a 1 million-token context window and always-on reasoning. It scores 43 on the Artificial Analysis Intelligence Index, the highest of any European model and 13th out of 178 models worldwide. Two constraints matter: it supports English and Spanish only, and it is API-only, since the weights have not been released. Its natural territory is complex, multi-step work where planning, tool use and sustained context matter.

Mistral Large 4, nicknamed "Le Chonk", takes the broadest approach. Its granular Mixture-of-Experts architecture holds 1.05 trillion parameters but activates only about 49 billion per token. A 1.6 billion-parameter vision encoder lets it read images as well as text, although it only generates text. Its training data covers more than 160 languages, including every official EU language. Mistral positions it for coding agents, document processing, cybersecurity, finance and industrial visual tasks. During the preview it is available through the API only, and its weights are scheduled for October 27, 2026 under a custom Mistral license. That is a notable shift from Mistral Large 3, which was released under Apache 2.0.

Kolibri is deliberately specialized. It contains 78.1 billion parameters, but only around 3.46 billion (about 4.4%) are active for each token. It was designed around German and English rather than adding German on top of an English model:

  • German makes up more than a fifth of its pre-training data, roughly 4.3 trillion tokens.
  • Its tokenizer is built for German compound words.
  • It reasons in German when it works in German.

It supports tool calling and adjustable reasoning effort, from none to high. Its weights are on Hugging Face under Apache 2.0, which makes it the only one of the three you can download and run on your own infrastructure today.

The important difference is not size. It is what each team chose to optimize for.

Quasar 438B Mistral Large 4 Kolibri
Developer Multiverse Computing (Spain) Mistral AI (France) Aleph Alpha (Germany)
Released Sept 2, 2026 Oct 6, 2026 (preview) Oct 3, 2026
Primary strength Reasoning, coding & agents General-purpose & multimodal German-English & sovereignty
Parameters 438B 1.05T total / 49B active 78.1B total / 3.46B active
Context window 1M tokens 1M tokens 1M max (262K recommended)
Input Text Text + images Text
Languages English, Spanish 160+, all EU languages German, English
Reasoning & tools Always-on reasoning, tool calling Function calling, agents, built-in tools Adjustable reasoning, tool calling
AA Intelligence Index 43 38 Not listed yet
API price / 1M tokens $0.60 in · $1.80 out $0.68 in · $0.07 cached · $2.09 out Open weights: infra cost
Weights Not released (API only) Oct 27, 2026, custom license Available, Apache 2.0
Self-hosting footprint Not possible ~1 TB GPU memory (8-bit) ~78 GB (FP8), e.g. 2× H100

Specifications and prices as published by each vendor at launch. Mistral Large 4 is in public preview and may change.

Benchmarks: strong results, but no single leaderboard

Only one score is directly comparable between two of the three models: the Artificial Analysis Intelligence Index, an independent composite of reasoning, knowledge and agentic evaluations.

  • Quasar scores 43, which means it remains the highest-scoring European model even after Mistral Large 4's release.
  • Mistral Large 4 scores 38, which still makes it the strongest open-weight model from outside China on that index.
  • Kolibri has not been listed yet.

Beyond that shared yardstick, each vendor emphasizes different evaluations, and the choice itself tells you what each model was built for:

  • Quasar 438B: 69.3 on Terminal-Bench v2.1, which measures real tasks in terminal environments rather than isolated coding questions. It also scores 75.0 on AA-LCR, which tests reasoning across long documents rather than simply accepting a long input. That combination points to coding agents, repository analysis and multi-step automation.
  • Mistral Large 4: 62% on DeepSWE v1.1 for long-horizon software engineering, 67% on the Finch finance benchmark, and 73% on DIOR-RSVG visual grounding. Most of these figures are self-reported and come from a preview checkpoint. On the live DeepSWE leaderboard, the best published configurations reach around 69%, so Large 4's result is strong but not a lead.
  • Kolibri: 96.9 on AIME 2025, 84.3 on GPQA Diamond and 85.9 on LiveCodeBench v6, for an overall score of 75.5 in English and 70.8 in German. These are Aleph Alpha's own evaluations, and they compare Kolibri against sparse models in its own weight class, such as Qwen3.6 35B-A3B and Mistral Small 4, rather than against 400B+ models. The meaningful reading is quality per unit of compute, not raw rank.

These results were produced with different evaluations, methodologies and reporting conventions, so they should not be merged into one leaderboard. A benchmark can tell you where a model is strong; it cannot tell you whether that strength matters for your application.

For production teams, the useful question is not "which model has the highest score?" It is which model produces the most useful result for the work you actually need done?

The 1M-token context: what fits vs. what you should send

All three models can reach roughly one million tokens of context. That makes it possible to pass an entire repository, a large document collection or a long-running agent history without aggressively splitting it into chunks.

But a maximum context window is not a recommendation to fill it on every request. Every additional token adds input cost and latency, and more context also means more irrelevant information competing with the evidence that actually matters.

Kolibri illustrates the distinction well. It supports up to 1,048,576 tokens, but its native trained context is 262,144 tokens. Aleph Alpha recommends staying at or below that length for latency-sensitive deployments and complex tasks, and that is the figure to use for production planning. The same principle applies to Quasar and Mistral Large 4: long-context recall should be tested on your own data, not assumed from the window size.

A large context window changes what an application can do. It does not remove the need to decide what the model should actually see.

A larger context window changes what is possible, not what should be sent.
More context More documents, code and history can fit into one request.
More noise Irrelevant or repeated content increases latency and spend.
Better context strategy Retrieval, selection and caching still determine production efficiency.

That decision becomes even more important once cost enters the picture.

Beyond token prices: what does a completed task cost?

Model comparisons usually stop at the price per million input and output tokens. That number is useful, but it can mislead for agentic and long-running applications.

Consider an AI coding agent. The first request includes a large slice of the repository, and the model proposes a change. A tool call runs the tests, which fail. The agent receives the error output, reasons again and retries, perhaps several times, before the task is done.

The application did not buy one output. It bought a workflow. A more useful production metric is:

Cost per successful task = total workflow spend ÷ successfully completed tasks

Total workflow spend includes:

  • input tokens, including repeated and uncached context
  • output and reasoning tokens
  • tool calls
  • retries and failed attempts
MODEL PRICING

Token price is only the starting point

1. Request Input + context
2. Generation Output + reasoning
3. Workflow Tools + repeated context
4. Recovery Retries + failed attempts
Production metric Cost per successful task

This changes how each model's pricing should be read.

Quasar costs $0.60 per million input tokens and $1.80 per million output tokens, and its reasoning is always on. That means output tokens include reasoning, not just the visible answer. In a measured call through Eden AI, a one-sentence question used 22 input tokens and 919 output tokens, for about $0.0017. That is cheap for a hard reasoning step and wasteful for simple extraction, so Quasar is best reserved for the steps that genuinely need it.

Mistral Large 4 lists $0.68 per million input tokens, $2.09 per million output tokens and just $0.07 per million for cached input. For agents that repeatedly send the same instructions, documentation or repository context, caching changes the economics. A 500,000-token prompt costs about $0.34 at the standard rate and about $0.035 from cache. Mistral's pricing page also notes that batch processing halves the price for non-urgent work.

Kolibri raises a different economics question, because it has no token price when you host it yourself. The relevant costs become GPU infrastructure, concurrency, power, engineering time and maintenance. Its sparse design is what makes self-hosting realistic:

  • Kolibri: the FP8 weights need about 78 GB of memory, which fits on two H100s or a single B200.
  • Mistral Large 4: the weights alone need roughly 1 TB at 8-bit, which means multi-GPU, often multi-node, deployments.

This is why comparing only the cheapest token rate can lead to the wrong conclusion. A more expensive model can be cheaper per task if it needs fewer retries, produces fewer invalid outputs or completes multi-step work more reliably. Conversely, an excellent model used on trivial requests makes an application more expensive than it needs to be.

The right objective is not minimum token price. It is minimum cost for an acceptable successful outcome.

Sovereignty is an architecture, not a label

"European AI" is an increasingly loose category. A model built by a European company does not make every deployment that uses it sovereign. Sovereignty depends on four things:

  • the model itself
  • where inference runs
  • where data is processed
  • whether an organization can change or move the system without depending on a single provider

We explored this gap in The Missing Ring in Europe's AI Sovereignty Chain.

AI sovereignty Control over the AI stack
Model Weights, training, licensing and intellectual property
Infrastructure Where inference runs and who controls the hardware
Data Where inputs are processed, stored and governed
Operations Deployment, monitoring, updates and provider independence

The three models sit at very different points on this map.

Kolibri covers the most layers. It was developed in Germany and trained on infrastructure in Germany and Finland. Its weights are public under Apache 2.0, so an organization can deploy it on hardware it controls or with a hosting partner it chooses. Aleph Alpha is also a signatory of the EU General-Purpose AI Code of Practice.

Quasar offers a different kind of value. It gives access to Europe's most capable reasoning model without requiring you to run any infrastructure. Because the weights are not released, however, control over the model and inference stays with the provider.

Mistral Large 4 is in transition. Today it is an API model served from Mistral's European infrastructure, with zero-data-retention options. From October 27, the weights become available under a custom license. That creates a self-hosting path, but both the license terms and the hardware requirements need to be evaluated before relying on it.

Even open weights are not a full compliance answer. GDPR compliance depends on how the whole application handles data, including access controls, retention, logging and monitoring, not on which model sits at its center.

Which model should you choose?

In practice, the choice follows the workload:

  • A software company building an autonomous coding agent will likely prioritize Quasar's reasoning and terminal performance.
  • A multinational building a multilingual assistant across European markets will value Mistral Large 4's language coverage and image understanding.
  • A public administration processing sensitive German documents may value deployment control above benchmark scores, which points to Kolibri.
Workload Best fit Why
Agentic coding Quasar 438B Reasoning, tool use, long context and strong terminal performance.
Complex long-context reasoning Quasar 438B 1M context paired with a 75.0 AA-LCR long-context reasoning score.
Spanish-language products Quasar 438B Spanish is a first-class language alongside English.
Multilingual enterprise applications Mistral Large 4 160+ languages, including every official EU language.
Multimodal workflows Mistral Large 4 Text and image input, strong visual grounding, 1M-token context.
Agents with large repeated context Mistral Large 4 Cached input at $0.07 per million tokens.
German-language applications Kolibri Trained and tokenized specifically for German and English.
Sovereign or self-hosted deployment Kolibri Apache 2.0 weights that run on two H100s or one B200.

Notice that each model wins several rows. That is the real takeaway: these choices are not mutually exclusive, and a production system does not need to commit to one model for every task.

Why the strongest architecture may use all three

Imagine an enterprise application that handles several kinds of requests:

  • Simple classification and extraction go to a small, inexpensive model.
  • A multilingual document containing images goes to Mistral Large 4.
  • A difficult reasoning or coding step goes to Quasar.
  • A German-language workflow involving sensitive internal data runs on a self-hosted Kolibri deployment.
APPLICATION One AI workflow
ROUTING LAYER Choose by task, cost, language & constraints
Quasar 438B
Hard reasoning
Agentic coding
Mistral Large 4
Multilingual
Multimodal
Kolibri
German / English
Sovereign workloads
Lightweight model
Classification
Extraction

This is where a model gateway becomes useful. Instead of embedding provider-specific integrations throughout an application, teams put a routing layer between the application and the models. That layer can then weigh capability, language, cost, latency, availability and data requirements on every request. Our guide to LLM routing strategies covers the main approaches.

The conceptual shift matters: model selection becomes a runtime decision rather than a permanent architectural commitment. An organization could start with Mistral Large 4 as its general-purpose route, send its hardest reasoning tasks to Quasar, and reserve Kolibri for workloads where German specialization or self-hosting matters. The same architecture provides fallback: if a model becomes unavailable, too slow or unsuitable for a request, another route takes over.

With Eden AI, Quasar is already callable as compactifai/quasar-438b through the same OpenAI-compatible endpoint as hundreds of other models. You can check the current availability of other models in the Eden AI model catalog. For workloads that must stay in Europe, the EU endpoint keeps prompts, files and outputs routed within Europe with zero data retention. That route can run alongside a self-hosted Kolibri deployment.

This is also why benchmarking should happen at the workflow level. Run the same real tasks through several routes and measure:

  • task success rate
  • accepted-output rate
  • number of retries
  • total tokens consumed, including reasoning
  • latency
  • tool-call reliability
  • cost per successful task
  • data and deployment requirements

The result is a model comparison much closer to what production actually looks like.

What this says about European AI in 2026

The most interesting conclusion is not that one of these three models is Europe's winner. It is that European AI is no longer converging on a single strategy:

  • Mistral is pursuing scale, breadth and multimodality.
  • Multiverse Computing has shown that a compression specialist can reach the top of Europe's capability rankings.
  • Aleph Alpha is building around specialization, deployment control and sovereignty.

The map has broadened, too. European AI was largely a French and German story for two years; its highest-scoring model is now Spanish. For a fuller view of the ecosystem, see our overview of the best European AI model providers.

For teams building AI products, that diversity is good news. The best model is not necessarily the one with the highest benchmark score or the lowest token price. It is the one that delivers the right combination of capability, reliability, economics and control for each task.

As applications become more agentic, the definition of cost will keep shifting from price per token to price per successful outcome. That may be the most important metric to watch this year.

FAQ

There is no single best model for every workload. Quasar 438B has the highest European score on the Artificial Analysis Intelligence Index (43) and is strongest for reasoning and agentic coding. Mistral Large 4 is the broadest choice for multilingual and multimodal applications. Aleph Alpha's Kolibri stands out for German-English workloads and sovereign, self-hosted deployment.
They optimize for different workloads. Quasar scores higher on the independent Artificial Analysis Intelligence Index (43 vs. 38) and is particularly strong for reasoning, long-context work and agentic coding. Mistral Large 4 accepts images, covers more than 160 languages and offers cheap cached input. The better choice depends on the tasks your application needs to complete.
Kolibri is designed specifically around German and English. German makes up more than a fifth of its pre-training data, and its tokenizer is optimized for German compound words. Mistral Large 4 is the stronger choice when German is one language among many in a broader multilingual application. Quasar does not support German.
Kolibri can be self-hosted today: its weights are available under Apache 2.0 and need about 78 GB of memory in FP8. Mistral Large 4's weights are scheduled for October 27, 2026 under a custom Mistral license, and they require roughly 1 TB of GPU memory at 8-bit. Quasar 438B is API-only, since its weights have not been released.
No. A production workflow can include multiple model calls, long contexts, reasoning tokens, tool calls, failed attempts and retries. A better metric is cost per successful task, which measures how much the system spends to produce an acceptable completed outcome.
Not necessarily. A large context window makes larger inputs possible, but sending unnecessary information increases cost and latency and can dilute the relevant evidence. Kolibri, for example, supports 1M tokens but recommends 262K for production. Retrieval, context selection and caching remain important even with million-token models.
Kolibri has the clearest sovereign-deployment profile of the three. Its weights are available under Apache 2.0 and can run on infrastructure the organization controls. Sovereignty still depends on the complete architecture, including infrastructure, data handling and operations, not on the model alone.
Not necessarily. Different models can be routed to different tasks based on capability, cost, latency, language and deployment requirements. A multi-model architecture also provides fallback when a primary model is unavailable or unsuitable for a request.
Quasar 438B is available on Eden AI as compactifai/quasar-438b through an OpenAI-compatible endpoint. Availability of other models should be checked in the Eden AI model catalog, as integrations are added regularly. Eden AI's EU endpoint keeps requests and data routed within Europe with zero data retention, and it can run alongside a self-hosted model such as Kolibri.

Articles similaires

AI NEWS
IA Générative
Mistral Large 4: Specs, Benchmarks, Pricing & Use Cases
10/6/2026
·
Written byClément Moreau
AI NEWS
IA Générative
Aleph Alpha Kolibri: Open-Weight German-English LLM for Sovereign AI
10/5/2026
·
Written byClément Moreau
AI NEWS
IA Générative
Google Gemini 4 Argon: Cost and Performance Benchmarks for Production AI Systems
10/1/2026
·
Written byClément Moreau
COMMENCEZ

Commencez à créer avec Eden AI

Une interface unique pour intégrer les meilleures technologies d’IA dans vos flux de travail.