AI NEWS
Generative AI
8 min reading

Mistral Large 4: Specs, Benchmarks, Pricing & Use Cases

Summarize this article with:

summary

Mistral Large 4 is Mistral AI's trillion-parameter, multimodal flagship: a sparse Mixture-of-Experts model with 49B active parameters, a 1M-token context window and open weights scheduled for October 27, 2026. Here is what it offers, how it performs and what it costs.

Mistral AI has released Mistral Large 4, a multimodal flagship model that takes its Large family to the trillion-parameter scale while keeping a strong focus on deployment flexibility. The model is nicknamed "Le Chonk." According to Mistral's documentation, it combines several things:

  • a granular Mixture-of-Experts (MoE) architecture with 1.05 trillion total parameters and 49 billion active parameters
  • a 1.6 billion-parameter vision encoder
  • a 1 million-token context window

Large 4 entered public preview through the Mistral API on October 6, 2026, and its open weights are scheduled for the end of October. Mistral positions it for demanding workloads such as software engineering, cybersecurity, finance, manufacturing and visual analysis. In its launch announcement, Mistral calls it the strongest open-weight model from the US or Europe on aggregated benchmarks.

Mistral Large 4

A frontier-scale model with sparse inference

1.05T
total parameters
49B
active parameters
1M
token context
1.6B
vision encoder
Text & image input Text output Sparse MoE Tools & agents

What is Mistral Large 4?

Mistral Large 4 is a general-purpose model available in public preview under the identifier mistral-large-4. It accepts both text and images as input and generates text. Mistral says it trained the model from scratch in its own European data centers, over roughly two months, on close to 4,000 NVIDIA Grace Blackwell GPUs. The training data covers more than 160 languages, including every official EU language (VentureBeat).

Large 4 is also the first model Mistral has released since its €3 billion Series D. It is a substantial step up from Mistral Large 3, which had 675 billion total and 41 billion active parameters.

Mistral Large 4 at a glance

DeveloperMistral AI
Model IDmistral-large-4 (public preview)
ReleaseAPI preview on October 6, 2026; open weights scheduled for the end of October, 2026
ArchitectureGranular Mixture-of-Experts
Total parameters1.05 trillion
Active parameters49 billion per token
Vision encoder1.6 billion parameters
Context window1 million tokens
Input / outputText and images in, text out
Languages160+, including all official EU languages
TrainingFrom scratch on ~4,000 NVIDIA Grace Blackwell GPUs in Mistral's own data centers
CapabilitiesStructured outputs, function calling, document Q&A, chat completions, batching, agents and conversations, built-in tools
API pricing (per 1M tokens)$0.68 input · $0.07 cached input · $2.09 output
Weights licenseCustom Mistral license (at weights release)

Specifications and pricing reflect Mistral's published information at launch and may change after the preview period.

The feature set targets application development rather than plain text generation. Large 4 is most relevant where the model is one component of a larger system, for example:

  • an agent that calls APIs
  • an assistant that processes documents
  • a coding agent working across a repository
  • an application that links visual inputs to language-based reasoning

Architecture: trillion-parameter capacity, 49B-parameter inference

In a dense model, essentially every parameter participates in every token computation. Large 4 works differently. Its parameters are distributed across many experts, and a router selects which experts process each token. Only about 49 billion parameters, under 5% of the total, are active at each step.

This separates model capacity from per-token computation. Mistral can scale how much the model can represent without making each inference step as expensive as running a dense trillion-parameter network.

Sparsity reduces compute, not memory. All 1.05 trillion parameters still have to be loaded to serve the model, which matters for self-hosting (covered below).

Architecture

Large capacity, selective computation

Input
Token
Router
Select experts
Expert A Active
Expert B Active
Expert C
Expert D
… other experts
Output
Next token

Illustrative view: for each token, the router activates only a small subset of experts. Large 4 holds 1.05 trillion parameters in total, but about 49 billion participate in any single inference step.

Vision: multimodal input, text output

The 1.6 billion-parameter vision encoder lets images enter the same model instead of going through a separate vision pipeline. A single workflow can take screenshots, diagrams, scanned documents, technical drawings or aerial imagery, reason over them, and then produce structured output or call tools. The output itself is always text.

Mistral highlights visual grounding as a particular strength: locating specific objects within an image. It also mentions converting technical drawings into CAD models.

A 1 million-token context window

Large 4 has four times the 256K context of Large 3 and Medium 3.5. That leaves room for large document collections, entire codebases and long-running agent sessions without aggressive chunking. The real benefit is less preprocessing when the relationships between distant pieces of information matter, not simply fitting more text into a prompt.

A large window does not make retrieval and context selection obsolete. Every unnecessary token adds latency and cost. Long-context recall should also be tested on your own data rather than assumed from the maximum window size.

Benchmarks: strong results, mostly vendor-reported

Mistral makes two headline claims. The first is that Large 4 is the best open-weight model from the US or Europe on aggregated benchmarks. The second is that it outperforms closed frontier models on visual grounding. These are the main preview results:

Preview benchmarks

Where Large 4 stands at launch

DeepSWE v1.1Mistral
62%
Long-horizon, agentic software engineering
GLM-5.3: 61% · DeepSeek V4 Pro: 57% in Mistral's comparison
FinchMistral
67%
Enterprise finance and accounting workflows
Tied with DeepSeek V4 Pro · Mistral Medium 3.5: 37%
Harvey Legal Agent BenchmarkMistral
15%
Task pass rate on legal agent tasks
Kimi K3: 12.9% on the public Vals.ai leaderboard
DIOR-RSVGMistral
73%
Visual grounding on remote-sensing imagery
GPT-6 Astra: 68% in Mistral's comparison
Dense200Mistral
42%
Dense visual grounding
Ahead of most general-purpose models compared
Artificial Analysis Intelligence IndexIndependent
38
Aggregated independent evaluation
Top open model outside China · GLM-5.3: 45 · MiMo-V2.6-Pro: 46

Preview checkpoint results. Mistral-reported scores have not yet been fully reproduced independently, and Mistral is continuing reinforcement learning before the final weights are released.

The results are competitive, but they need context. Most of the numbers come from Mistral's own materials and were produced on a preview checkpoint.

The independent picture is mixed:

  • Artificial Analysis: On its Intelligence Index, Large 4 scores 38. That puts it ahead of every other Western open-weight model, but still behind leading Chinese open models such as MiMo-V2.6-Pro, GLM-5.3 and Kimi K3.
  • Coding: On the live DeepSWE leaderboard, the best published configurations put GLM-5.3 and Kimi K3 around 69%, and top closed models around 74%. Large 4's 62% is a strong result, not an outright lead.

Expect these numbers to move once the final weights are out and outside evaluators can test them.

Where Mistral Large 4 fits in production

Not every application needs a trillion-parameter model. Large 4's real value is that it combines capabilities that would otherwise require several models or extra infrastructure.

Software engineering and coding agents. The 1M context helps when a task depends on relationships across many files, specifications and existing implementation details. Function calling and agent support let the model take part in workflows instead of just generating snippets. DeepSWE is also one of the benchmarks Mistral emphasizes most.

Cybersecurity. This is one of Mistral's most strategic targets. Its argument is that security teams need models they can control for dual-use but legitimate defensive tasks, such as code scanning and defensive testing, which closed providers may refuse. During the preview, Mistral says it monitors API traffic and works with cybersecurity partners privately.

Long-context document and knowledge workflows. Finance, legal and enterprise knowledge work often means processing large volumes of material and returning structured outputs or tool calls. The Finch and Harvey results point in this direction. For regulated or high-stakes use cases, however, model capability is not the same as domain-level reliability. These applications still need validation, access controls, monitoring and human oversight.

Multimodal and industrial workflows. Manufacturing environments mix technical documentation, drawings, structured data and domain-specific processes. Mistral specifically highlights:

  • technical drawings and CAD conversion
  • satellite and aerial imagery
  • semiconductor and chip design

Here, a single multimodal model can act as an interface between different forms of information. Each deployment still requires its own evaluation.

Pricing and the production trade-off

Mistral lists Large 4 at $0.68 per million input tokens, $0.07 per million cached input tokens and $2.09 per million output tokens.

Caching makes a large difference for applications that repeatedly send the same context. Agents that work against a stable system prompt, shared documentation or recurring reference material benefit most. For example, a 500,000-token prompt costs about $0.34 per request at the standard input rate, versus about $0.035 when it is served from cache. Mistral's pricing page also notes that batch processing halves the price for high-volume, non-urgent work.

Here is how Large 4 compares with the rest of Mistral's current lineup:

Model Architecture Parameters Context Input / 1M Output / 1M
Mistral Medium 3.5 Dense 128B 256K $1.50 $7.50
Mistral Large 3 Granular MoE 675B / 41B active 256K $0.50 $1.50

Pricing and specifications reflect Mistral's published model information and may change over time.

The comparison is less straightforward than "newest is best":

  • Against Large 3: Large 4 costs roughly 35–40% more per token. In return, it offers four times the context window and far more total capacity.
  • Against Medium 3.5: Large 4 is cheaper on both input (less than half the price) and output (about 72% less). For teams currently on Medium 3.5, it may be worth testing for cost reasons alone.

For production teams, the right metric is cost per successful task, not cost per million tokens. A pricier model can be cheaper overall if it needs fewer retries, produces fewer invalid structured responses or completes multi-step tasks more reliably.

Open weights, licensing and sovereign AI

Large 4's weights are scheduled for release on October 27. That follows a roughly three-week testing period with developers, cybersecurity leaders and government authorities.

Two details matter for teams planning to self-host:

  • License: The weights are expected under a custom Mistral license, not Apache 2.0 like Large 3. Review the terms before building commercial products on them.
  • Naming: "Open-weight" is the accurate description here, not "open source."

Mistral is also positioning Large 4 around sovereign AI. The model is deployable from Europe through Mistral's own cloud infrastructure, with zero-data-retention options. Once the weights are public, organizations will be able to run and customize it on infrastructure they control.

In enterprise settings, model selection is rarely about benchmark scores alone. Data governance, deployment options, infrastructure control and the ability to switch providers can weigh just as much as raw capability.

Self-hosting a model of this size is still a serious undertaking. At 8-bit precision, the weights alone take roughly 1 TB of GPU memory, and about 2 TB at 16-bit, before accounting for the KV cache needed for long contexts. That means multi-GPU, often multi-node, deployments. For most teams, the API will remain the practical way to use Large 4, at least initially.

The practical takeaway

Mistral Large 4 is best understood as a large-capacity, sparse, multimodal model built for serious production workloads, not as another larger chatbot. Each part of its design does a specific job:

  • the MoE architecture separates total capacity from active computation
  • the 1M context window makes large-context applications practical
  • the vision encoder brings images and text into one workflow
  • function calling, structured outputs, agents and batching connect it to real application infrastructure

It also arrives as a preview. The benchmarks are largely self-reported, the final checkpoint is still being tuned, and the weights are three weeks away. The right next step is not to adopt it everywhere. It is to benchmark it on the tasks that matter to you: your documents, codebases, images, tool calls and production constraints.

That is where a model like Large 4 can be properly evaluated. A unified layer such as Eden AI makes it practical to compare it side by side with the models you already use.

FAQ

Mistral Large 4, nicknamed "Le Chonk", is Mistral AI's flagship multimodal model, released in public preview on October 6, 2026. It uses a granular Mixture-of-Experts architecture with 1.05 trillion total parameters and 49 billion active parameters, alongside a 1.6 billion-parameter vision encoder and a 1 million-token context window.
Large 4 has 1.05 trillion parameters in total, but its sparse MoE architecture activates only about 49 billion of them for each token. This gives the model very large overall capacity without the per-token computation of a dense trillion-parameter model. All parameters still need to be loaded in memory to serve it.
Large 4 is an open-weight model rather than open source. Its weights are scheduled for release on October 27, 2026, under a custom Mistral license, unlike Mistral Large 3, which was released under Apache 2.0. Until then, the model is available through the Mistral API.
Mistral Large 4 supports a 1 million-token context window, four times the 256K of Mistral Large 3 and Medium 3.5. This makes it suitable for large documents, extensive codebases and long-running agent workflows, although retrieval and caching remain important to control latency and cost.
Yes, on the input side. A 1.6 billion-parameter vision encoder lets Large 4 process images alongside text, but the model generates text only. Mistral highlights visual grounding, satellite imagery and technical drawings as particular strengths.
Mistral reports 62% on DeepSWE v1.1, 67% on the Finch finance benchmark, 15% on Harvey's Legal Agent Benchmark and 73% on DIOR-RSVG visual grounding. Independently, it scores 38 on the Artificial Analysis Intelligence Index, the highest among open models from outside China but behind leading Chinese open models such as GLM-5.3 and MiMo-V2.6-Pro.
Mistral lists Large 4 at $0.68 per million input tokens, $0.07 per million cached input tokens and $2.09 per million output tokens. That is slightly more than Large 3 but significantly less than Medium 3.5. Pricing can change, so verify the current rates before deployment.
Once the weights are released on October 27, 2026, yes, subject to the license terms. Hardware needs are substantial: the weights alone require roughly 1 TB of GPU memory at 8-bit precision, which means multi-GPU or multi-node deployments. For most teams, the API is the practical starting point.
Large 4 is particularly relevant for coding and agentic workflows, cybersecurity, long-context document analysis, finance and legal work, and multimodal industrial use cases such as technical drawings and satellite imagery.
Eden AI provides a unified API for accessing and comparing models from multiple providers. Availability of a specific model should be verified in the Eden AI model catalog, as integrations are added and updated regularly.
Not automatically. Benchmark Large 4 on representative production workloads and compare quality, latency, tool-call reliability, context handling and cost per successful task against the models already in use, keeping in mind that the current version is a preview.

Similar articles

AI NEWS
Generative AI
Quasar 438B vs Mistral Large 4 vs Kolibri: Best European LLM in 2026
10/7/2026
·
Written byClément Moreau
AI NEWS
Generative AI
Aleph Alpha Kolibri: Open-Weight German-English LLM for Sovereign AI
10/5/2026
·
Written byClément Moreau
AI NEWS
Generative AI
Google Gemini 4 Argon: Cost and Performance Benchmarks for Production AI Systems
10/1/2026
·
Written byClément Moreau
let’s start

Start building with Eden AI

A single interface to integrate the best AI technologies into your products.