Addis PulseStudio

Mistral Ships le Chonk: A Trillion-Parameter Model Built to Patch Vulnerabilities the Frontier Refuses to Touch

Mistral Large 4 is in public preview now, open weights land by month's end, and the cyber benchmark is where the story actually lives.

4 min read843 words

What happened

Mistral opened a public preview of Mistral Large 4 β€” officially "le Chonk," informally ML4 β€” on 6 October 2026, with the API live on Mistral Studio and open weights scheduled by the end of the month. The model carries one trillion total parameters with 49 billion active, is natively multimodal, and is Mistral's largest release to date.

Context

Mistral Large 3, released in December 2025, scored 9 on the Artificial Analysis general benchmark. ML4 scores 38 on the same test, just behind DeepSeek 4.1 Flash, a 552B model. That jump in ten months is the through-line: Mistral has been closing the gap with the frontier while keeping training and inference on European hardware. The 160-language training corpus, including every official EU language, and the end-to-end European deployment under European law are not afterthoughts. They are the product.

How it works

ML4 is a mixture-of-experts architecture: one trillion total parameters, but only 49 billion are active per forward pass, which is what makes serving a trillion-parameter model on a single inference stack plausible. It is natively multimodal, meaning image and text inputs share a single representation rather than a bolted-on encoder. Training ran on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, and the preview is served on that same infrastructure.

The API exposes two reasoning levels: "none" and "high," which control how much internal deliberation the model performs before emitting a response. In Simon Willison's SVG-generation test, the "high" setting produced 2,717 output tokens versus 3,275 at "none," suggesting the extra compute buys tighter output rather than a longer one. Customisation and RL fine-tuning run through Mistral Forge, the same pipeline Mistral uses internally.

Our read

The number everyone will quote is the 82% on the Artificial Analysis Cyber Index, where ML4 reproduces and patches a real vulnerability in open-source software. Mistral calls it the highest score of any model tested. The more interesting data point sits next to it: Claude Opus 5.5 and GPT-6 Astra score near zero on the same task because they refuse it. The gap is not capability. It is a policy boundary. Mistral is demonstrating that a model can be pointed at a live CVE and produce a working patch, and that the refusal behaviour of the US frontier labs is a design choice, not a technical limit.

The red-teaming regime they describe β€” cybersecurity leaders, vetted partners, and state authorities accessing the model with reduced moderation during preview β€” reads as Mistral pre-emptively answering the question of who controls this once the weights are out. The answer, so far, is a curated group operating under European oversight.

What the announcement does not resolve is general-purpose reasoning. Willison calls ML4 "maybe about 6 months behind the frontier," and the Surge AI blind evaluation agrees: second of five at 3.74 out of 5, behind Claude Opus 5's 4.22. The 49.8% Coding Agent Index, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max, is solid but not a frontier number. ML4 is a specialist with a very specific edge, wrapped in a generalist's framing.

What this changes

The preview API is live on Mistral Studio today. If you are building on Mistral's stack and need the multimodal path or a larger parameter base, you can point existing API calls at the new endpoint and start benchmarking. The two reasoning levels are the first thing to test: "none" for latency-sensitive work, "high" for tasks where output quality matters more than token count.

If your deployment requirement is European data sovereignty, ML4 is the first trillion-parameter model you can run end-to-end under European law. That is not a small thing for a studio serving EU clients.

What is not available yet: the open weights. They land by the end of October. Until then, you are locked to Mistral's API for inference, and the Forge customisation pipeline is the only path to a tuned variant. Nothing changes for a local-stack workflow this week.

License

The sources do not state a licence for the open weights scheduled for release by the end of October 2026. Check the model card when the release lands before building anything commercial on it.

Key takeaways

  • ML4 is a 1-trillion-parameter MoE with 49 billion active, natively multimodal, in public preview now with open weights due by end of October.
  • The 82% AA Cyber Index score is the standout result; US frontier models score near zero on the same task due to refusal, not capability.
  • On general reasoning, ML4 trails the frontier by an estimated six months; it is a specialist, not a generalist replacement.
  • Two reasoning levels ("none" and "high") let you trade latency against output quality on a per-request basis.
  • European data sovereignty is the structural differentiator: training, inference, and legal jurisdiction all sit in the EU.

Sources

  1. Introducing Mistral Large 4 β€” tier 1
  2. Mistral Large 4 β€” tier 2
  3. Introducing Mistral Large 4: Le chonk β€” tier 2
  4. Mistral Large 4 β€” tier 3
mistralcybersecurityllmopen weights

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
4
cluster pair
clef:27b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20261007T162720Z