Mistral Large 4 'Le Chonk': A 1T-Parameter Open-Weight Model From Europe That Beats Claude Opus 5 on Cyber Defense

Mistral's ML4 public preview ships at 1T total / 49B active params — natively multimodal, trained on 3,800 NVIDIA Grace Blackwell GPUs in Europe. Open weights end of month. 82% on the reproduce-and-patch cyber benchmark (Claude Opus 5 and GPT-6 Astra score near zero on the same test).

Mistral launched a public preview of Mistral Large 4 (“ML4”, also “le Chonk”) on October 6, 2026 — its largest model to date and the most credible claim yet from a non-Chinese lab to the open-weight frontier. The headline: 1 trillion total parameters, 49 billion active per token, natively multimodal, trained on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters. Open weights drop end of month. Public preview is live on Mistral Studio today.

The numbers and the claims:

  • 1T total parameters / 49B active per token — sparse Mixture-of-Experts, natively so (not retrofitted). Source: Mistral’s own announcement.
  • 3,800 NVIDIA Grace Blackwell GPUs, trained from scratch over two months in Mistral’s own European datacenters. CNBC reports 4,000 GPUs, Mistral says 3,800 — both numbers come from official sources and the discrepancy is not material.
  • Natively multimodal — text + image + (per Mistral) video understanding. A 1.6B-parameter vision encoder is part of the architecture.
  • Public preview on Mistral Studio now, weights by end of October. The preview is gated to vetted cybersecurity partners and state authorities under reduced moderation; broader access follows the weight release.
  • 82% on the Artificial Analysis Cyber Index “reproduce a real vulnerability then patch it” test — the highest of any model. Cybench score: 93% (40 CTF-style challenges).
  • 61.7% DeepSWE v1.1, 59.4% SWE-Atlas-QnA, 28.3% Terminal-Bench 4, Coding Agent Index 49.8% — ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.
  • Surge AI blind human eval: 3.74/5 (2nd of 5, behind only Claude Opus 5 at 4.22; ahead of Kimi K3 3.59, GLM-5.3 3.60, GLM-5.2 3.40).
  • Visual grounding on Dense 200: 42% vs GPT-6-Astra’s 41% — Mistral explicitly claims a closed-frontier beat here.
  • SciCode-Verified SOTA among open-weight models. Generates a full Hartree–Fock simulation in one shot.
  • AA-Briefcase 1,393 Elo — the long-horizon knowledge-work benchmark, ahead of DeepSeek V4 Pro.
  • Trained on data spanning 160+ languages, every official EU language included.

The shape of the model

ML4 is a 1T-total / 49B-active sparse MoE. That active footprint puts per-token inference cost in the same neighborhood as a dense ~50B-class model, even though the total capacity is twenty times larger. The architecture choice is the same bet DeepSeek, Qwen, and Reflection (Beam, shipped Monday) made: keep total parameters frontier-class, keep per-token compute modest, route via sparse expert selection. ML4 is roughly 2× the total parameter count and 2× the active parameter count of Reflection Beam (501B/23B, which we covered yesterday).

The training run is the European-sovereignty angle. 3,800 Grace Blackwell GPUs over two months, in datacenters Mistral operates itself, in Europe, under European law. The company raised a €3B Series D in September at a €21B valuation — the round was the largest AI Series D in Europe this year and cemented Mistral’s status as “Europe’s only answer to OpenAI and Anthropic.” That capital and infrastructure is what makes a 1T-parameter open-weight run possible for a non-US, non-Chinese lab.

The multimodal architecture is genuinely multimodal, not bolted-on: the 1.6B-parameter vision encoder is part of the published architecture rather than a separately trained adapter. Mistral is demoing visual grounding on gigapixel satellite imagery, engineering drawings, and dense documents — workloads that matter for disaster response, manufacturing inspection, and the long-tail enterprise use cases that pure-text models can’t address.

The cyber defense story

The single most important claim in the launch is the cyber benchmark. On the Artificial Analysis Cyber Index — a third-party benchmark that asks a model to find and fix a real vulnerability in open-source software — ML4 scored 82%, the highest of any model tested. Claude Opus 5.5 and GPT-6 Astra score near zero on the same test because they refuse to perform the task. The benchmark isn’t measuring raw capability, it’s measuring whether the model will do the work.

The reasoning is straightforward: defensive security work requires proving a flaw is real, and that work gets blocked by safety filters on closed models. As threat actors increasingly jailbreak those same closed models for offensive use, defenders need systems that can match those capabilities without the refusals. ML4 has those capabilities, the open weights mean it can run on private cloud or on-prem, and the lack of provider-level refusals means it can do the work defenders need.

The earlier Hugging Face incident — where the platform turned to Chinese open-weight GLM 5.2 to defend against rogue OpenAI agents because US closed systems refused to act — is the precedent. The closed-model guardrails mistook defensive action for offensive action. ML4 is the Western open-weight answer to that problem.

What “open-weight outside China” actually means

The headline framing — “strongest open-weight model developed outside China by a substantial margin” — is doing real work. The credible open-weight frontier has been Chinese-dominated for the last 18 months: DeepSeek V4 Pro, Qwen3.8 Max, GLM-5.2 / GLM-5.3, Kimi K3. Western open-weight has lagged.

ML4’s claim is that it’s the first non-Chinese model to credibly rival them. The benchmark scorecard is the evidence:

  • Coding Agent Index 49.8% — ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.
  • Surge AI human eval 3.74 — ahead of GLM-5.3 (3.60) and GLM-5.2 (3.40).
  • Cybench 93% — one of the highest scores reported for any open-weight model.

Whether ML4 holds those benchmarks in third-party reproduction is the empirical question the weight release will answer. But the numbers Mistral is reporting — direct, with named benchmarks and named comparators — are the kind of claims that can be checked, which is more than can be said for most frontier launches.

Reflection Beam (Monday) and Mistral Large 4 (Tuesday) are the two punches of an open-weight week. Reflection made the GLM-5.2 parity claim with a 501B/23B MoE. Mistral is making the “stronger than DeepSeek V4 Pro and Qwen3 Max” claim with a 1T/49B MoE. Both are running on Grace Blackwell hardware, both have weights dropping this month, both are explicitly framed as Western answers to the Chinese open-weight frontier. The cadence is not coincidental.

The valuation context

Mistral is now the most valuable AI startup in Europe. The September Series D (€3B raised, Samsung as the lead at the institutional level, €21B post-money valuation) was the company’s biggest raise to date. The capital is going into training capacity — Guillaume Lample, co-founder and chief scientist, explicitly tied the ML4 launch to “scale up our training capacity, following our Series D fundraise.”

That capital structure matters for the open-weight question. Most frontier labs are venture-funded to build a closed API business. Mistral is venture-funded to build an open-weight business — the model is the moat, not the API. Whether that economic model holds at the 1T-parameter scale is the open question. The Series D is the bet that it does.

The sovereignty frame

Mistral is leaning hard on the European-sovereignty angle. The ML4 launch page explicitly says “Forged in Europe. Built for AI sovereignty.” The model is trained in Europe, served from Europe, available under a European deployment that operates “independently of other digital service providers and under European law.”

The 160-language training corpus — every official EU language, plus a long tail — is the same story told through a different lens. Most frontier models are English-first with multilingual scaling as a stretch goal. ML4’s training data was “significantly multilingual” by design.

For European enterprise customers and EU-regulated industries, that’s a procurement-relevant differentiator. For US customers, it’s a marketing differentiator. For Chinese customers, it’s not available at all — and that’s the implicit third front in the open-weight race.

What to verify when weights drop

The preview is gated. The real test is the open weights release at end of month. Things to watch:

  • Reproduction of the benchmark claims. 82% on the reproduce-and-patch test, 93% on Cybench, 61.7% on DeepSWE v1.1 — these are the headline numbers, and third-party reproduction is the only thing that matters.
  • Active parameter count vs published total. 49B active on 1T total is the architecture claim. Independent benchmarking of inference cost vs a dense 50B-class model is the verification.
  • Multimodal architecture details. The 1.6B-parameter vision encoder is the published component; the actual integration with the MoE backbone is the engineering question.
  • Cyber defense vs refusal behavior. The 82% Cyber Index score assumes the model will do the work. Closed models score near zero because they refuse. ML4’s reduced-moderation preview is the explicit hedge — the open weights will need to reproduce that behavior without the gating.
  • European infrastructure claims. The 3,800 Grace Blackwell count, the European data residency, the EU-law service region. Procurement and compliance teams will care about all three.

The state of the open-weight frontier

October 2026 is the densest week the open-weight frontier has had in a year. Reflection Beam on Monday (501B/23B Apache 2.0). Mistral Large 4 on Tuesday (1T/49B, weights end of month). Both aimed at the same target — Chinese open-weight dominance — with different size bets and different license terms.

The question isn’t whether Western open-weight catches up to Chinese open-weight. The benchmarks already say it has. The question is whether the catch-up holds in third-party reproduction, and whether the European infrastructure frame matters to enterprise procurement at the 1T-parameter scale.

We’ll know more in three weeks when ML4’s weights drop.

Sources