The headline is real. So is the fine print.
Mistral released Large 4 on 6 October 2026. A 1-trillion-parameter mixture-of-experts model, multimodal input, native fluency across 160+ languages, and open weights planned before the end of October. The announced context window is 512k tokens — Artificial Analysis reports 520k tokens for the preview; the two sources disagree, and the discrepancy has not been resolved publicly. Mistral’s announcement framed the model as the strongest open-weight option for cybersecurity, finance, and manufacturing. That framing holds up in some places. In others, it needs a careful read.
Large 4 is a genuine step forward for European AI — but the benchmark picture is more textured than a launch blog suggests. Operators choosing a model for production work deserve the unflattened version.
Where the scores are cleanest
Cybersecurity is the strongest ground. Mistral reports 93% on Cybench and 82% on the Artificial Analysis Cyber Index’s reproduce-and-patch test, which it says is the highest of any model at launch. Several closed models score near zero on that test because they refuse the task entirely. For security teams that need a model willing to engage with adversarial material under controlled conditions, open weights plus self-hosting is the operational point, not a footnote.
The safety profile reinforces this. On Lakera’s public B3 AI Security Benchmark, Large 4 resists 93.3% of attacks — Mistral reports no higher score among competitors at launch. On the KORA Benchmark for responsible engagement, it scores 1.691 out of a maximum 2.0, the highest Mistral has measured among open-source models. Those figures speak directly to what a security-conscious operator needs to assess before deploying a model on internal data.
Knowledge work is the other area with meaningful external validation. Mistral commissioned third-party evaluator vals.ai, which found the model exceeds GPT-6-Astra on both legal and financial tasks. On Harvey AI’s Legal Agent Benchmark, VentureBeat notes that Vals.ai’s public leaderboard shows Kimi K3 at 12.92%, MiMo V2.6 Pro at 10.83%, and GLM-5.3 at 8.33% — matching the rounded competitor figures in Mistral’s chart. If Large 4’s reported 15% task-pass rate was produced under the same methodology, it leads those open-weight rivals. That conditional matters: it is an external cross-check, not an independent replication.
Where the picture gets complicated
Coding benchmarks are solid but require context. Large 4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas Codebase QnA, and 28.3% on Terminal-Bench 4.0, for 49.8 on the Coding Agent Index — per AI Release Tracker. Artificial Analysis ran those evaluations before the harness went public. Benchmark configuration matters: the DeepSWE result for DeepSeek V4 Pro 0813 shifts depending on which agent harness is used, as VentureBeat’s analysis notes. Direct comparisons require matching methodology, not just matching benchmark names.
The broader intelligence index tells a more grounded story. Artificial Analysis scores Large 4 Preview at 38.4 on its Intelligence Index — a composite of 10 evaluations covering reasoning, knowledge, mathematics, and coding. That is a large jump over Mistral Medium 3.5 (14 points) and Large 3 (9 points). But among open models, it currently ranks eighth. All seven models ahead of it are Chinese: Xiaomi’s MiMo-V2.6-Pro leads at 46.3, followed by Z.ai’s GLM-5.3 at 44.8 and Moonshot AI’s Kimi K3 at 43.6.
Mistral’s claim that Large 4 outperforms every open model from the United States and Europe holds up. Trending Topics reports the previous strongest non-Chinese open model was Motif 3 from South Korea at 33.6 points; Large 4 at 38.4 clears that bar. But “best open model from the US and Europe” and “best open model” are different claims. Conflating them in a procurement decision is a real risk.
The counterpoint worth taking seriously
The strongest objection to deploying Large 4 today is that the weights are not yet public. Artificial Analysis currently lists the preview as proprietary. Mistral has stated weights will arrive by end of October 2026 — that is a plan, not a release. Until the weights ship and the community can reproduce results under consistent harness configurations, some benchmark comparisons remain provisional. Operators building on open-weight assumptions should treat the preview API as exactly that.
Pricing adds another variable. At $1.36 per million input tokens and $4.18 per million output tokens, Large 4 is cheaper than GLM-5.3 and Kimi K3 on a cost-per-task basis. Artificial Analysis notes a 90% cache discount, but that applies only to workloads with high prompt reuse — not to diverse agentic tasks, where the full per-token rate applies. The cost-per-task figures in the Trending Topics comparison table show DeepSeek V4 Pro 0813 at $0.67 per task versus Large 4 at $1.13 per task — a meaningful gap at volume, though per-task and per-token costs are not directly interchangeable and your workload mix will determine the real difference.
What this means for operators
The clearest case for Large 4 today is a security-sensitive professional workflow: legal document analysis, financial spreadsheet automation, or internal threat research. Those are the three areas where third-party evaluations — vals.ai (commissioned by Mistral), Harvey AI’s benchmark, Lakera’s B3 — provide the strongest external signal. The cybersecurity story is genuinely differentiated: a model that scores 82% on reproduce-and-patch while remaining open-weight and self-hostable is unusual. The combination of that score with 93.3% attack resistance on B3 is not matched by any competitor in the source material.
General-purpose coding agents are a weaker fit right now. The Terminal-Bench score of 28.3% reflects real limits in long-horizon CLI tasks. For cost-sensitive inference at scale, the cost-per-task gap between Large 4 and DeepSeek V4 Pro is worth modelling before committing to a provider — particularly if your workload does not benefit from prompt caching.
Large 4 is also explicitly positioned as a foundation. Mistral says it will serve as the base for a new generation of specialised models built for specific industries. Operators in finance, legal, or manufacturing should track fine-tuned derivatives as they arrive, rather than evaluating only the base model.
One concrete next step: before the weights drop, run Large 4 through your own eval harness on the specific task type you care about — using the same prompt structure, tool configuration, and agent harness you would use in production. Pay particular attention to harness version if you are comparing against published DeepSWE or Terminal-Bench figures; configuration differences have already produced measurable score shifts across models in this cohort. Vendor benchmarks establish a baseline. Your workload is the actual test.
— Eagentix
Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.
Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.
Sources
- Source: Mistral Large 4 — Benchmarks, Specs & Release Date (
- Source: Introducing Mistral Large 4 | Mistral (
- Source: Mistral debuts Large 4 ‘Le Chonk’, a 1-trillion parameter text output model with high benchmarks planned for open weights release | VentureBeat (
- Source: Mistral Large 4 Preview – Intelligence, Performance & Price Analysis | Artificial Analysis (
- Source: Mistral Large 4 on Artificial Analysis: Eighth Among Open Models | Trending Topics (
- Source: Mistral Large 4 benchmarks : r/LocalLLaMA (
- Source: Mistral Docs (
