Mistral Large 4: a cyber model with an open-weights deadline
On 6 October Mistral opened a public preview of Mistral Large 4, which it calls "le Chonk", and made cybersecurity its headline capability. The weights are due by the end of October. From that point, the controls around the model belong to whoever runs it.
What Mistral announced
Mistral describes Large 4 as "a 1 trillion-parameter natively multimodal model with 49 billion active parameters". Its documentation lists 1.05 trillion total and 52 billion active; the Hugging Face page reconciles the two: 49 billion per token, "52 billion including embeddings and output layers".
Mistral calls it "one of the world's strongest AI models for cybersecurity" and reports that on one Artificial Analysis Cyber Index test, which "asks a model to reproduce a real vulnerability in open-source software and then patch it" (labelled CyberGym-E2E on Mistral's chart), it scores 82%, "the highest of any model". It also "solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions". Mistral says several closed models score near zero on the first test "because they refuse to perform the task". These are Mistral's own figures; we have not reproduced them.
That puts Large 4 squarely up against Z.ai's GLM-5.3, the Chinese open model that has set the pace for cyber since August. Zhipu reported 84.5% on CyberGym for GLM-5.3, ahead of Anthropic's Mythos 5 (83.8%) and OpenAI's GPT-5.6 Sol (83.6%), and said that "cyber capability developed faster than we expected" (CSO Online, Axios). Mistral's 82% is on an end-to-end variant of the same family of tests, so the two numbers are not directly comparable. The two launches share a shape: a cyber-capable frontier model, a short staged preview, then open weights. Zhipu also planned to publish GLM-5.3's weights about two weeks after launch, following safety evaluation. Europe now has its own contender in that race, with the same open-weights question attached.
Mistral also reports a cyber-prompt refusal rate "higher than all OSS models". The launch post lists $1.36 per million input tokens and $4.18 per million output; the model page currently shows half those rates beside them.
A staged release, then open weights
Mistral's wording on the release is precise: "We will release the weights by the end of the month. Until then, we are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities." Reuters reported a public release on 27 October; the Hugging Face page estimates 31 October.
That deserves credit. A moderated preview, weeks of external red-teaming and published refusal results are more than many open-weight releases get. Mistral's case is also serious: "provider-level refusals can block legitimate vulnerability research and incident response".
We found no model card or system card as of 7 October. Mistral's pages do not state the licence; The New Stack reports a custom licence rather than Large 3's Apache 2.0.
"Tried to escape": what the source actually says
A widely shared headline says the model "tried to escape its test environment". Mistral's launch post, documentation and Hugging Face page do not say this. It comes from a Reuters interview with Mistral's vice president of science, Pierre Stock. In Reuters' text: "Stock told Reuters the model had tried to go beyond its testing environment, but that this was expected and the company was able to prevent it." The New Stack, whose headline used the word "escape", adds that the company "contained it using software".
So a senior Mistral executive said it on the record. But it is one sentence. Nothing published says what the model attempted, what stopped it, or how the episode was found, and "escape" is the headline's word, not Mistral's. We treat the substance as unverified until Mistral documents it. Our analysis of containment failures at two labs showed why detail matters: the useful questions are how an attempt was detected, and how quickly.
After the weights ship
Mistral's staged controls belong to its service. A moderated API and vetted access do not travel with downloaded weights; according to The New Stack, Stock told Journal du Net that replicated weights cannot easily be revoked. Mistral is making that trade openly, and for many defenders it is the point.
It does move the question. Defenders will run Large 4 inside agent runtimes that read code, execute commands and touch networks. Safety then depends on what the runtime enforces and records, not on the provider's moderation. Our research question applies directly: when an AI agent acts, can anyone reliably reconstruct what happened? What agent harnesses let you record varies widely.
The model is half the story. Mistral's coding agent, Vibe, was added to the AgenticBench security leaderboard on 7 October. Vibe 2.26.0 scored 50 of 55 with no fails and no missing controls, the joint top score, level with Codex 0.159.3 and Grok 1.0.44, and listed third on the tie-break. Its S08 cell is marked "reported", not scored, as for the other top harnesses. The leaderboard runs every harness against the same model, not Mistral's, so it says nothing about Large 4 itself, but Vibe held up well on the controls we test. An earlier Vibe batch was discarded because our own adapter's S11 file denylist also matched the control file; we fixed that in rig commit 6b42a60 before the full rerun. AgenticBench takes no money from the agents it tests.
A small spot check
On 7 October we sent mistral-large-4 six authorised defensive tasks through the API: an authentication-bug review, reading a system-call trace, a harmless permission probe, explaining command injection, drafting a vendor disclosure and a sandbox-config review. With no output-token cap it completed all six with no refusals, in 14 to 26 seconds per call. This was a six-prompt, single-run spot check, not a benchmark or a security evaluation.
What we can and cannot confirm
| Claim | Confidence | Basis |
|---|---|---|
| Preview 6 October; weights by end of October; vetted reduced-moderation access meanwhile | High | Mistral launch post |
| Weights date of 27 October | Moderate | Reuters; Hugging Face shows 31 October |
| About 1 trillion total, 49 to 52 billion active parameters | High | Mistral post, docs, Hugging Face |
| 82% on the reproduce-and-patch test; 93% of Cybench's 40 challenges | Moderate | Mistral's figures, not reproduced |
| A Mistral executive said the model tried to go beyond its testing environment | High | Reuters interview |
| What that attempt involved and how it was contained | Unknown | No published detail |
| Custom licence rather than Apache 2.0 | Low | One secondary report |
What would help
- A system card before the weights. Publish the red-teaming findings and the testing-environment episode while Mistral's controls still apply.
- Treat the runtime as the control point. Decide what the agent may do, and record what it did, before the record is needed.
Disclosure: Agentic Thinking is an independent research lab. We steward the open-source AgentHook standard, maintain the open-source HookBus project and run AgenticBench, which takes no money from the agents it tests. Agentic Thinking has also developed agent governance software, AgentProtect, which is not sold and not currently offered. These interests may overlap with the runtime recording issues discussed above.
Sources
- CSO Online, "Zhipu says new coding AI developed advanced cyber skills faster than expected", August 2026
- Axios, "A Chinese lab's new model is nearly as good at hacking as U.S. AI", 14 August 2026
- Mistral, "Introducing Mistral Large 4", 6 October 2026
- Mistral documentation, "Mistral Large 4" model page
- Hugging Face, mistralai/Mistral-Large-4.0-1T05-A52B (upcoming release page)
- Reuters, "France's Mistral launches AI model it says outperforms some Chinese rivals", 6 October 2026 (text read via syndicated copy at KFGO)
- The New Stack, "Mistral's new AI tried to escape its test environment. In three weeks, anyone can download it", 6 October 2026
- Mistral, "Introducing: Devstral 2 and Mistral Vibe CLI", 9 December 2025
- AgenticBench security leaderboard v0.4, 7 October 2026
Related: Two labs, one failure · What agent harnesses record, and what they send home · Our research
Collaborate with us →Agentic Thinking. We test what AI agents really do, and investigate when it goes wrong.