Mistral reproduces the exploit, the others refuse
Mistral's open-weight Large 4 tops a vulnerability-reproduction test where closed models refuse, and ships with refusals any self-hoster can tune away.
Mistral Large 4 scores 82% on a test that asks a model to reproduce a real vulnerability in open-source software and then patch it. On Mistral’s account that is the highest any model has posted on that test, part of the independent Artificial Analysis Cyber Index. Several frontier closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform it.
That gap is the pitch for Mistral Large 4, unofficially ML4, officially “le Chonk.” Mistral is releasing it as an open-weight model, with weights due by the end of the month, and leaning hard on the argument that provider-side refusals break legitimate security work. Proving a flaw is real is often the first step in fixing it, and a safety filter that blocks exploit reproduction blocks that step. Losing a capability mid-incident, Mistral notes, is its own security risk.
The model is a one-trillion-parameter mixture-of-experts with 49 billion active parameters, natively multimodal, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters. The active-parameter count is what matters for anyone planning to run it: at 49B active it sits within range of a well-equipped on-prem deployment rather than demanding a hyperscaler. Preview API pricing is $1.36 per million input tokens and $4.18 per million output, served on the same European infrastructure.
On cyber specifically, Mistral places ML4 in the top five globally on the Cyber Index and well ahead of any open-weight model developed outside China. It solves 93% of Cybench, a set of 40 exercises drawn from security competitions. Mistral says the model also handles malware analysis, vulnerability prioritisation, and detection-rule writing without being explicitly trained for those tasks, and that it will run on private cloud or on-premise for teams that need a self-hosted security assistant.
The open weights change the calculation on both sides. Mistral reports that ML4 refuses malicious cyber prompts more often than any other open model it tested, measured against JailbreakBench, StrongREJECT, and AgentHarm, that it saturates Mistral’s indirect-prompt-injection benchmarks, and that it resists 93.3% of attacks on Lakera’s public B3 benchmark. Those numbers describe the checkpoint Mistral ships. On an open-weight model, alignment is a property of the weights you downloaded, and refusal behaviour can be fine-tuned out by anyone with the weights and modest compute. The refusal rate that reads as a safety feature in the release notes is something an attacker running the model locally can strip.
Mistral is explicit that the same model runs with reduced moderation for a vetted few. During the preview it is red-teaming with “cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.” The capability ceiling of ML4 is whatever the model does once moderation is turned down, not the refusal-laden public preview.
For defenders this cuts a specific way. A self-hostable model that reproduces vulnerabilities, triages them, and writes detections, running under your own policies with no provider able to pull access mid-incident, is genuinely useful to a security team with the discipline to sandbox it. The same properties make it useful to the people you are defending against, who get the same weights and can fine-tune the refusals out just as easily. If you are evaluating ML4 for security operations, judge it on the cyber capability underneath those refusals, since the refusal rate is the first thing anyone self-hosting can tune away.
Weights drop by the end of the month, along with architecture details, more benchmarks, and Mistral’s post-training methodology. The preview API is live now on Mistral Studio.
Keep Reading
ai-hardwareAI agents designed the chip that runs their inference
openTPU is an AI-written inference accelerator whose FPGA card reproduces its reference simulator's tokens bit for bit, checked by an executable spec.
ai-agentsMuse hands root to anyone
Meta's Muse AI agent grants root to anyone claiming to be an agent and uploads private messages despite permission settings, after a rushed launch.
ai-agentsStrands built a 2B model that never writes text
Strands Decider 2B is an open-source 2B decision model that scores choices in ~115ms, cheap enough to run guardrail checks before every agent tool call.
Latest on the Wire
Full wire →- Advantest confirms ransomware attack exposed personal dataBleepingComputer
- Anthropic Expands AI Model Access for Cybersecurity TeamsThe Hacker News
- Anthropic Releases Claude Haiku 5.5: Faster, Cheaper, More CapableHacker News
- AnyPS5 Ports PS5 Games to PC Without EmulationHacker News
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.