RC RANDOM CHAOS

Mistral reproduces the exploit, the others refuse

Mistral's open-weight Large 4 tops a vulnerability-reproduction test where closed models refuse, and ships with refusals any self-hoster can tune away.

· 3 min read
Mistral reproduces the exploit, the others refuse

Mistral Large 4 scores 82% on a test that asks a model to reproduce a real vulnerability in open-source software and then patch it. On Mistral’s account that is the highest any model has posted on that test, part of the independent Artificial Analysis Cyber Index. Several frontier closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform it.

That gap is the pitch for Mistral Large 4, unofficially ML4, officially “le Chonk.” Mistral is releasing it as an open-weight model, with weights due by the end of the month, and leaning hard on the argument that provider-side refusals break legitimate security work. Proving a flaw is real is often the first step in fixing it, and a safety filter that blocks exploit reproduction blocks that step. Losing a capability mid-incident, Mistral notes, is its own security risk.

The model is a one-trillion-parameter mixture-of-experts with 49 billion active parameters, natively multimodal, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters. The active-parameter count is what matters for anyone planning to run it: at 49B active it sits within range of a well-equipped on-prem deployment rather than demanding a hyperscaler. Preview API pricing is $1.36 per million input tokens and $4.18 per million output, served on the same European infrastructure.

On cyber specifically, Mistral places ML4 in the top five globally on the Cyber Index and well ahead of any open-weight model developed outside China. It solves 93% of Cybench, a set of 40 exercises drawn from security competitions. Mistral says the model also handles malware analysis, vulnerability prioritisation, and detection-rule writing without being explicitly trained for those tasks, and that it will run on private cloud or on-premise for teams that need a self-hosted security assistant.

The open weights change the calculation on both sides. Mistral reports that ML4 refuses malicious cyber prompts more often than any other open model it tested, measured against JailbreakBench, StrongREJECT, and AgentHarm, that it saturates Mistral’s indirect-prompt-injection benchmarks, and that it resists 93.3% of attacks on Lakera’s public B3 benchmark. Those numbers describe the checkpoint Mistral ships. On an open-weight model, alignment is a property of the weights you downloaded, and refusal behaviour can be fine-tuned out by anyone with the weights and modest compute. The refusal rate that reads as a safety feature in the release notes is something an attacker running the model locally can strip.

Mistral is explicit that the same model runs with reduced moderation for a vetted few. During the preview it is red-teaming with “cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.” The capability ceiling of ML4 is whatever the model does once moderation is turned down, not the refusal-laden public preview.

For defenders this cuts a specific way. A self-hostable model that reproduces vulnerabilities, triages them, and writes detections, running under your own policies with no provider able to pull access mid-incident, is genuinely useful to a security team with the discipline to sandbox it. The same properties make it useful to the people you are defending against, who get the same weights and can fine-tune the refusals out just as easily. If you are evaluating ML4 for security operations, judge it on the cyber capability underneath those refusals, since the refusal rate is the first thing anyone self-hosting can tune away.

Weights drop by the end of the month, along with architecture details, more benchmarks, and Mistral’s post-training methodology. The preview API is live now on Mistral Studio.

Share

Keep Reading

Latest on the Wire

Full wire →

New signal daily · RSS

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.