Anthropic, OpenAI, and DeepMind grade their own models' danger
AI labs publishing dangerous-capability evals turns safety disclosure into marketing, and the self-graded scorecards drift toward a race to the bottom.
Anthropic, OpenAI, and Google DeepMind now publish documents rating their own models on how much help they give with building bioweapons, running cyberattacks, and operating without a human in the loop. Anthropic uses AI Safety Levels. OpenAI grades models low, medium, high, or critical across cybersecurity, CBRN, persuasion, and autonomy. Google DeepMind tracks “critical capability levels.” These are safety artifacts. Read from a boardroom, they are also capability brochures. “Our model is dangerous” and “our model is powerful” turn out to be the same sentence in two fonts.
That double meaning is the problem worth watching, and it is a systems problem, not a villain problem. No lab has to act in bad faith for the incentives to pull safety standards downward. The structure does the pulling on its own.
A warning label that reads as a sales pitch
When a lab reports that its newest model gives “meaningful uplift” to someone attempting a dangerous task, two audiences hear two things. Regulators and safety researchers hear a risk disclosure. Investors, enterprise buyers, and competitors hear a benchmark: this model is closer to the frontier than the last one. Danger is a proxy for capability, and capability is what the market pays for.
Watch what happens after a scary eval result gets published. The stock of attention moves toward the lab that produced the scariest number, because a model capable of assisting with a bioweapon synthesis pathway is, by implication, a model capable of a great deal of legitimate work. The safety report becomes a recruiting tool and a fundraising slide. A team that wanted to communicate caution ends up advertising reach.
This matters because the feedback loop rewards the wrong output. If demonstrating threat raises your valuation and your hiring pipeline, then producing threatening demonstrations becomes something you are quietly paid to do. The signal that was supposed to slow the field down speeds it up.
The persuasion category makes the bind obvious. A model rated high on persuasion is a model that can move people’s beliefs and actions at scale, which is a direct threat when the operator is a disinformation network and a direct selling point when the customer runs marketing, sales, or political consulting. The same measured behavior lands as a hazard on one slide and a revenue projection on the next. There is no clean way to publish the hazard without also publishing the pitch, and the lab knows both audiences are reading.
The scorecard is self-graded
The frameworks that assign these danger ratings are voluntary, and each lab writes its own. Anthropic decides what counts as ASL-3. OpenAI decides where “high” ends and “critical” begins. DeepMind sets its own capability thresholds and its own mitigations for crossing them. There is no shared rubric, no external body that certifies the grade, and in most cases no requirement that a third party can reproduce the test.
Compare that to how other high-consequence industries handle risk. A pressure vessel gets certified against ASME standards by an inspector who does not work for the manufacturer. A drug clears trials whose endpoints the FDA defines, not the drug’s maker. Aircraft get airworthiness certificates from regulators who can ground the fleet. AI safety frameworks have none of that scaffolding yet. The company shipping the product also writes the exam, grades it, and decides whether a failing score blocks the launch.
This matters because a self-graded test drifts toward whatever answer the grader needs. When the commercial pressure is to ship, the threshold for “safe enough” is a number the shipping party controls. Ask of any framework: who defined this threshold, and has anyone outside the company reproduced the result that it hangs on. If the answer to both is “the lab,” you are reading marketing that cites itself.
What a race to the bottom actually looks like here
The phrase gets used loosely, so here is the concrete mechanism. A lab runs an eval, gets a result that trips one of its own danger thresholds, and then ships the model anyway with mitigations that the same lab designed and judged sufficient. That decision sets a precedent. The next lab facing a similar result now has cover: the bar was cleared this way once, and the market did not punish it, so clearing it the same way is defensible.
METR, an outside evaluation group, has run autonomous-replication and task-completion tests on frontier models from multiple labs under time pressure before release. Time pressure is the tell. When an evaluator gets a few weeks with a model before a launch date that marketing has already fixed, the evaluation cannot function as a brake. It functions as a checkbox. A brake is something that can actually stop the car. If the launch date is immovable, the eval was never a brake.
Each individual decision looks reasonable from inside the company that makes it. The mitigation was real. The residual risk was judged acceptable. The competitor already shipped something comparable. Stack those reasonable decisions across four or five labs over three years and the collective floor for “safe enough” sinks, one defensible step at a time. No one chose the bottom. Everyone chose the next step down.
The demos drift toward theater
There are two kinds of dangerous-capability evaluation, and they pull in opposite directions. One kind is built to block a release: hard thresholds, adversarial testing, red teams empowered to say no, results that can actually delay a launch. The other kind is built to generate a headline number: a striking demonstration that the model did something alarming under lab conditions, packaged for a blog post and a press cycle.
The second kind is cheaper, faster, and better for engagement, so the incentive gradient runs toward it. This is safety-washing, the AI version of a company touting a recycling program while lobbying against emissions rules. The visible artifact says caution. The underlying behavior optimizes for reach.
You can tell the two apart by asking one question: what outcome of this eval would have stopped the release. If a bad result would have delayed the launch, cut a capability, or killed a feature, it was a brake. If every possible result ends with the model shipping on schedule and only the framing changes, it was theater. A red team that has never once blocked a release is a decoration, not a control.
What breaks if this keeps running
Three things erode, and they are the exact things a functioning safety regime depends on.
Shared thresholds go first. As long as every lab defines “dangerous” differently, no one can say whether the field as a whole is getting safer or more reckless, because there is no common yardstick to measure against. A “high” at one lab and a “critical” at another may describe the same behavior or wildly different ones, and outsiders cannot tell which.
Third-party verification goes next. If evaluators only ever get rushed, pre-launch access under NDA, independent testing collapses into rubber-stamping. The knowledge of how these models actually behave under adversarial pressure stays inside the companies that profit from shipping them, and the public gets the summary the company chose to write.
Trust goes last, and it takes the longest to rebuild. The first time a lab publishes a reassuring safety report and a serious harm surfaces that the report waved off, every future report from every lab loses credibility at once. The whole point of publishing evals is to earn the public license to keep building. Spend that credibility on marketing and it does not come back on demand.
Questions worth asking about any alarming AI demo
When a lab announces that its model reached a frightening new capability, treat the announcement the way you would treat a vendor’s own benchmark, because that is what it is. A short list separates disclosure from performance.
Who defined the danger threshold this result is measured against, and does anyone outside the company use the same one. Did the model ship anyway, and if so, who decided the mitigations were enough and whether that person’s compensation depends on shipping. Can an independent group reproduce the eval, or does verifying it require access only the lab controls. How long did outside evaluators have with the model before the launch date was set, and could their findings have moved that date. What specific result would have stopped the release, and has that ever actually happened.
If the answers show a company grading its own exam against a bar it set, on a schedule it fixed, with mitigations it judged, and no result could have changed the launch, the demonstration is a capability announcement wearing a safety costume. That is not a claim that any specific lab is acting in bad faith. It is a description of what the incentives produce when no shared standard, no independent grader, and no enforceable brake sits between the eval and the ship date. Those three things are the whole fix, and none of them exist yet at the level the technology now requires.
Contains a referral link.
Keep Reading
AI safetyResignations are signals, not scandals
How to read a high-profile resignation from an AI safety lab like Anthropic - what it signals, what to ignore, and what to watch in the months after.
AI governanceUnsealed court briefs turned training data into evidence
Unsealed briefs in the authors' case against Microsoft and OpenAI reveal training-data provenance as a board-level liability surface, not a technical detail.
AI safetyA 1938 law now points at AI critics
How labeling AI critics as 'foreign agents' under FARA-style rules chills safety research, disclosure, and open discourse - and what researchers can do.
Latest on the Wire
Full wire →- 16,000 Supabase databases exposed due to misconfigurationsBleepingComputer
- 2026: The Year AI Coding Agents Took OverSimon Willison
- 80,000+ Firms Hit by Stolen AI Logins in Supply-Chain Attack SurgeBleepingComputer
- AI Agent's Auto-Reply Blunder Worsens Delivery No-ShowSimon Willison
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.