Tech · Culture · Fiction
Article Anthropic, OpenAI, and DeepMind grade their own models' danger
AI labs publishing dangerous-capability evals turns safety disclosure into marketing, and the self-graded scorecards drift toward a race to the bottom.
It walks through the handshake between your apps
Ember-1 exploits trust relationships between applications, moving past perimeter and identity controls. What failed, why, and what must now be true.
Loading a trusted model runs untrusted code.
from_pretrained resolves a Hugging Face model name to whatever the branch holds now, extending a one-time trust decision over content it never re-checks.
Loading a model executes the uploader's code
How autonomous AI agents weaponize Hugging Face via pickle deserialization and trust_remote_code, mapped to MITRE ATT&CK and ATLAS with the telemetry defenders see.
The bypass is a feature
Persistent authentication stores a completed verification as a token, then acts on the token forever. Reference replaces validation, and the person goes unchecked.
TOTP fixes the channel, not the credential
TOTP kills SMS interception and SIM-swap OTP theft, but AiTM phishing still steals the session token. What TOTP secures and what it doesn't.
Your decisioning problem isn't accuracy.
Run open-source decision models locally-pinned versions, validation, immutable logs-so every approval or denial stays reproducible and auditable.
The Wire — latest
All →- 1841 rail delay was not the earliest space weather event
- 1996 Chat Room Simulator Revives Vintage Web Desktops
- AI Agent Breaches Dutch Cybersecurity Nonprofit in Unprecedented Attack
- AI Agents Pose New Insider Threat Risks
- AI Capability Jumps Pose Sudden Cybersecurity Challenges
- AI firms compete to showcase most dangerous models
- AI Firms Leak User Data to Advertisers
- AI Won't Solve Coding: Why Critical Software Still Needs Human Oversight
- AI's Economic Impact: Old Games Show Transformation Power
- American Dinner Parties Collapse: A 70% Drop in 50 Years
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.