Security researchers say Anthropic's Fable blocks even routine, harmless requests
Original source
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
Hacker News →Anthropic launched Fable on Tuesday as a publicly available, scaled-down counterpart to Mythos, its restricted cybersecurity model. The release immediately drew criticism from security professionals who say the model’s safety guardrails are so aggressive they make it nearly useless for everyday work. IBM X-Force researcher Valentina Palmiotti reported that Fable refuses anything even loosely connected to security, including reading a blog post, while others complained that simple code reviews or requests for secure coding practices trip the filters. When triggered, the model halts the conversation, citing safety flags for cybersecurity or biology topics, and falls back to Claude Opus 4.8.
The restrictions reflect Anthropic’s worry that a capable model could help build malware or, in the biology case, bioweapons. Critics argue the implementation is crude — veteran researcher Matt Suiche described it as essentially keyword-based, penalizing anything in the lexical neighborhood of ‘cybersecurity,’ including ordinary software engineering questions about writing secure code. Still, Suiche defended the company’s caution as reasonable for an early release, predicting the guardrails will loosen as frontier labs build closer ties with the newer wave of AI-focused security firms.
The episode highlights a structural tension in how AI companies handle dual-use capabilities. Anthropic gates full cybersecurity functionality behind its Cyber Verification Program, mirroring OpenAI’s Trusted Access for Cyber, while Mythos itself rolled out under Project Glasswing to a vetted set of organizations — recently expanded to hundreds across 15 countries. The complaints suggest that coarse public-facing filters risk alienating the defenders these models are ostensibly meant to help, while verified-access programs become the de facto path to real capability.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.