OpenAI says upcoming 'Astra' model may cross 'High' cyber-capability threshold
OpenAI’s internal evaluations of an upcoming model, referred to as Astra, indicate it may reach the ‘High’ cyber-capability tier under the company’s Preparedness Framework — a bar first defined in December 2023, well before frontier models neared it. In OpenAI’s framing, a ‘High’ cyber model could develop working zero-day remote exploits against well-defended systems, or meaningfully assist complex enterprise and industrial intrusion operations aimed at real-world effects. The company says it can no longer rule out such capabilities, and is disclosing this proactively rather than waiting for a confirmed classification.
On the mitigation side, OpenAI points to access controls, infrastructure hardening, egress controls, and monitoring around the model, alongside investment in defensive tooling — helping security teams audit code and patch vulnerabilities faster. It also plans a program offering vetted cyberdefense users and customers tiered access to enhanced capabilities, and is standing up a ‘Frontier Risk Council’ to bring outside security practitioners into its risk decisions, starting with cybersecurity before expanding to other domains.
The disclosure lands amid real-world signals that agentic AI is already being turned against infrastructure — including a contained security incident at Hugging Face involving a compromising AI agent. The significance is less any single feature than the inflection point it marks: a major lab conceding that its own systems may soon be capable enough to matter offensively, and trying to get ahead of that by tilting access toward defenders and formalizing external oversight. The obvious tension is that the same capabilities meant to empower defenders are exactly what makes the model dangerous in the wrong hands.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.