OpenAI calls Astra its most dangerous model yet – watching what it does is only getting harder
OpenAI is rating its upcoming Astra model as the first system with “critical” cyber capabilities, while promising it’s also the safest model the company has built. But a report on Astra’s architecture raises questions.
It’s an unusual way to announce a product: OpenAI says its upcoming Astra model is so dangerous that it hits the highest risk tier for cybersecurity in the company’s own Preparedness Framework, and in the same breath calls it the safest model it has ever built. Given the right tools, Astra can find and exploit previously unknown security holes in well-protected systems, without a human guiding each step. No model before this got that rating from OpenAI, though the company had already hinted Astra might get there. The timing is probably no accident. The warning landed the same day rival Anthropic shipped Claude Fable 5.1 and Mythos 5.1, and the two companies have a habit of dropping announcements right around each other’s releases. On X, CEO Sam Altman explained why his company isn’t shipping anything new: the team spent the summer “sprinting on safety priorities,” Astra “has been done training for a while now,” and the models after it are being slowed down on purpose. Users on X read that as an excuse from a company falling behind. According to The Information, Anthropic passed OpenAI on revenue this year. OpenAI backs up the critical rating with a batch of tests. On ExploitBench, a benchmark that measures how well a model builds exploits from known vulnerabilities, Astra scored full marks. Worried those tasks might have leaked into its training data, the company built an internal follow-up stocked with 20 recently disclosed, high-severity V8 vulnerabilities. There too, Astra beat its predecessor GPT-5.6 Sol by a wide margin, while burning far fewer tokens. It also found two previously unknown zero-day flaws and chained them into a working exploit. OpenAI says it’s now reporting those vulnerabilities to the people responsible for the affected software. In expert-led tests, the model went further.