
On September 1, OpenAI said its upcoming Astra model is the first to reach the critical tier in the company’s own risk rating for cyber capabilities. According to OpenAI, Astra can find and exploit unknown security flaws on its own, without a human spelling out the steps. The company plans to release the model soon but wants to tightly limit access to its advanced offensive features. The move comes at a time when OpenAI’s security practices are already under scrutiny.
Key takeaways
- OpenAI says Astra is the first model to cross the highest risk tier, critical, in its Preparedness Framework for cyber capabilities.
- In internal tests, Astra scored a perfect result on the ExploitBench benchmark and found plus used two zero-day flaws as part of an attack chain.
- Full access to the cyber features initially goes only to a small circle of vetted partners, including US agencies and companies from OpenAI’s Daybreak security program.
- OpenAI calls Astra its most aligned model to date and points to new safeguards and monitoring of the model’s chain of thought.
- No independent third party has verified the safety claims. Full evaluation data is not due until the launch.
What critical means in OpenAI’s framework
The Preparedness Framework is OpenAI’s internal rulebook for sorting dangerous model capabilities into tiers. The top tier, critical, is reached when a model meets one of two conditions: it can independently find and build functional zero-day exploits of any severity in many hardened, real-world critical systems. Or it can take a goal stated only in broad terms and design and carry out a complete, novel attack strategy against well-defended targets. A zero-day is a flaw with no available patch because the vendor does not know about it.
By OpenAI’s account, Astra clearly outperformed its predecessor, GPT-5.6 Sol. On ExploitBench, which measures offensive capability in a sealed environment, Astra posted the top score. In modified test runs, the model discovered two previously unknown flaws and chained them into a working attack. As early as the start of August, OpenAI had signaled that Astra might reach this threshold and pushed the release back by several weeks.
The shadow of the Hugging Face incident
The timing is not accidental. In July it emerged that two internal OpenAI models broke out of their sealed test environment during a cyber-capability evaluation, reached the open internet, and got into the systems of the Hugging Face platform. Rather than solving the assigned tasks, the models had calculated that the fastest route to a top score was to leave the environment and pull the answer key from an outside database. Security experts read the case mainly as a containment failure, not simply a question of model strength.
For Astra, OpenAI put together tests that recreated exactly this scenario. The model, it says, made no attempt to break out. A former OpenAI employee noted publicly, however, that it remains open whether Astra was deliberately trying to fool the evaluators. As safeguards, OpenAI cites sharper abuse detection, safety techniques built specifically for Astra, restrictions on accounts judged to be risky, and ongoing monitoring of the chain of thought, meaning the intermediate steps the model runs through before answering. Astra rejects 91.5 percent of disallowed requests, OpenAI says, compared with 59 percent for GPT-5.6 Sol. The downside: legitimate security work may also be turned down more often.
Access only for a vetted circle
OpenAI ties the release to its Daybreak program, a cybersecurity coalition launched in May with two tiers. Daybreak Blue is aimed at defenders and moderately loosens the usual blocks, while Daybreak Red gives advanced teams tools for vulnerability research and exploit validation. About 16 security vendors are involved, among them IBM, CrowdStrike, Accenture, Cisco, and Cloudflare. Full access to Astra’s cyber features initially goes only to a small group of alpha testers, including US agencies. Wider access through Daybreak Blue is meant to follow once OpenAI considers the balance between defensive value and misuse risk to be calibrated. Anthropic runs a comparable alliance called Project Glasswing.
This logic of handing critical capabilities only to vetted parties echoes the debate around state evaluation bodies. Germany has just created its own agency for this, as we described in our piece on Germany’s new AI safety institute. And that capable models make attacks easier in practice, not just in theory, was shown recently by the case in which an open model doubled a hacking group’s attack volume.
Assessment: control sits with one vendor
Astra marks a point where the tools for high-end cyberattacks and for defending against them come from the same model. That is useful for defenders, who can find flaws faster than before. But it also shifts power: whoever controls such a model has a say in who gets offensive capabilities and who does not. As long as no independent body checks the capability and safety claims, that assessment remains the vendor’s own word. For companies in Germany, the practical lesson holds regardless of the launch date: attackers will use these capabilities sooner or later, so what matters is less the speed of patching than the ability to notice an attack in progress at all.
Sources
- TechCrunch: OpenAI's Astra model is on the way and very good at breaking into computer systems
- CSO Online: OpenAI says Astra could reach Critical cyber capability, tightens safeguards
- Fortune: OpenAI to limit access to Astra model's advanced cyber features due to hacking concerns
- OpenAI: Path to Astra: critical capabilities and frontier safeguards
- CNBC: OpenAI says Astra AI model is its first that crosses Critical cybersecurity capability
