OpenAI Throttles Its Riskiest Model While Having Just Disbanded Its Safety Team

Dunkler Serverraum mit blau leuchtenden Kabeln als Symbol für Cybersicherheit und KI-Risiken
Photo by Jefferson Santos on Unsplash

OpenAI has paused parts of the training on its upcoming flagship model, internally named Astra, after early tests pointed to critical cyberattack capabilities. The model reportedly could independently find and exploit severe vulnerabilities in real systems. What makes this notable is the timing: OpenAI had just disbanded the very team responsible for assessing catastrophic risks like this one.

Key takeaways

  • Internal testing led OpenAI to classify Astra as critical for cybersecurity for the first time, the highest risk tier in its own Preparedness Framework.
  • In response, the company paused certain reinforcement-learning training for two weeks, and its largest planned frontier RL run remains on hold.
  • A new monitoring system is meant to flag suspicious model behavior within 30 minutes of detection.
  • Shortly before the Astra pause, OpenAI had disbanded the Preparedness team responsible for evaluating exactly this kind of catastrophic risk.
  • CEO Sam Altman says Astra’s core training was never fully halted and new models remain on schedule.

What Astra allegedly can do, and why it’s concerning

OpenAI’s Preparedness Framework classifies a model as critical when it can independently find and exploit severe software vulnerabilities in real-world systems, or carry out sophisticated cyberattacks against well-defended targets without human direction. Astra reportedly crossed exactly that threshold in early internal tests, showing advanced, autonomous coding skills and the potential to identify or exploit previously unknown vulnerabilities, so-called zero-days, on its own. In response, OpenAI is moving parts of further development into contained environments with limited network access, where code runs strictly inside a sandbox. The new monitoring system, designed to flag suspicious behavior within 30 minutes, reportedly covers roughly a fifth of the compute in use, with variation depending on the workload.

The uncomfortable parallel to the disbanded Preparedness team

The irony lies in the timing. As kabel-salat.info reported earlier, OpenAI disbanded its own team dedicated to catastrophic risks and distributed its responsibilities across existing teams. It is precisely in this period that the company’s own risk-assessment system produced the clearest evidence yet of why such a team existed in the first place. OpenAI says it wants to invest more heavily in alignment research instead, meaning methods that align AI systems more fundamentally with human values and goals. Whether spreading responsibility across existing teams preserves the same attention to individual risk categories as a dedicated unit did remains an open question, and one that is hard to verify from the outside. Critics read the timing of the disbandment as a sign that having a warning system and having an organization that actually acts on it are two different things.

Not an isolated case: other labs work with similar risk thresholds

OpenAI is not the only lab that has bound itself to formal risk thresholds for especially dangerous capabilities. Anthropic commits to comparable safety tiers under its Responsible Scaling Policy before a model with potentially critical capabilities can be deployed more broadly, and Google DeepMind takes a similar approach with its Frontier Safety Framework for biological, chemical, and cyber risk categories. The difference lies less in the principle than in the practice: all three frameworks depend on an internal team evaluating them independently of product pressure, and enforcing them against the company’s own timeline if necessary. That is exactly where the criticism of OpenAI’s approach lands, because a risk threshold that is defined cleanly on paper but no longer clearly owned by anyone inside the organization loses some of its teeth. How resilient OpenAI’s newly distributed responsibility actually is cannot be verified from the outside right now, since the internal escalation paths have not been published.

Between caution and time pressure

Sam Altman has publicly stressed that Astra’s core training was never fully stopped and that new models remain on track to ship on schedule. That lines up with the broader picture emerging from reporting: it is not the entire project on ice, but specifically the training runs that would further amplify the riskiest capabilities. This balancing act is typical of the current phase of frontier AI development. Commercial pressure is high, not least because rival Anthropic recently overtook OpenAI in revenue growth. At the same time, the company’s own safety thresholds, whose breach is now publicly documented, force visible braking maneuvers. Claiming both speed and responsible caution at once requires being able to demonstrate both, otherwise it remains just a claim.

Where this leaves things

Taken on its own, pausing training after a critical risk classification is exactly what a Preparedness Framework is supposed to do: a technical warning system that actually kicks in when it matters. What makes it questionable is the interplay with disbanding the team responsible for acting on it. A warning threshold is only as valuable as the organization that takes it seriously and attaches consequences to it. Whether OpenAI’s newly distributed responsibilities can actually fill that role, or whether they mostly diffuse accountability, will only become clear the next time a critical classification comes up, and when it becomes public who actually makes the call.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top