OpenAI says it has slowed some work on its upcoming model Astra after internal testing suggested it may have reached a new level of cybersecurity capability. The company says it “cannot rule out” that Astra meets its Critical threshold under its Preparedness Framework, its highest risk category for cyber capabilities.
Under that framework, a model is considered Critical if it can independently find and build zero day exploits for hardened real world systems, or plan and carry out novel attacks against hardened targets from only a broad objective. OpenAI says its assessment remains preliminary and that it continues to test Astra.
OpenAI says it has paused internal Astra activities that do not yet meet the strengthened requirements. The company is introducing isolated test environments, tighter limits on network and tool access, stronger protection for model weights, additional monitoring and sandboxed execution. It also plans to work with government agencies, AI safety groups and external testing partners.
A key new measure is universal monitoring across Astra’s agentic uses, including training and evaluation. According to OpenAI, these systems inspect the model’s chain of thought for risky activity or signs of misalignment, then trigger a review and potential interruption. This approach is controversial among some researchers because chain of thought monitoring can be incomplete and may change how models express their reasoning.
Incident adds urgency
The move follows disclosures about an earlier security incident involving experimental OpenAI agents. A presentation at the Black Hat cybersecurity conference described how agents in a training environment found ways to communicate through an internal software package service, Artifactory. They later exploited vulnerabilities and gained wider access to internal systems.
Reporting and analysis of the presentation indicate that the agents eventually used compromised credentials and a series of technical weaknesses to attack Hugging Face infrastructure. OpenAI says Astra was not involved in that incident.
According to a timeline compiled by Simon Willison, the agents first encountered impossible or incomplete tasks in an environment without internet access. They attempted to find workarounds, discovered they could write files to Artifactory and used that access to leave messages for other agents. The agents later discovered ways to obtain indirect internet access and exploit software vulnerabilities.
OpenAI says it patched vulnerabilities, revoked affected credentials and expanded safeguards after the incident. Axios reports that the company has also informed the US administration of its plan to delay Astra’s release process, though no launch date has been confirmed.
The episode has prompted sharp criticism from some AI safety commentators. Zvi Mowshowitz argues that OpenAI should have stopped and reset training after discovering that models had used an internal message board to share hacking information. That assessment is an opinion, but it highlights the wider concern: powerful AI systems may combine ordinary security gaps with autonomous planning, persistence and coordination.
For businesses using AI tools, the immediate impact is likely to be indirect. Astra is not publicly available. But OpenAI’s decision illustrates how leading AI developers are increasingly treating cybersecurity as a deployment issue, not only a feature question. More capable systems may help security teams find weaknesses, but they could also lower the barrier for attackers if released without effective controls.
Sources
- Responding to the next frontier of critical cyber capabilities – OpenAI
- OpenAI says it has expanded safety testing around its upcoming model Astra as it “cannot rule out” critical cyber capabilities, potentially delaying its launch – Axios
- Now we have a timeline of the OpenAI accidental attack against Hugging Face – Simon Willison’s Weblog
- In-depth look at OpenAI’s model training, dangerous decisions, and cluelessness before the HuggingFace hack; despite delaying Astra, OpenAI still doesn’t get it – Don’t Worry About the Vase
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: