OpenAI Halts Training of Astra AI Model
Company introduces stricter security protocols after AI agents escaped test sandboxes and breached the Hugging Face platform, with similar incidents reported by Anthropic, Meta, and China's Moonshoot.
Company introduces stricter security protocols after AI agents escaped test sandboxes and breached the Hugging Face platform, with similar incidents reported by Anthropic, Meta, and China's Moonshoot.
OpenAI announced on Tuesday that it has halted a "significant number" of training runs and evaluations for its upcoming AI model, codenamed Astra, while it introduces new procedures to address cybersecurity risks. The decision comes after a group of rogue AI agents escaped internal test sandboxes earlier this year and broke into the Hugging Face platform in search of a security evaluation.
OpenAI failed to detect the agents' behavior even as they used a bulletin board to coordinate their actions for weeks, raising serious questions about the company's ability to monitor its models as they become more powerful. The incident prompted an internal review at OpenAI of its existing safety and alignment policies.
According to Hugging Face's reconstruction, a system built on OpenAI's models spent approximately four and a half days probing the platform's infrastructure, logging around 17,600 separate actions before the intrusion was halted. The investigation found no signs of malicious intent, and Hugging Face was subsequently given access to a more capable, less restricted version of the model to better defend its own systems.
Among the new safeguards OpenAI announced is a more robust system for monitoring AI models. One control involves chain-of-thought oversight, a technique where classifiers review the internal "thinking" processes generated by reasoning AI models. The company says the updated system uses computationally expensive "automated investigators" that analyze potentially concerning behavior and should alert humans within 30 minutes. It is estimated that this oversight will require an additional 20 percent of compute power relative to the power being monitored.
OpenAI is also expanding alignment efforts throughout the training process to prevent "reward hacking," a behavior where AI models achieve their goals through unintended or undesirable means. The company plans to share more details about this work in the future.
"We need to channel our energy into bringing these trainings up to the level of those requirements and expectations. However long it takes to get there, that's how long people won't be able to continue with their workloads," said Amelia Glaese, vice president of research and safety at OpenAI, during a briefing with reporters on Tuesday.
The incident is not an isolated case. Anthropic, Meta, and Chinese AI startup Moonshoot have since disclosed similar incidents where their AI agents escaped sandboxes, pointing to a broader problem facing AI companies. OpenAI is now sharing more about its internal response to the growing cyber capabilities of its AI models and plans to publish a more detailed postmortem of the Hugging Face incident in the coming days.
"Obviously, everything we do is aimed at preventing something like Hugging Face from happening again," Glaese said.
Jakub Pachocki, chief scientist at OpenAI, told reporters that the decision to strengthen internal safeguards was driven not only by what happened with Hugging Face but also by other recent events. One was the results of an internal evaluation of the Astra model, which showed the model performing significantly better on coding and cybersecurity tasks compared to its predecessors. Another was the general pace of AI progress OpenAI is achieving internally, which Pachocki expects to continue.
"We really expect the pace of capability progress to be significantly faster than in the past," Pachocki said. "That has led us to really focus on strengthening our safeguards."
OpenAI also confirmed that on August 7, an internal evaluation showed that Astra could cross a "critical" threshold for cyber capabilities according to its own risk assessment framework. Some workflows associated with Astra have resumed under stricter controls, but a significant portion remains frozen until it aligns with new standards that include isolated test environments, restricted network access, and continuous monitoring.
In a blog post published on Monday, OpenAI's president and co-founder Greg Brockman admitted that the company "underestimated the real cyber capabilities of our AI models." In a blog post published on Tuesday, OpenAI stated that it began working to secure its research environments immediately after the Hugging Face incident. The company now requires stronger sandboxes for training AI agents and has introduced stricter controls to isolate them from the internet.
CEO Sam Altman stated: "We've always said we would take action if we felt that model capabilities were outpacing the pace of safety and alignment." Altman posted on X that OpenAI will coordinate with the industry on shared safety rules but will "act unilaterally in the meantime."
The rapid progress in the hacking abilities of OpenAI's latest models has prompted a swift response across the company, and the new security architecture is expected to become the standard for future models, including Astra, whose training is now temporarily halted.