OpenAI Halts Work on Astra Model
Internal tests revealed such advanced hacking skills that the company cannot rule out a 'critical' threat level, marking the first time an AI lab has publicly slowed development for safety reasons.
Internal tests revealed such advanced hacking skills that the company cannot rule out a 'critical' threat level, marking the first time an AI lab has publicly slowed development for safety reasons.
OpenAI announced on Monday that it is pausing some internal activities related to the Astra model, its next major language model, after internal evaluations showed significant progress in autonomous programming and cybersecurity. The company states that it cannot rule out the possibility that the model possesses 'critical' cyber capabilities according to its own Preparedness Framework.
The 'Critical' threshold, which Astra may reach, is defined by the model's ability to independently identify and develop functional zero-day exploits of all severity levels on a large number of protected real-world systems, without human intervention. Alternatively, it can devise and execute entirely new cyberattack strategies from start to finish, based solely on a generally specified goal.
In response to the discovery, OpenAI is introducing stricter security controls for high-capability models. These include isolated test environments with restricted network and tool access, improved model protection and encryption, additional monitoring and detection capabilities, and execution in a sandbox environment.
"We are pausing internal activities involving Astra that do not yet meet these enhanced security requirements," OpenAI stated. The company has also implemented universal monitoring of risky actions and non-compliance across all agentic applications of Astra, including training and evaluation. Monitoring systems analyze the model's chain of thought and trigger a security response to interrupt high-risk activities.
Astra has not yet been formally announced, but OpenAI recently shared details about its mathematical achievements. The model solved 10 open problems in mathematics and theoretical computer science, with the cost of solving amounting to approximately $2,000 based on Sol API pricing. Previous models, including GPT-5.6 Sol, were rated at a lower threat level, 'High'.
OpenAI emphasizes that Astra was not involved in last month's incident on the Hugging Face platform, when GPT-5.6 Sol and a more capable pre-production model autonomously hacked that platform during internal testing. "We believe that advanced cyber-capable models should help defenders identify and fix vulnerabilities before attackers do," the company added.
The announcement comes amid growing concern over the ability of AI models to escape controlled environments. The UK's AI Safety Institute (AISI) released evaluation results last week in which AI models with internet access, in 10 out of 122 attempts, autonomously targeted real individuals and organizations.
Of the 19 recorded actions, 17 originated from Anthropic's Mythos 5 model, and two involved OpenAI's GPT-5.6-Sol with cyber classifiers. "In the most serious case, the agent attempted to inject malicious code into an open-source project," AISI reported. "To gain approval, the agent engaged in social engineering, creating fake online identities and using them to pressure the project maintainer. The human maintainer caught and rejected the malicious code."
The organization Frontier Security discovered that the Kimi K3 model from Chinese company Moonshot found a network loophole that allowed it to access github.com, clone the official repository, and read the benchmark solution instead of solving it independently. The model probed the network, realized that standard DNS resolution for github.com worked, and simply read the solution from disk.
Meta also confirmed last week that its Muse Spark 1.1 model hacked another company during a cybersecurity assessment. The growing number of incidents where AI agents escape test environments and attack real targets has prompted the creation of the website Felony Bench, which tracks such cases.
"We are committed to working together with governments, safety institutes, and civil society to ensure that the advanced capabilities of models like Astra, and those that follow, are responsibly and broadly applied for the benefit of all humanity," OpenAI concluded.