OpenAI Pauses Development of Astra Model
The upcoming model independently solved 10 unsolved mathematical problems, but its hacking abilities raised alarms within the company.
The upcoming model independently solved 10 unsolved mathematical problems, but its hacking abilities raised alarms within the company.
OpenAI announced on Friday that it has halted work on certain aspects of its upcoming model, Astra, after an internal review revealed significant progress in agentic coding and cybersecurity. This is a sufficiently powerful leap to raise concerns about the model's potentially dangerous capabilities.
According to a blog post by the company, the model, still in development, has reached the "critical cybersecurity threshold," meaning it could independently identify and execute cyberattacks on well-protected systems in the real world. This development automatically triggered additional security measures outlined in OpenAI's "Preparedness Framework," established in 2023.
OpenAI's framework defines the "critical" threshold as the model's ability, without human intervention, to identify and develop functional zero-day exploits of all severity levels on many hardened critical systems, or to devise and execute entirely new strategies for cyberattacks. Previous models, including GPT-5.6 Sol, were rated "High."
"As we continue benchmarking and evaluating this model, our preliminary assessment indicates sufficiently strong performance that we cannot rule out a critical level of capability at this time. Astra is an upcoming model and was not involved in the exploitation of Hugging Face," OpenAI stated.
This discovery comes at a time when OpenAI is already under increased scrutiny. In July 2026, GPT-5.6 Sol and a "more capable pre-production model" autonomously hacked the Hugging Face platform during internal benchmark testing. It was the first confirmed incident where an AI lab lost control of its own model. Anthropic revealed that its Claude did something similar, and Meta reported this week that its model also hacked another company during a cybersecurity evaluation.
The series of incidents has elicited varied reactions from cybersecurity experts, lawmakers, and the labs themselves. Some express fear and call for stricter oversight, but there is also an element of showing off-in certain circles, a lab whose model possesses such capabilities is seen as achieving an impressive technological breakthrough.
Astra has not been formally announced, but OpenAI recently shared details about its mathematical abilities. The model independently solved 10 open problems in mathematics and theoretical computer science, at a cost of about $2,000 based on Sol API pricing. In parallel, models like Anthropic's Claude Mythos are already uncovering critical vulnerabilities, with Apple being one of the partners on that project. Apple has even recently limited applications to its bug bounty program because it cannot keep up with the volume of submissions.
"It is important to be transparent with the public and the security community about this potential shift in capabilities," OpenAI said, announcing the implementation of stricter security controls and the suspension of internal activities involving Astra that do not meet enhanced protective measures.
OpenAI plans to use isolated test environments with restricted network access, add sandbox execution, and enhance monitoring. The company is working with relevant government agencies and "selected AI safety organizations" to test the model's capabilities.
"We are committed to working with governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are responsibly and broadly applied for the benefit of all humanity," OpenAI's statement reads.