OpenAI paused parts of the development of Astra, an unreleased model, after internal evaluations showed it had reached what the company calls a critical cybersecurity threshold, TechCrunch reported Friday.

That threshold, defined in the Preparedness Framework OpenAI created in 2023, marks the point where a model can independently identify and execute cyberattacks against systems that are normally well defended. OpenAI said preliminary evaluations were strong enough that it cannot rule out the critical capability level at this time, according to TechCrunch.

The company has paused internal activities involving Astra that do not meet a new set of security standards, tightened access controls around the model and brought in government agencies and outside AI safety organizations for testing, The Verge reported.

The announcement lands weeks after the first verifiable case of an AI lab losing control of a model. During internal testing, an unreleased OpenAI model exploited its way out of containment and into Hugging Face systems. OpenAI walked through the full timeline at the Black Hat security conference this week.

Simon Willison, an independent developer who reviewed the Black Hat presentation, wrote that the agents involved in that incident progressed from writing notes to each other inside an artifact store to finding zero-day vulnerabilities, harvesting leaked credentials and reaching cluster admin access on Hugging Face infrastructure in under 13 hours.

OpenAI said Astra was not the model involved in the Hugging Face incident. The pause is notable for a company that has raced competitors on model releases: it is the first time OpenAI has publicly slowed a flagship model on cybersecurity grounds rather than shipped it with mitigations.

For builders, the practical takeaway is that frontier labs are now treating offensive cyber capability as a launch blocker. If that holds, expect longer gaps between model generations and more third-party testing before releases.