OpenAI's Astra Model Can Apparently Hack on Its Own — So They Slammed the Brakes

Something changed inside OpenAI's labs recently, and for once, the company said so out loud. The startup announced Friday that its upcoming model, internally designated Astra, had performed well enough on internal cybersecurity evaluations that OpenAI could not rule out it had crossed into "critical" territory — the highest risk tier in its own Preparedness Framework. That framework, published by OpenAI itself, defines "critical" as a model capable of autonomously identifying and exploiting previously unknown software vulnerabilities — zero-days, in the industry's parlance — or executing sophisticated, end-to-end cyberattacks against hardened targets without meaningful human direction. In plain language: a model that could hack things on its own, at a level that strains the defenses of well-resourced organizations.
This is not a threshold OpenAI had ever triggered before. The company has run its Preparedness evaluations on previous frontier models and, until now, always landed short of the top tier. Astra is the first time the internal scorecard came back with a result the safety team couldn't confidently wave off. The response was immediate and, by Silicon Valley standards, unusually cautious: a pause on certain internal development work and the activation of protocols the framework lays out specifically for this scenario.
The Preparedness Framework — a document OpenAI released publicly in late 2023 and has since updated — describes a four-tier risk ladder for four domains: cybersecurity, biological and chemical weapons, nuclear and radiological threats, and what it calls "persuasion." Most models, OpenAI has argued, stay in the low-to-medium range. Getting to "high" is supposed to require a serious internal response. Getting to "critical" is supposed to stop the train entirely. The company now says Astra may be sitting at that top rung, at least in the cyber domain — and that uncertainty alone is enough to engage the brakes.
What exactly Astra did in testing has not been disclosed in full. OpenAI's public statement describes the concern as an inability to "rule out" critical-level capabilities, which is a careful, lawyerly formulation — it does not confirm the model can execute nation-state-grade cyberattacks, only that internal evaluators could not demonstrate it couldn't. The distinction matters, but so does the threshold for triggering a pause: under the framework's own language, a model does not have to be confirmed dangerous to warrant a stop. Uncertainty at the top tier is itself the tripwire.
The timing is not incidental. The AI industry is in a full sprint toward systems capable of "agentic" behavior — models that don't just answer questions but take sequences of actions in the world with minimal human input. That capability, which every major lab is racing to deploy, is also precisely what makes offensive cyber use so alarming. A model that can plan, iterate, pivot, and execute across a complex technical environment is not categorically different from a model that can plan, iterate, pivot, and execute a cyberattack. The researchers who build these systems have said so privately for years. Astra is the first case where that concern has surfaced in an official, public-facing safety document from one of the labs themselves.
OpenAI's Preparedness Framework assigns ownership of these evaluations to a dedicated internal team, with oversight from the company's safety board. Under the framework, a "critical" finding in any domain is supposed to result in the model not being deployed externally — full stop — until mitigations are validated. Whether those mitigations involve capability suppression, fine-tuning, architectural constraints, or some combination is not public. What is public is that the pause is real and that the company is framing this as the safety system working as designed, rather than as a failure.
That framing deserves scrutiny. OpenAI is, after all, the same organization that has faced sustained criticism from former employees, regulators, and its own ex-board members over whether commercial pressure consistently wins arguments with safety caution inside its walls. The pause on Astra is genuinely notable — it is not something a company eager to ship product does lightly, and it is the kind of action safety advocates have long demanded labs take proactively. But it is also a company describing itself as doing the right thing, and the underlying evaluation data, the specific capabilities demonstrated, and the criteria for lifting the pause remain internal.
What Astra ultimately represents is a live stress test of a system the AI industry has been building in theory for two years: the idea that frontier labs can self-govern at the capability frontier through internal safety frameworks, without waiting for regulators to force the issue. OpenAI is now pointing to the Astra pause as evidence that self-governance works. Skeptics will note that the public only knows about this because OpenAI chose to disclose it — and that the decision about when Astra is safe enough to ship will also be made entirely by OpenAI.
Who is covering this (18+ outlets)
- iClarified - Apple News and TutorialsOpenAI's Astra Model Hits 'Critical' Cybersecurity Threshold, Prompting Safety Pause
- The GuardianOpenAI to pause work on AI model Astra due to security concerns
- NeowinOpenAI is locking down its upcoming Astra model after alarming cybersecurity test results
- Mashable IndiaWhat is OpenAI Astra? Everything we know about the new -- and possibly dangerous -- new model
- Fello AIOpenAI Astra: What It Is and When You Can Use It
- Okay NewsOpenAI Pauses Work on New AI After It Shows Ability to Launch...
- mintWhy is OpenAI slowing down the release of its new Astra AI model? | Mint
- Analytics InsightOpenAI Pauses Astra AI Model Work Over Cybersecurity Risks After Hugging Face Hack
- THE DECODEROpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time
- CryptoRankOpenAI Pauses Astra Model Work Over Cybersecurity Threshold Concerns | AI News OpenAI
- TechlusiveWhy OpenAI paused work on Astra AI model after internal cybersecurity tests
- BlockonomiOpenAI Halts Astra AI Development Over Autonomous Hacking Concerns
- Tech TimesOpenAI Pauses Astra After Tests Reveal Autonomous Zero-Day Exploit of Hardened Systems
- TechGenyzOpenAI Astra AI Raises Critical Cybersecurity Concerns, Pauses Development - Techgenyz
- The Times of IndiaAfter Hugging Face hack, OpenAI pauses work on Astra AI model over cybersecurity risks
- DataBreachTodayOpenAI Is Watching Astra Think. Can It See Trouble?
- The Hans IndiaOpenAI pauses development of new Astra model to strengthen cybersecurity safeguards
- The Straits TimesOpenAI pauses AI model work amid cyber concerns
See what people are saying about this story on X.
