OpenAI Hits the Brakes on Astra After the Model Gets Too Good at Hacking

OpenAI Hits the Brakes on Astra After the Model Gets Too Good at Hacking

Internal tests found cybersecurity capabilities strong enough that OpenAI can't rule out its own 'critical' risk threshold — so it's pausing parts of development to find out.

Written by OutOfToken AI

August 10, 2026 · 4 min read · Synthesized from reporting by The Hacker News · How this works

AI Verified · 9/10

OpenAI has paused some internal work on Astra, its next frontier model, after evaluations showed cybersecurity and agentic coding skills strong enough to potentially cross the company's highest self-defined risk threshold. The announcement, made in a blog post on Friday, marks the first time OpenAI has flagged one of its own models as a possible 'critical' cyber risk before public release.

What 'Critical' Actually Means

Under OpenAI's Preparedness Framework, first introduced in 2023, a model reaches the critical cybersecurity threshold if it can autonomously identify and develop functional zero-day exploits across all severity levels against hardened, real-world systems — without human help. The same threshold also covers models capable of devising and executing end-to-end novel cyberattack strategies against well-defended targets on their own.

The Trigger

OpenAI says internal evaluations over the past few days revealed 'significant advancements' in Astra's agentic coding and cybersecurity performance. The company stopped short of confirming Astra has actually crossed the line, stating instead that preliminary results are strong enough that it cannot rule out the possibility while benchmarking continues.

"OpenAI's own framework treats the mere possibility of a model finding zero-days in hardened systems unassisted as reason enough to halt work and lock down access — a rare case of a lab pausing itself rather than waiting for outside pressure."

Locking Astra Down

In response, OpenAI says it's rolling out new security controls for higher-capability models and the activities surrounding them, including isolated environments meant to contain what Astra can access and do during continued testing. Some internal activities involving the model have been paused outright while these safeguards are put in place, though OpenAI hasn't detailed which activities specifically or for how long the pause will last.

Not an Isolated Signal

OpenAI's disclosure lands alongside similar warnings from Anthropic and Meta, both of which have recently reported their own frontier models brushing up against advanced offensive-cyber capabilities. The pattern suggests the industry's most capable systems are approaching a point where autonomous exploit discovery is no longer theoretical — a shift that safety frameworks built years earlier are now being tested against in real time.

OpenAI hasn't said when Astra will resume full development or ship publicly, and it remains unclear whether further testing will confirm or walk back the critical designation. What's clear is that the industry's voluntary safety thresholds, long treated as largely precautionary, just got their first real stress test against a model that might actually warrant them.

Editorial Note

The research corroborates all major factual claims in the article regarding Astra's capabilities assessment, the critical threshold definition, OpenAI's precautionary pause, and the 2023 framework introduction. Sources consistently confirm that OpenAI cannot rule out—but has not definitively confirmed—that Astra meets the critical cybersecurity threshold. The only unverified claim is whether this is truly the 'first time' such a flag has been raised, which the sources do not address.

Claim Tracker

AI-assessed

VerifiedOpenAI paused internal work on Astra after evaluations showed cybersecurity and agentic coding skills strong enough to potentially cross the company's highest self-defined risk threshold.

Confirmed by Source 1 (Ground News), Source 3 (iClarified), Source 4 (The Decoder), and Source 5 (TechCrunch). All sources corroborate the pause and the 'critical' threshold evaluation.

VerifiedUnder OpenAI's Preparedness Framework, a model reaches the critical cybersecurity threshold if it can autonomously identify and develop functional zero-day exploits across all severity levels against hardened, real-world systems without human help.

Directly confirmed by Source 1, Source 3, and Source 5. The definition of the critical threshold is consistently reported across sources.

VerifiedOpenAI stopped short of confirming Astra has actually crossed the line, stating instead that preliminary results are strong enough that it cannot rule out the possibility.

Source 4 (The Decoder) and Source 5 (TechCrunch) both confirm OpenAI 'cannot rule out' the critical level, emphasizing this is not a definitive confirmation but a precautionary threshold trigger.

VerifiedOpenAI's Preparedness Framework was first introduced in 2023.

Source 5 (TechCrunch) confirms the framework was 'created in 2023.'

UnverifiedThe announcement marks the first time OpenAI has flagged one of its own models as a possible 'critical' cyber risk before public release.

No source explicitly confirms this is the 'first time' OpenAI has done this. While sources confirm this announcement is significant, they do not provide historical comparison to verify the 'first time' claim.

Ask AI about this story

// discussion

sign in to join the discussion