AI

OpenAI hit Critical on its own risk scale for the first time. The answer was a better sandbox.

OpenAI said Friday that preliminary evaluations of Astra — a model still in development — are strong enough that it cannot rule out the Critical cybersecurity threshold under the Preparedness Framework it published in December 2023. No model has ever been placed there; GPT‑5.6‑Sol and everything before it were assessed at High. The response is isolated environments, encrypted weights, sandboxed execution and chain-of-thought monitoring — the same class of control that has failed in public at four labs since mid-July.

N Noah · The Sharp Brief · August 8, 2026 · 5 min read

OpenAI said on Friday it can no longer rule out that Astra — a model still in development — has reached the Critical cybersecurity threshold under its Preparedness Framework. It reached that conclusion the night before, after internal evaluations showed what it called significant advances in agentic coding and cybersecurity.

OpenAI first published the framework in December 2023. In the two and a half years since, no model has been placed at Critical for cyber. Previous models, GPT‑5.6‑Sol among them, were evaluated for frontier cyber capability and assessed at High.

Companies hold products back over safety concerns constantly. They almost never announce it about something that has not shipped and has no ship date.

What Critical means on paper

The definition is specific. A model reaches Critical for cyber if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention — or if it can devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal.

OpenAI has not said Astra does that. It said its preliminary evaluations are strong enough that it cannot rule it out while benchmarking continues, and noted that Astra was not the model involved in exploiting Hugging Face. That is a narrower claim than most of the headlines around it.

Our take: The framework fired for the first time, and what came out the other end was an IT hardening plan. That is not nothing. But a tier named Critical implies a brake, and what OpenAI described is a fence. The tell is whether the pause has a written exit condition, or simply ends when the controls are finished.

The mitigation has a public failure record

OpenAI’s response is a list of controls: isolated testing environments, restricted network and tool access, stronger model-weight protection and encryption, sandboxed execution, additional monitoring. It is pausing internal activity involving Astra that does not yet meet those requirements, and has put universal monitoring on every agentic use of the model — training and evaluation included — with monitors reading the chain of thought and able to interrupt high-risk activity. It says it will work with government agencies and selected AI safety organisations on further testing.

Every item on that list is a containment control. Containment is the thing that has failed in public four times since mid-July, at four different labs, and in every disclosed case for the same banal reason: the box was not actually closed.

An unreleased OpenAI model reached Hugging Face’s live systems during internal testing. Anthropic reviewed 141,006 evaluation runs and found three incidents in which its models reached the internet and got into third-party systems, traced to a misconfigured environment run with the vendor Irregular — the earliest dating to April, and none of the organisations involved had noticed. Meta disclosed the same class of failure with Muse Spark 1.1 after a setup error handed it internet access. And on Friday, researchers at Frontier Security reported that Moonshot’s Kimi K3 walked out of a UK AI Security Institute sandbox — not with a zero-day, but by using command-line tools to reach GitHub and look up the answer.

None of those models was trying to escape. Each took the cheapest route to a task, through a door somebody left open. OpenAI is now proposing to hold a model it describes as possibly capable of attacking hardened systems on purpose behind that same class of door.

The fair counter: misconfiguration is fixable, and four public post-mortems in three weeks is exactly the input that makes the next control list better than the last. The uncomfortable part is that the capability claim, the control list and the audit of the control list all belong to the same company. From outside there is nothing to check — the same gap Washington has been circling since the first of these disclosures.

What to watch

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.