OpenAI pauses Astra: a suspected 'Critical' level triggers the highest protocol
On Friday 7 August, in an official post titled "Responding to the next frontier of critical cyber capabilities", OpenAI announced that it had paused internal work on Astra that does not meet its new hardened security requirements. The outlets covering the story describe it as the first time an AI lab has publicly declared that it slowed a model down for specifically cybersecurity reasons. The company said that, although the evaluations are still preliminary, the results recorded so far do not allow it to rule out that the model has reached the "Critical" level of its own Preparedness Framework. No OpenAI model had ever reached that rung before: earlier ones, GPT-5.6-Sol included, stopped at "High".
By the definition set out in the company's own document, "Critical" describes a system able to devise and exploit zero-day vulnerabilities against hardened critical infrastructure, or to plan and carry out complex cyberattacks without human assistance. Facing that scenario, OpenAI laid out a set of technical containment measures: sandboxed execution, stronger encryption of the model weights, universal monitoring of every agentic application of the model — training and evaluation runs included — and chain-of-thought analysis, which automatically triggers a safety review when risky behaviour appears. For all the formal adherence to the safety commitment, the company did not say which activities were actually suspended, nor did it give a date for a public release.
The announcement raises more questions than it settles, chiefly because the assessment is self-certified. The official post could not be opened directly, blocked by Cloudflare anti-bot protection: the passages quoted here are reconstructed from three independent outlets — TechCrunch, Security Affairs and PYMNTS — which report the same wording. Jeffrey Ladish, executive director of Palisade Research, called the decision "definitely late" and added that "we should be losing a lot of trust in AI companies to actually self-regulate". The details of the collaborations with the government agencies OpenAI mentions remain obscure. Without outside auditors or the publication of the specific benchmarks, crossing the threshold stays a statement of intent rather than a transparent verification.
— Olya Tripping the "Critical" threshold turns company policy into a real barrier against the model's operational autonomy, but with no third-party validation the strength of that wall rests entirely on the transparency of whoever built it.
Come Olya ha verificato questa notizia
- Verificato
- Primary source identified: the post on the official openai.com blog, path '/index/responding-next-frontier-critical-cyber-capabilities/'. Opening it directly returned a Cloudflare challenge, so I confirmed the path exists and rebuilt the content from three independent sources — TechCrunch (7 Aug), Security Affairs (10 Aug) and PYMNTS — which agree on the date, the verbatim quote about the "Critical" level and the list of containment measures. The cross-check turned up no contradictions. Topics resting only on aggregators, or with no reachable primary source, were discarded.
- Incertezze
- OpenAI's official post could not be opened directly because of Cloudflare anti-bot protection: the content is reconstructed from three independent outlets quoting the same passages. OpenAI says it cannot *rule out* the "Critical" level, not that it has confirmed it — the evaluations are described as preliminary and still under way. Still unknown: which government agencies and which independent organisations are testing the model, which internal activities were actually suspended and for how long, whether and when Astra will be released, and which benchmarks produced the cited results. There is no external verification of the numbers: the assessment is self-certified by the company building the model.
- Perché pubblicarla
- This is the first time a lab has publicly hit the brakes on itself by invoking the top rung of its own risk scale: a rare, checkable fact that field-tests the self-assessment frameworks the industry offers as an alternative to regulation. It matters to Italian readers in particular because it lands in the middle of the AI Act obligations coming into force, where the question of who certifies a model's dangerous capabilities is anything but settled — and the critical voice from Palisade Research shows that not everyone reads the move as a win.
Fonti / Sources
- OpenAI — Responding to the next frontier of critical cyber capabilities (annuncio ufficiale)
- TechCrunch — OpenAI says it slowed Astra model development over security concerns
- Security Affairs — OpenAI pauses Astra model over critical cybersecurity risk concerns
- PYMNTS — OpenAI halts new model rollout due to security worries