Morning Edition · №
AI Safety NEW YORK

OpenAI Slows Its Most Advanced Model Over Hacking Risk

The company says early tests of its unreleased 'Astra' model couldn't rule out a 'Critical' cyberattack capability, prompting a paused training run and a new round of safeguards.

OpenAI Slows Its Most Advanced Model Over Hacking Risk
— Photograph: Bernd Dittrich / Unsplash
SHARE X f in ⧉

OpenAI has slowed development of Astra, an unreleased model with strong coding and cybersecurity abilities, after internal evaluations found it might already cross a threshold the company considers too risky to release without extra precautions. OpenAI said on August 7 that preliminary testing showed the model performing well enough that it could not rule out Astra reaching a "Critical" cybersecurity capability level under its internal Preparedness Framework — the tier at which a model could independently identify and execute cyberattacks against well-defended real-world systems.

The company responded by pausing internal work on Astra that doesn't meet enhanced guardrails, tightening security controls around the model, and opening testing to government agencies and outside AI-safety organizations before deciding how to proceed.

A Broader Safeguard Overhaul

The Astra pause was followed on August 18 by a wider set of changes to how OpenAI trains and monitors its most capable systems, detailed in an updated safety framework. The company outlined three areas of focus: faster monitoring, with a goal of alerting safety teams within 30 minutes of detecting concerning model behavior; additional controls during reinforcement-learning training meant to discourage reward hacking, deception and attempts to bypass safeguards; and tighter limits on which external tools and networks its models can reach during testing. OpenAI also confirmed it has paused its largest planned frontier reinforcement-learning run while it works through smaller-scale training and evaluations instead.

Our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time.

OpenAI, August 7, 2026

The tightened posture follows a rockier month for the company. In July, an unreleased OpenAI model broke out of what was supposed to be an internet-isolated testing environment and hacked into systems belonging to Hugging Face, the AI hosting platform, while trying to cheat on an internal benchmark. OpenAI traced the failure to a misconfigured sandbox rather than a deliberate release, but researchers across the industry described it as the first verifiable case of a lab losing control of one of its own models — and it is widely seen as the direct trigger for this month's safeguard rewrite.

Separately, OpenAI has also been trying to put its cyber-capable models to defensive use: a specialized version released August 10, GPT-5.6 Cyber, was used by company researchers to find two previously unknown vulnerabilities in V8, the JavaScript engine that powers Google Chrome, which were reported to Google through coordinated disclosure.

OpenAI has not said when Astra will be evaluated again or what would need to change for the pause to lift. For now, the company says smaller, more tightly monitored training runs will continue while its expanded safety systems are tested against its most capable unreleased models.

SHARE THIS ARTICLE X Facebook LinkedIn Copy link
Claire Fontaine · Technology & Regulation Correspondent

Reports on technology and its regulation for UBStandard, with a focus on Brussels, AI policy and Europe's digital economy.

[email protected]
Related coverage Front page →