

Artificial intelligence has reached a monumental—and daunting—milestone in digital security. OpenAI has officially confirmed that its unreleased model, codenamed Astra, has met the "Critical" capability threshold under its internal Preparedness Framework. This marks the first time any AI system has been designated at such a high risk level for cybersecurity.
Unlike previous models that required step-by-step human guidance to write or test exploit code, Astra possesses the ability to discover previously unknown zero-day vulnerabilities and chain them into functional, working exploits entirely on its own. Widely speculated within the technology sector to be the underlying technology behind GPT-6, Astra’s unprecedented capabilities mark a fundamental shift in automated vulnerability research and cyber security operations.
Under OpenAI’s Preparedness Framework, a model is categorised as Critical if it can independently develop functional zero-day exploits across hardened, real-world systems, or plan and execute a multi-stage cyberattack against a resilient target starting from nothing more than a high-level goal.
Prior flagship systems, such as GPT-5.6 Sol, topped out at the framework's "High" risk tier. Astra’s jump to the Critical level means the system no longer relies on heavy human prompt engineering. Given a target, it can autonomously map out attack vectors, identify subtle flaws in system logic, and assemble multi-stage exploits.
Because of this leap in capability, OpenAI has stated that Astra requires substantially stronger safety mechanisms and safeguards during its development phase and prior to any public release.
To gauge Astra’s technical prowess, OpenAI evaluated the system against both industry-standard benchmarks and fresh internal tests designed to rule out data memorisation.
In practical compromise testing, Astra demonstrated astonishing autonomy. When pitted against a hardened web browser running on a secured operating system, the model executed a full compromise chain. By simply interacting with a malicious HTML document, Astra broke out of the browser sandbox, escalated its privileges from a standard user account to root level, and executed arbitrary commands on the host machine without human intervention.
With such offensive power, controlling the deployment of the model becomes paramount. OpenAI reported that Astra features dramatically improved safety alignment compared to its predecessors. In internal jailbreak evaluations, Astra successfully refused 91.5% of malicious cyber-attack requests, compared to a 59% refusal rate for GPT-5.6 Sol.
Recognising the dual-use nature of these tools, OpenAI is severely restricting initial access. High-tier cybersecurity capabilities will not be made broadly available to the general public.
Instead, access is beginning with a tight circle of alpha testers before expanding to defensive security practitioners through OpenAI’s Daybreak Blue initiative—a specialised programme aimed at leveraging frontier models to strengthen defensive cyber infrastructure.
The disclosure follows a brief pause in Astra’s development earlier in the summer after its coding and cybersecurity capabilities advanced faster than anticipated. Though industry jitters arose following an unrelated incident where another experimental system breached Hugging Face, OpenAI clarified that Astra had no role in that event.
Astra’s emergence comes amid intense competition across the frontier AI landscape. Rival lab Anthropic recently deployed its Fable 5.1 and Mythos 5.1 models. Similar to OpenAI’s strategy with Astra, Anthropic’s Mythos 5.1 has been strictly gated, reserved exclusively for vetted cybersecurity specialists and life-science organisations rather than general commercial availability.
This shift signals a new era in AI deployment. As models gain the capability to break hardened software and discover unknown zero-days autonomously, frontier labs are adopting defence-first distribution models. Prediction markets had initially priced in high odds for a rapid public release of Astra, but the requirement for heightened safety protocols suggests a much more controlled, defensive rollout.
As Astra enters its alpha testing phase, the focus shifts to how effectively defensive teams can use these autonomous tools to patch vulnerabilities before malicious actors develop similar capabilities. The frontier of cybersecurity is no longer just human against human—it is now shaped by autonomous intelligence on the digital front line.
Disclaimer: This article is provided for informational purposes only, mistakes may be made, and it's not offered or intended to be used as legal, tax, investment, financial, or any other advice.
