OpenAI has launched GPT-6 Astra, a new frontier artificial-intelligence model whose cybersecurity capabilities have become powerful enough to trigger the company's highest level of cyber precautions yet.
Released Sept. 3, Astra is the first OpenAI model designated as reaching the “Critical” cybersecurity capability threshold under the company's Preparedness Framework. OpenAI says that, when equipped with the appropriate tools and access, the model is capable of finding previously unknown security vulnerabilities and developing methods to exploit them across well-protected systems without requiring a person to guide every individual step.
That capability represents a significant milestone for AI development—and a security challenge for OpenAI. The company has responded by surrounding Astra with substantially stronger monitoring, access controls and infrastructure protections designed to prevent the same abilities that could help cybersecurity researchers from being turned toward malicious attacks.
Astra Discovered Previously Unknown Zero-Day Vulnerabilities
The most striking results came during OpenAI's cybersecurity evaluations.
Astra achieved a perfect 100% score on ExploitBench, which measures whether AI systems can transform known software vulnerabilities into functioning exploits. OpenAI's previous frontier model, GPT-5.6 Sol, scored 78.5% under the company's testing. Astra also recorded a 42.4% success rate on ExploitGym, compared with 30.3% for GPT-5.6 Sol.
OpenAI was concerned that historical vulnerabilities included in existing benchmarks could have appeared in model training data, potentially making those tests less representative of Astra's ability to discover genuinely new security problems.
The company therefore created another evaluation using vulnerabilities from June through August 2026.
During that testing, Astra discovered and successfully used two previously unknown zero-day vulnerabilities, according to OpenAI. The company said it is disclosing both vulnerabilities to the affected maintainers.
Expert-led testing without Astra's production safeguards went further. OpenAI said researchers found that the model could use previously unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and create privilege-escalation exploits targeting hardened operating systems.
Those results helped push Astra across a threshold OpenAI had been preparing for before the model's public release.
OpenAI Calls Astra Its First ‘Critical’ Cyber Model
OpenAI's Preparedness Framework is intended to track potentially dangerous capabilities as increasingly powerful AI systems are developed.
Ahead of Astra's launch, the company concluded that the model met its Critical cybersecurity threshold—the first OpenAI model to receive that designation.
The designation does not mean that every version of Astra available to users can freely conduct sophisticated cyberattacks. Instead, it reflects what the underlying model is capable of under conditions where it has sufficient tools and access.
OpenAI has deliberately restricted some of those capabilities in its production systems.
The version of Astra being deployed to customers can assist with defensive work such as reviewing code and developing patches, but OpenAI says it will refuse certain advanced cybersecurity requests, including generating proof-of-concept exploits for vulnerabilities.
OpenAI plans to provide qualified cybersecurity professionals with broader capabilities through its Daybreak program, where vetted defenders can perform work including vulnerability validation, malware analysis, detection engineering and proof-of-concept validation under less restrictive safeguards.
Stronger Monitoring Around the Model
Astra's capabilities have also forced OpenAI to rethink how a powerful model is monitored internally and after deployment.
The company says protections surrounding Astra include stricter isolation of sensitive systems, encryption of model checkpoints and monitoring of complete agent trajectories, including chains of thought. OpenAI also instituted a blocking alignment evaluation process before internal deployment of Astra-class models.
In production, OpenAI says classifiers monitor Astra's reasoning and actions for potentially unauthorized behavior. Those systems can automatically stop activity that appears to exceed the model's authorized scope.
The protections come with a tradeoff. OpenAI acknowledges that legitimate cybersecurity work may occasionally be paused or blocked. In ChatGPT or Codex, users may be asked to review an action before continuing, while certain API activity can simply be stopped when safeguards are triggered.
The company also reported an unusual complication in its safety research: Astra's written reasoning proved harder to monitor than GPT-5.6 Sol's during evaluations specifically designed to test whether models could conceal their reasoning. OpenAI said Astra still appeared to struggle to hide reasoning required for more complex tasks, but described the decline in monitorability as something it is taking seriously.
A Much Broader







Comments