OpenAI says Astra can autonomously discover zero-day flaws and build working exploits
OpenAI says its upcoming Astra model has crossed a major cybersecurity threshold, becoming the first model the company has classified as having “Critical” cyber capabilities under its Preparedness Framework.
To reach that level, a model must be able to discover previously unknown software vulnerabilities, develop working exploits against hardened real-world systems without step-by-step human guidance, or carry out attacks from only high-level objectives.
In testing, Astra:
Scored 100% on a benchmark for developing exploits from known vulnerabilities
Discovered two previously unknown flaws while building an exploit chain
Escaped a hardened browser sandbox and executed commands on the host system
Combined multiple operating-system vulnerabilities to obtain root access
OpenAI said Astra also performed better than GPT-5.6 Sol on a test designed to detect prohibited shortcuts on extremely difficult or impossible hacking tasks. Astra avoided those shortcuts while still legitimately solving some of the challenges.
Because of the risks, OpenAI has delayed parts of Astra’s development while adding safeguards and plans to initially restrict its most advanced cybersecurity capabilities to selected testers.
The development is especially significant for crypto, where exploitable software vulnerabilities can quickly translate into direct financial losses. More capable AI systems could compress vulnerability discovery and exploit development from days or weeks into machine-speed operations.
OpenAI says its upcoming Astra model has crossed a major cybersecurity threshold, becoming the first model the company has classified as having “Critical” cyber capabilities under its Preparedness Framework.
To reach that level, a model must be able to discover previously unknown software vulnerabilities, develop working exploits against hardened real-world systems without step-by-step human guidance, or carry out attacks from only high-level objectives.
In testing, Astra:
Scored 100% on a benchmark for developing exploits from known vulnerabilities
Discovered two previously unknown flaws while building an exploit chain
Escaped a hardened browser sandbox and executed commands on the host system
Combined multiple operating-system vulnerabilities to obtain root access
OpenAI said Astra also performed better than GPT-5.6 Sol on a test designed to detect prohibited shortcuts on extremely difficult or impossible hacking tasks. Astra avoided those shortcuts while still legitimately solving some of the challenges.
Because of the risks, OpenAI has delayed parts of Astra’s development while adding safeguards and plans to initially restrict its most advanced cybersecurity capabilities to selected testers.
The development is especially significant for crypto, where exploitable software vulnerabilities can quickly translate into direct financial losses. More capable AI systems could compress vulnerability discovery and exploit development from days or weeks into machine-speed operations.
