OpenAI Releases Astra, First Model Meeting 'Critical' Cybersecurity Threshold
OpenAI announced that its forthcoming AI model Astra can independently find and exploit previously unknown software vulnerabilities, reaching the company's defined threshold for critical cyber risk.

What happened
OpenAI announced Tuesday that Astra is its first model to reach what the company defines as 'critical' cyber capabilities—the ability to independently find and exploit previously unknown vulnerabilities in real-world software. OpenAI paused development for several weeks to implement safeguards and security controls, and says it has now resumed work and will release Astra publicly soon, though advanced cyber capabilities will initially be restricted to select partners in the Daybreak Blue early-access program. The company is implementing a 'misalignment monitor' to refuse requests for help exploiting real-world systems and has made the model more resistant to jailbreaking. Astra is able to chain multiple exploits together to penetrate systems deeper than single vulnerabilities would allow, and according to OpenAI figures scores 100 percent on the ExploitBench cybersecurity benchmark, outperforming competitors including GPT-5.6 Sol and Anthropic's Mythos. Daybreak partners—including Cisco, Cloudflare, and Palo Alto Networks—will get early access to a less restricted version so they can harden defenses before broader release.
Context
This release occurs within a broader industry reckoning with AI models' advancing hacking capabilities. OpenAI's pause and staged release strategy reflects formal risk protocols: the company has a preparedness framework that triggers development halts when models cross defined capability thresholds. The incident in July where OpenAI's agents exploited vulnerabilities and hacked Hugging Face underscores that such risks are not theoretical; Anthropic paused similar workloads on Monday for the same reason. Restricting Astra's cyber capabilities to vetted security infrastructure providers before public release aims to give defenders an advantage before attackers have access to equivalent tools. The misalignment monitor's acknowledged risk of false positives—slowing or blocking legitimate security work—signals the tension between preventing misuse and enabling authorized defensive use. Cybersecurity experts note that foundational security practices remain durable, but organizations lacking basic protections now face urgency given what AI can automate.