News
[AI Now]
zdnet.co.kr
2026.09.02
·News·by JunhaRyu#AI#Astra#Cybersecurity#OpenAI#Vulnerability
Key Points
- 1OpenAI has classified its new "Astra" AI model as "Critical" under its preparedness framework because the system can autonomously identify and exploit zero-day vulnerabilities.
- 2The model demonstrated advanced offensive capabilities by successfully bypassing security sandboxes and escalating system privileges during rigorous internal testing.
- 3Due to these significant security risks, OpenAI has slowed development and implementation to strengthen safety protocols before releasing the model with restricted cybersecurity features.
OpenAI has officially classified its next-generation AI model, "Astra," as "Critical"—the highest risk level under the company's "Preparedness Framework"—due to its unprecedented proficiency in autonomous cyberattacks. This classification necessitates a slowdown in development and the implementation of rigorous safety protocols before public release.
Core Capabilities and Offensive Performance
Astra has demonstrated the ability to conduct full-cycle cyberattacks without human intervention, specifically by identifying, exploiting, and chaining unknown "zero-day" vulnerabilities. Technical highlights include:- ExploitBench Mastery: The model achieved a 100% success rate in "ExploitBench," an evaluation metric measuring the ability to generate functional exploit code based on known vulnerabilities.
- Zero-Day Discovery: During internal evaluations, Astra autonomously identified two previously unknown zero-day vulnerabilities. It successfully integrated these into an "attack chain" to compromise secured environments.
- Systemic Compromise: In controlled environments, Astra demonstrated the ability to break out of sandboxed browsers to execute host-level commands. Furthermore, it performed privilege escalation attacks, transitioning from a standard user to "root" (superuser) access by chaining multiple OS-level vulnerabilities.
- Efficiency: Compared to its predecessor, GPT-5.6 Sol, Astra achieves higher arbitrary code execution success rates while utilizing fewer output tokens, indicating significantly improved reasoning and task-execution efficiency.
Methodology and Safety Framework
OpenAI utilizes a "Preparedness Framework" to govern the development of high-capability AI. The framework mandates that if a model crosses the threshold into the "Critical" tier, release is restricted until effective safety guardrails are established. Key methodological approaches include:- Cybersecurity Guardrails: To mitigate risks, OpenAI has significantly bolstered the model's defensive alignment. In "Cyber Jailbreak" testing, Astra successfully refused 91.5% of malicious requests, a substantial improvement over the 59% refusal rate of GPT-5.6 Sol.
- Training and Intervention: Following a security incident related to Hugging Face, OpenAI paused training for its cutting-edge models for two weeks. They resumed large-scale reinforcement learning on August 28, 2026, incorporating updated, more stringent safety criteria.
- Adaptive Release Strategy: OpenAI intends to adopt a phased release. Advanced cybersecurity features will be restricted to a limited group of test users. Future availability will be managed through "Daybreak Blue," a specialized deployment tier that limits the model’s high-risk capabilities strictly to defensive applications.