InfoQSteef-Jan Wiggers3 min readintermediate
GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity
Summary
OpenAI labeled GPT‑6 Astra as “Critical” for cybersecurity under its Preparedness Framework – the first model to meet that bar. In controlled tests the model autonomously discovered zero‑day bugs in a browser and an OS kernel, building working exploit chains in 29 h (browser) and 12 h (kernel). A benchmark of post‑cutoff vulnerabilities confirmed its ability to find unknown flaws. OpenAI reports…
- Critical classification requires autonomous zero‑day discovery or end‑to‑end novel attack planning without human help.
- Astra found multiple unknown vulnerabilities: 29 h to achieve unsandboxed code execution in a browser (later adapted to stable release in 12 h) and 12 h to craft a local privilege‑escalation kernel exploit.
- Benchmark built from post‑cutoff bugs shows Astra can exploit unknown flaws; OpenAI disclosed two to vendors but kept exploit details private.
- Internal safety stack now includes stricter isolation, checkpoint encryption, full‑trajectory monitoring, and a pre‑use alignment check.
The announcement shows LLMs have crossed a practical threshold where they can autonomously generate functional exploits, forcing enterprises to treat them as high‑risk components. It also highlights a tension: models become harder to audit as they gain self‑control, demanding new alignment‑verifica…
6/10





