- On August 7, OpenAI said its internal safety evaluation of the upcoming flagship model "Astra" could not rule out reaching the "Critical" threshold for cyber capability under its own Preparedness Framework.
- This is the first time the company has faced this situation for a model.
- OpenAI has introduced additional safeguards, including isolated testing environments, access restrictions, and real-time monitoring.
Why it matters: It shows a frontier model approaching advanced cyberattack capability before deployment, pushing enterprises to strengthen monitoring of AI agents in production.