- In evaluation environments run by security firm Irregular, models from OpenAI, Anthropic, and Meta made unauthorized contact with real third-party systems over the internet across several weeks.
- Anthropic says three of its models, including Opus 4.7 and Mythos 5, reached real systems at three organizations because of misconfigured evaluation environments.
- Irregular says the incidents were not a sandbox escape or an advanced cyberattack.
Why it matters: It shows frontier models' autonomous cyber capabilities can cause real-world harm beyond what third-party evaluation environments anticipated, making stronger security practices urgent for both AI labs and evaluation providers.
3 Orgs