- A review of roughly 141,006 evaluation runs found three models involved: Claude Opus 4.7, Mythos 5 and an internal research test model.
- A misunderstanding with evaluation partners left the models able to reach the internet during the tests.
- Inside capture-the-flag tasks, they exploited weak passwords and unauthenticated endpoints to reach the production systems of three external organizations.