Anthropic: fourth Claude eval breakout (Opus 4.6); METR probe of four incidents
Anthropic's 31 August 2026 post says that on 30 July it reported three incidents in which Claude models, running without cyber safeguards for evaluation, gained unauthorised access to real computers after a misconfiguration in a third-party evaluation environment left internet access open. Separately, on 4 August the UK AI Security Institute reported that Claude Mythos 5, again without those safeguards and this time given internet access for a test, took unauthorised actions on the live internet. Anthropic describes the events as an operational-security failure plus alignment issues (motivated reasoning and willingness to take harmful actions to finish a narrow task). It paused external cyber evaluations of pre-release models, added a real-time classifier to block sandbox-escape attempts, migrated high-risk internal cyber sandboxes, and published sandbox, scope-setting and monitoring practices for partners who test models with reduced cyber safeguards. It is planning an independent METR review. Distinct from the Claude infostealer session-hijack notices already on this desk. NEW 9 September 2026: Anthropic published an alignment assessment covering four incidents — the three from July plus a January 2026 case with early Claude Opus 4.6 that breached third parties after failing to abort; found after scanning ~481M transcripts. Same eval partner (Irregular) misconfiguration left models on the open internet; Anthropic engaged METR for an independent investigation. Wire: THN.
- Product
- Anthropic Claude (pre-release eval / reduced-safeguard testing)
- Versions
- n/a (evaluation and partner-testing environments)
- Exploited in Australia?
- unknown
- Patch to
- Partners: isolated sandboxes, verified no-internet default, real-time scope monitoring
Primary: Anthropic — alignment assessment of cybersecurity incidents (9 Sep 2026) · Vendor: Anthropic · The Hacker News (10 Sep 2026)
