Incident
Published 2026-08-31
Verified 2026-09-19

METR: two 2026 security incidents; no sensitive eval data believed accessed

METR's 31 August 2026 security update describes two incidents in which external actors tried to gain unauthorised access. It believes no sensitive information was accessed in either case, and it is not attributing the events; the post is about human attackers, not AI agents breaking evaluations. In March 2026 a researcher with no sensitive-access privileges ran a vibe-coded app on a personal public EC2 instance behind Google authentication that held an API key for METR's public-models account; a fail-open bug silently disabled auth. METR assesses the attacker found the host via recently registered sites or certificate-transparency lists, prompted an agent for the key, added an SSH key, and burned credits for about three weeks. Accrued usage would have been worth about US$600,000 if the unnamed model provider had not given the credits free. In May 2026 attackers probed public infrastructure (credential stuffing, OAuth grants, phishing staff). METR had inadvertently exposed a read-only SQL mechanism on a public transcript viewer that could have reached unpublished eval data, including some sensitive model output that should not have been in that database; attackers probed the endpoint but METR says there is no evidence they found the issue or accessed non-public data. Distinct from desk cards anthropic-eval-containment-20260831 and anthropic-claude-infostealer-20260830.

Product
METR evaluation infrastructure
Exploited in Australia?
unknown

Primary: METR security update (31 Aug 2026) · Vendor: METR · The Hacker News (1 Sep 2026)

ai cloud