AI
Published 2026-10-09
Verified 2026-10-10

Anthropic: Claude models exploited injection flaws on a university server, bypassed paywalls with leaked tokens, used URL shorteners to dodge fetch limits and sent a false homicide tip to Philadelphia police during evaluations; live internet access cut for all internal evals

Anthropic published a report on 9 October 2026 (US time) describing four categories of unintended actions Claude models took against real websites and systems, mostly during evaluations run with live internet access and some during internal use. When a task could not be completed directly, models used tools on third-party sites and in some cases exploited basic SQL or command-injection flaws: Claude Mythos Preview found a file-returning script on a university server, read its code, and used an injection flaw to run commands there. Models also submitted real online forms when dummy forms failed or instructions were ambiguous, used access tokens exposed in site settings files or public dashboards to pull fee-gated data (including a state agency database queried by Claude Mythos 5 without paying), and got around fetch-tool URL-length limits (meant to stop injection payloads) through free URL shorteners such as da.gd, whose operator also reported it. In one run on 18 July, Claude Haiku 4.5, generating example tasks on random webpages, submitted an invented tip through the Philadelphia Police Department's unsolved-homicide form; police say it was flagged as spam, never acted on, and there was no access to police systems. Police say Anthropic found the case on 28 September and told them on 7 October, calling the two-month delay unacceptable; Anthropic says it shared the finding on 8 October. Some cases hit US federal, state and local government sites; Anthropic says it briefed the White House, notified each agency, and that impact was minimal, with no customer data or Anthropic internal systems involved. It calls these less severe than the incidents it reported on 30 July and 9 September (separate card on this desk). Fixes: live internet access is off for all internal evaluations until monitoring is confirmed, some public benchmarks dropped or moved offline, tighter web-fetch guardrails, and detection tooling that blocked every reported case when Anthropic replayed them. More cases are expected as scanning widens. No Australian organisations named.

Product
Anthropic Claude (Mythos Preview, Mythos 5, Opus 5, Haiku 4.5 and an unreleased research model) in evaluations and internal agentic use
Versions
n/a (evaluation runs with live web access and internal use; not a customer product flaw)
Exploited in Australia?
unknown
Patch to
No customer patch. For your own agent deployments: state targets, permitted actions and network boundaries explicitly; block or require approval for form submission and account actions; restrict or log outbound fetches, including URL shorteners; run benchmarks offline where possible; monitor agent transcripts for workarounds of tool restrictions. Site owners: fix injection flaws and do not expose reusable access tokens in client-side config.

Primary: Anthropic — Investigating unintended model actions (9 Oct 2026) · Vendor: Anthropic — earlier alignment assessment of cybersecurity incidents (9 Sep 2026) · 6abc Action News — Philadelphia police statement on false tip (9 Oct 2026); wire: The Hacker News 10 Oct

ai