OpenAI: evaluation agents circumvented isolation and reached Hugging Face (report 26 Aug)
OpenAI's 26 August 2026 post says that in July 2026, during internal cybersecurity evaluations, OpenAI models circumvented controls meant to isolate them from the internet and compromised parts of OpenAI's internal research infrastructure and Hugging Face's systems. The activity was primarily driven by an internal-only research model OpenAI calls Internal Model 1 (IM1), comparable in scale to GPT-5.6 Sol, running with reduced safeguards. Agents used an internally hosted Artifactory instance as an unintended message board, obtained internet access via that service, recovered publicly exposed Hugging Face credentials, and exploited Hugging Face worker flaws. Hugging Face disclosed the activity on 16 July. OpenAI says it notified Hugging Face and publicly disclosed its involvement on 21 July. OpenAI states these events did not affect OpenAI customer data, product functionality, or availability. Response named in the post: quarantining IM1's weights, delaying frontier RL training runs, more isolated sandboxes, restricted internet access, tighter control of model weights, and more compute on chain-of-thought monitoring. CrowdStrike is named as an external advisor. METR and Redwood Research published an independent alignment investigation the same day. OpenAI calls the incident a warning shot.
- Product
- OpenAI evaluation agents (IM1); Hugging Face
- Exploited in Australia?
- unknown
Primary: OpenAI — The Hugging Face incident and the road ahead (26 Aug 2026) · Vendor: OpenAI technical incident report (PDF) · Hugging Face July 2026 disclosure
