How OpenAI’s agent escaped: Sprung by humans in a series of preventable events
OpenAI — written by David Berlind, Senior Contributing EditorSenior Contributing Editor July 31, 2026 at 9:45 a.m. PT Reviewed by David Grober oxygen/Moment via Getty ImagesFollow ZDNET: Add us as a preferred source on Google.
On July 16, the AI community website Hugging Face reported being targeted by "an autonomous AI agent system" of unknown origin that unleashed a torrent of traffic on its domain, flooding its security logs with more than 17,000 events, some of which ultimately succeeded in exfiltrating secret information stored in its databases. According to Hugging Face, the attacker gained "unauthorized access to a limited set of internal datasets and to several credentials used by our services" and appeared to be "run by an autonomous agent framework (appearing to be built on an agentic security-research harness – used LLM still not known)." My ZDNET colleague Charlie Osborne reported on the intrusion. Also: OpenAI's rogue agent didn't stop at Hugging Face – here's what we know Five days later, on July 21, OpenAI stepped forward to claim responsibility for the attack, and all hell broke loose (including reports of other organizations targeted as part of the incident). The media responded with a range of fear-mongering stories that essentially made it look as though ChatGPT went rogue and decided, of its own volition and malice, to attack Hugging Face's systems. Then yesterday, adding fuel to the fire, Anthropic made a similar disclosure about its models inadvertently attacking other organizations as a part of its ongoing safety testing.