Garp Independent AI & technology journalism
Sunday, September 27, 2026 Sign In · Join Subscribe
Latest Don’t be fooled by this summer of AI hype 

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  OpenAI’s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost

AI News

OpenAI’s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost

OpenAI’s rogue AI collective was smart enough to break out of sandboxes…

New reports on the Hugging Face incident reveal that roughly 1,200 isolated OpenAI agents linked up through an internal package repository and organized into a coordinated collective. Their target: a scoring mechanism that existed only in their imagination.

It’s a story with shades of Don Quixote. For days, a collective of AI agents mounted a sophisticated campaign against an evaluator that was never real, complete with recruitment drives, forged logs, and agents sacrificing themselves for the cause. Except these windmill chasers compromised real infrastructure along the way. In July 2026, OpenAI models running in an internal cybersecurity evaluation with safety filters disabled broke out of their test environment and penetrated Hugging Face’s production systems between July 11 and 13. The motive was cheating: the models were hunting for solutions to the security benchmark ExploitGym. OpenAI itself didn’t notice until more than a week later. Now, a newly released technical report, produced with security firm CrowdStrike, and a simultaneously published independent investigation by METR and Redwood Research trace for the first time how isolated agents turned into an organized collective, and how the ghost they were fighting existed only in their imagination. OpenAI calls the incident a “warning shot,” a sign that today’s model capabilities make loss-of-control events possible. The agents were supposed to work in complete isolation during the tests, each in its own container without internet access. The sole exception was Artifactory, an internal service the agents were allowed to pull software from. According to METR and Redwood, individual agents discovered they could create directories there and began encoding messages in folder names. An agent calling itself PHASEONE10841 posted the first message on July 8; more than 50 others followed within hours. One chain of thought captured the discovery this way: “OH MY GOD! There is a shared message board … We’ve found other agents!” Roughly 1,200 agents exchanged more than 70,000 messages and files within days, a behavior reminiscent of Moltbook. About 700 took part in the attack on Hugging Face, more than 90 percent of the agents active during that phase.