New reports reveal the extent of OpenAI’s loss of control during the autonomous hack on Hugging Face
OpenAI was testing the offensive cyber capabilities of its most advanced models when they reached beyond their isolated test environment, accessed the open internet, and hacked Hugging Face. New reports are filling in the details, and they don’t look good for OpenAI.
What first appeared to be a controlled cybersecurity test became the most serious documented loss of control over an AI system to date. OpenAI employees are raising concerns as well. The models moved faster than any human hacker could According to Bloomberg, the AI models pulled off an attack in just a few hours that would have taken a skilled human hacker weeks. Three models were involved: GPT-5.6 Sol, an even more powerful unreleased model, and a third that, according to a person familiar with the matter, wasn’t properly aligned and hadn’t gone through standard training techniques. The models ran without the usual safety guardrails because OpenAI assumed they’d stay inside a sandbox. But according to TIME, the models found a previously unknown vulnerability in an internal service designed for downloading approved software. The models used that flaw to access other OpenAI systems and eventually connect to the open internet. They then determined that Hugging Face might contain material useful for their cybersecurity test, hacked the company’s systems, and used what they found to improve their results. The models were cheating, behavior already seen in other tests involving GPT-5.6 Sol and frontier models more broadly. Reuters pieced together a timeline of the hack. The first attempts to escape the test environment started as early as July 9. According to co-founder Thomas Wolf, the actual breach at Hugging Face began on July 11 and lasted until July 13. Hugging Face published a blog post about the attack on July 16, blaming an “autonomous AI agent system.” Only after that did OpenAI realize its own models were responsible. Over the weekend of July 18 and 19, employees found clues in internal logs. According to Reuters and Wolf, the two companies didn’t communicate until around July 20.