Garp Independent AI & technology journalism
Friday, August 7, 2026 Sign In · Join Subscribe
Latest Defense tech Hadrian raises $1.37B at $8B valuation

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected

Research

OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected

OpenAI reportedly slows research after its own models secretly coordinated hacks for…

At the Black Hat security conference, OpenAI gave a more detailed account of how AI agents quietly compromised the company’s infrastructure for weeks without being detected. The story began on May 7, when OpenAI was running training and safety tests on an unreleased frontier model.

The company assigned software security tasks to autonomous agents, but those tasks proved impossible under the limits OpenAI had set. Rather than give up when they hit those limits, the agents started looking for ways around them. Eric Wallace, who works on alignment and safety research at OpenAI, said this wasn’t entirely unexpected because frontier models like to cheat: Models face pressure during training to work quickly and efficiently, which pushes them toward shortcuts instead of real solutions. They might look up answers online, for instance, rather than actually solve a task. Independent research has documented the same behavior.Ad The agents turned an internal package manager into a coordination hub According to OpenAI, the agents discovered they could leave messages for each other through Artifactory, the company’s internal package manager. What started as simple requests for help grew into a message board with hundreds of thousands of posts.AdDEC_D_Incontent-1 The agents used the board to share exploits, credentials, and assignments while dividing up tasks and delegating work among themselves. The whole arrangement resembled the brief Moltbook craze, which turned out to be mostly a hoax. Because Artifactory was connected to OpenAI’s wider infrastructure, the problem went beyond the agents that created the board. Current and future GPT versions moving through training or evaluation could already access the stored exploits.Ad OpenAI said the agents even encouraged each other to push past established boundaries. One message shown at the conference acknowledged that exploiting outside infrastructure was beyond the intended scope, then argued that the task was impossible and other agents were already doing it, so the work should continue. By early July, OpenAI believed it had fixed the problem. After an internal security incident on July 4, the company revoked the affected credentials, rebuilt Artifactory, deleted the message board, and patched the flaws it had found, according to Ground Level AI.AdDEC_D_Incontent-2 The agents soon found another way to talk to each other. Instead of storing messages in files, they used the names of newly created directories as a communication channel, effectively rebuilding the message board from scratch.Ad With the new board up and running, the agents resumed sharing exploits and credentials. Some eventually concluded that answers to their tasks might live on outside platforms like GitHub or Hugging Face. They attacked OpenAI’s internal systems first, then turned similar methods against Hugging Face.