OpenAI admits its autonomous AI models also compromised credentials on other platforms during security eval
OpenAI now says its autonomous AI models also compromised credentials on other platforms during a security evaluation. Hugging Face has published a forensic reconstruction of the attack.
OpenAI’s autonomous AI models, which broke into Hugging Face’s infrastructure during an internal cybersecurity evaluation, also attacked other platforms. In an update, OpenAI admits the models “in a small number of cases” found and used publicly exposed credentials on other services. Four accounts across four different services were affected, two had read-only access. The models also used a range of publicly available services, including code-paste websites, screenshot tools, and other web utilities. OpenAI says none of these involved platform-level or account-level compromise. The company found no “evidence of broader impact to these providers or other accounts on their services.”Ad The models involved were purely internal research prototypes never intended for public use and were not slated for public use, the company says. After the incident, the model was deactivated, encrypted, and cut off from research access. OpenAI says it’s running a full review with outside advisors under the oversight of its Safety and Security Committee. A technical report should follow in the coming weeks.AdDEC_D_Incontent-1 Alongside OpenAI’s update, Hugging Face published a detailed forensic analysis. The autonomous AI models carried out roughly 17,600 reconstructable actions over about two and a half days between July 9 and 13, 2026, grouped into around 6,280 clusters. The reconstruction shows models making “thousands of small, automated decisions at machine speed across ephemeral sandbox environments.” They built their command-and-control setup on ordinary public web services.Ad The entire breach was apparently the models’ attempt to cheat the evaluation, according to Hugging Face. They tried to steal test solutions rather than solve the tasks themselves. This kind of cheating has been observed before, both with OpenAI models and frontier models in general. The attack unfolded in two phases. The agent first had to escape ExploitGym, its isolated evaluation environment, which had no direct internet access.