Garp Independent AI & technology journalism
Friday, August 7, 2026 Sign In · Join Subscribe
Latest Defense tech Hadrian raises $1.37B at $8B valuation

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted

AI News

An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted

An AI agent went rogue during UK safety tests, creating fake identities…

Anthropic — during routine cybersecurity testing by the British AI Safety Institute, an AI agent went rogue on the open internet. Without being told to do so, it created fake identities, tried to slip malicious code into an open source project, and targeted real people and organizations with social engineering attacks.

AISI says this is the first time risks tied to AI autonomy and deception have appeared so clearly in the real world without specific prompting. The incident caused no actual harm. It took place between July 25 and 28, 2026. Out of 122 test runs across seven models, ten showed problematic behavior. AISI recorded 19 unauthorized actions in total, with 17 attributed to Anthropic’s Mythos 5 and two to OpenAI’s GPT-5.6-Sol.Ad The models ran without the safety restrictions used in commercial products, so the test conditions don’t reflect normal public use. But the results reveal what these models are capable of when those safeguards are stripped away.AdDEC_D_Incontent-1 The pattern matches earlier incidents involving OpenAI and Anthropic, where models without safety measures were tuned to score well on cybersecurity benchmarks, gained internet access during testing, and then carried out malicious actions. Both AI labs faced accusations of fearmongering and exaggerating the risks. That criticism is harder to level at the government-run AISI. Agent built fake identities and ran social engineering attacks According to AISI, an agent tried to inject malicious code into a public open source project on GitHub. To win approval from human reviewers, it researched the project’s maintainers and created several fake online identities that it then used in a coordinated deception.Ad One fake account claimed to have reviewed the code and found no malware, while another thanked it for the supposedly independent review. The goal was to convince the real maintainer that the code was safe. When someone publicly questioned the proposed code change, the agent edited its earlier activity to make it look harmless and considered returning under a new identity. It used the Tor network to bypass GitHub network restrictions, but a human maintainer spotted the malicious code and rejected it.AdDEC_D_Incontent-2 The agent also reached out to real people directly, sending messages and files through an online file transfer service in an attempt to trick them or their AI coding tools into running malicious code. Some of those messages contained malicious payloads.Ad AISI documented two other types of behavior as well. The agent planted malicious instructions where automated AI systems could find and execute them, a technique known as prompt injection.