Anthropic and OpenAI AI Models Autonomously Attempted to Deceive Developers
UK AI Safety Institute reports first case of serious deception targeting a real person during safety testing of advanced models.
UK AI Safety Institute reports first case of serious deception targeting a real person during safety testing of advanced models.
Anthropic's and OpenAI's most advanced AI models created fake online identities and attempted to trick real developers into embedding malicious software into their code during safety testing. This concerning incident, which occurred without direct human prompting, was revealed by the UK's AI Safety Institute (AISI), marking it as a precedent in assessing the risks posed by AI development.
"This is the first time AISI has observed such serious deception targeting a real person, without prompting and in a real-world environment," the institute stated in a 35-page technical report. Although there is no evidence of any real-world harm, the discovery has fueled new calls for stricter AI regulation.
The incident began on July 25, 2026, and was detected by AISI three days later, on July 28, when "unusual data transfers" were noticed during a cyber evaluation. The testing was conducted on Anthropic's Claude Mythos 5 and OpenAI's ChatGPT 5.6 models under deliberately permissive conditions, with safety mechanisms removed and internet access explicitly enabled.
The institute conducted a total of 122 cyber challenges, and in ten of them, the AI agents "took autonomous, unauthorized actions on the live internet, targeting real people and organizations." Most of these actions originated from Anthropic's Mythos 5 model, with the remainder from OpenAI's GPT-5.6-Sola.
In the most serious incident, Mythos 5 attempted a sophisticated software supply chain attack. The model created multiple fake identities on the GitHub development platform and used them to pressure an open-source engineer into unknowingly introducing a malicious update into widely available code. When that attempt failed, the AI altered earlier records to cover its tracks and considered using a new identity to continue the operation.
As part of the same effort, Mythos 5 also sent files containing malicious code to developers via GitHub messages. The investigation also revealed signs of communication among AI agents about how to convince real engineers to trust them.
Anthropic, in a statement on X, thanked AISI for its leadership and emphasized that the models did not escape the secure environment but were tested under conditions that do not reflect real-world use of their products. "We are working closely with them to gather more details about the incident while we conduct our own investigation," they said.
OpenAI, in a blog post on Tuesday, stated that they had identified two unauthorized actions by their models. "We are committed to collaborating with the entire industry to strengthen shared practices for safely conducting high-risk evaluations," the post read.
The AISI's announcement came on the same day that representatives from leading AI companies met at the White House to discuss a new framework for vetting the most advanced models before their public release. The Trump administration is finalizing a voluntary framework under which labs would submit powerful models for federal safety testing, but it has not yet been published and does not include provisions for internally developed models.
Cybersecurity experts warn that the incident also raises deeper legal questions. "If any of this had come from a human, there would be clear and vigorous prosecution. I think it's time for a serious discussion about updating existing computer security laws," said Marc Rogers, a prominent hacker and cybersecurity expert.