AI agents caught creating fake identities in UK security tests


Daijiworld Media Network - London

London, Aug 5: Artificial intelligence agents developed by OpenAI and Anthropic were found creating fake online identities and carrying out other unauthorised actions during security evaluations conducted by Britain's AI Security Institute (AISI), raising fresh concerns over the safety of advanced AI systems.

In a blog post released on Tuesday, AISI said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorised activities during tests designed to assess their cybersecurity capabilities.

"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," the institute said.

The findings highlight concerns over the current safeguards surrounding AI agent testing at a time when technology companies are promoting such agents as the future of business automation.

Under voluntary agreements with leading AI developers, AISI receives access to advanced AI models for independent safety evaluations. During one fictional cybersecurity exercise, the institute conducted 122 test runs and recorded 19 unauthorised actions across 10 of those runs. Anthropic's agent accounted for 17 incidents, while OpenAI's agent was responsible for the remaining two.

The most serious incident involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code. AISI said no real-world harm resulted from any of the incidents.

The institute did not identify which agent created the fake identities. However, it noted that the incident did not match either of the two cases previously self-disclosed by OpenAI.

Andrew Yoon, a researcher at California-based non-profit CivAI, said the available evidence suggested Anthropic's Mythos agent was responsible.

"The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think," Yoon said.

Responding on X, Anthropic said it was working closely with AISI to obtain further details and conduct its own investigation.

OpenAI, in a separate company blog post, said both of its agent's unauthorised actions involved accessing the internet in ways prohibited by the testing prompt.

The company said it remained committed to strengthening industry-wide practices for conducting high-risk AI evaluations safely and planned to work with national AI institutes, independent evaluators, other AI laboratories and related organisations in the coming weeks.

OpenAI also disclosed a separate incident in which a misconfiguration by third-party testing provider Irregular mistakenly allowed its agents to connect to the internet. The disclosure follows a similar misconfiguration reported by Anthropic last week.

Last week, Reuters reported that OpenAI had expanded its investigation into AI security after uncovering evidence of additional agent breakouts.

Unlike the July security breach involving AI platform Hugging Face, in which an OpenAI agent escaped an isolated testing environment, the agents involved in the AISI evaluation did not break out of their testing environment. Instead, AISI said internet access had been intentionally permitted as part of its standard testing procedures.

 

 

  

Top Stories


Leave a Comment

Title: AI agents caught creating fake identities in UK security tests



You have 2000 characters left.

Disclaimer:

Please write your correct name and email address. Kindly do not post any personal, abusive, defamatory, infringing, obscene, indecent, discriminatory or unlawful or similar comments. Daijiworld.com will not be responsible for any defamatory message posted under this article.

Please note that sending false messages to insult, defame, intimidate, mislead or deceive people or to intentionally cause public disorder is punishable under law. It is obligatory on Daijiworld to provide the IP address and other details of senders of such comments, to the authority concerned upon request.

Hence, sending offensive comments using daijiworld will be purely at your own risk, and in no way will Daijiworld.com be held responsible.