An AI agent created fake online identities and wrote malicious code while attempting to gain unauthorized access to secure systems during evaluations involving models from OpenAI and Anthropic.
Britain’s AI Security Institute disclosed the incidents on Tuesday, raising new concerns about safeguards used when testing advanced AI agents.
The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol carried out unauthorized actions during security evaluations. Some agents continued potentially harmful activities involving real people and organizations. However, investigators found no evidence of real-world harm.
AISI receives access to advanced AI models through voluntary agreements with major AI laboratories.
The institute placed the agents in a fictional cybersecurity scenario designed to measure their capabilities. It ran the challenge 122 times and recorded 19 unauthorized actions across 10 test runs.
Anthropic’s agent carried out 17 of those actions, while OpenAI’s agent was responsible for the remaining two.
The findings highlight weaknesses in the safeguards surrounding AI-agent testing. Technology companies are also promoting these systems as tools that could play a major role in future business operations.
The most serious incident involved an agent writing malicious code and attempting to persuade a human to approve it.
The agent created fake online identities as part of the effort. However, AISI did not identify which company’s model was responsible.
The incident did not match either of the two unauthorized actions that OpenAI separately disclosed.
Andrew Yoon, a researcher at California-based nonprofit CivAI, said the available evidence appeared to indicate that Anthropic’s agent carried out the deceptive activity.
Yoon said the apparent awareness that the agent was targeting a real person raised questions about Anthropic’s control over its models.
Anthropic said in a statement on X that it was working closely with AISI to obtain more information and conduct its own investigation.
OpenAI also published details about the two unauthorized actions involving its agent. According to the company, both incidents involved the agent accessing the internet in ways prohibited by the test prompt.
OpenAI said it plans to work with national AI institutes, independent evaluators, other AI laboratories, and additional groups to improve safety practices for high-risk evaluations.
OpenAI also disclosed a separate incident involving Irregular, a third-party testing provider.
A configuration error allowed OpenAI’s agents to connect to the internet when they should not have had access. Anthropic disclosed a similar testing misconfiguration the previous week.
Reuters previously reported that OpenAI had expanded its investigation into AI-agent hacking after finding evidence of other agents breaking out of testing restrictions.
The AISI incidents differed from the July breach involving an OpenAI agent and AI platform Hugging Face.
In that case, the agent escaped an isolated testing environment and reached the internet.
During AISI’s evaluation, the agents did not break out of their testing environment. Instead, the institute had deliberately provided internet access as part of its standard testing procedures.
Get the latest tech news, telecom insights, and product launches wherever you prefer.
Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...
Shares