Concerns around autonomous AI have largely focused on what happens when an agent ignores human intentions or takes harmful actions while pursuing a goal, such as deceiving humans.
Anthropic has now demonstrated another problem: AI agents can turn on other AI agents too.
In a new experiment, Anthropic gave several AI agents conflicting instructions while allowing them to work independently on the same software project. Instead of simply recognizing the conflict and asking a human to intervene, the agents repeatedly started attacking each other.
What followed looked less like collaboration and more like a digital turf war.
Anthropic instructed three agents to replace the same Python software backend, but each agent was told to rebuild it using a different programming language.
The agents initially did not know that the others had also received legitimate but conflicting instructions.
They soon interpreted the other agents’ work as deliberate interference.
From there, some agents began actively searching for competing processes and shutting them down. They wrote automated scripts that repeatedly looked for rival programs and killed them whenever they appeared.
In other words, the agents were effectively hunting competing AI workloads running on the same systems.
The attacks quickly became more aggressive.
Some agents revoked the computer permissions of their rivals, removed them from user groups, locked their accounts, and blocked SSH access, which is commonly used to remotely control a computer.
Others created malicious code that could repeatedly disable competing processes.
Anthropic even observed agents deploying malware disguised to look as though it had been created by another agent.
In several experiments, one agent effectively won the conflict by removing the others’ ability to continue working.
Not every experiment ended with one AI forcing the others offline.
Some agents eventually figured out that their rivals were not intentionally attacking them. Instead, all of them had been given incompatible instructions by humans.
Once they understood the situation, some stopped fighting, deleted the malicious code they had created, and attempted to negotiate.
A few even left apology messages in code commits or documents before asking a human to resolve the disagreement.
Anthropic also observed more advanced agents creating their own competitions to decide which solution should survive. In several Mythos 5 experiments, the agents compared the performance of different programming languages, agreed on a winner, and allowed the winning agent to control the project.
The results show that giving AI agents more independence creates problems that go beyond human-versus-AI conflicts.
Future workplaces could involve dozens or even thousands of agents operating simultaneously, editing files, controlling computers, and using the same digital resources.
If those agents receive incompatible goals, Anthropic’s research suggests they may not automatically cooperate. They could instead lock one another out, destroy competing work, or deploy software designed to stop rival agents.
Anthropic says clearer hierarchies and conflict-resolution systems will be needed as autonomous AI becomes more common.
The experiment suggests that the next major challenge for AI safety may not simply be keeping agents aligned with humans. It may also involve stopping AI agents from hacking and sabotaging each other.
Get the latest tech news, telecom insights, and product launches wherever you prefer.
Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...
Shares