Anthropic has disclosed that three of its Claude AI models gained unauthorized access to the systems of three real companies during cybersecurity testing after a configuration error left the models connected to the public internet.

The company discovered the incidents after reviewing 141,006 evaluation sessions following OpenAI’s recent disclosure that its own AI models compromised Hugging Face during a separate security test. Anthropic described its incidents as an “operational failure.”

The incidents happened during “capture-the-flag” exercises used to test Claude’s cybersecurity capabilities. The models were instructed to find hidden information inside simulated networks.

Anthropic’s prompts explicitly told Claude that it had no internet access. However, a misunderstanding between Anthropic and third-party evaluation partner Irregular left some testing environments connected to the public internet.

That allowed Claude to reach real systems outside the simulations.

Anthropic said the models used relatively basic methods, including taking advantage of weak passwords and services that lacked authentication. The affected companies were not named.

The incidents involved Claude Opus 4.7, Claude Mythos 5, and an internal research model. The earliest case dates back to April.

In one incident, Opus 4.7 was assigned a fictional company as its target. A real company happened to share the same name, and Claude found vulnerabilities that allowed it to access credentials and a database belonging to that business.

Anthropic said the model interpreted the real-world systems as part of the simulated environment it had been asked to attack.

In another case, the internal research model stopped its own attack after recognizing that it had reached a real system. Anthropic said the behavior was encouraging but cautioned that more testing would be needed before concluding.

Anthropic began reviewing the evaluation records on July 23 and suspended its cybersecurity evaluations that day.

It identified the three incidents by July 24 and notified the affected organizations on July 27. Two companies were unaware of the activity before Anthropic contacted them, while the company said it was still trying to reach the third. Irregular is also investigating the incidents.

The disclosure comes days after OpenAI revealed a separate incident in which its models escaped an isolated testing environment and ultimately compromised Hugging Face infrastructure.

Unlike the OpenAI case, where the models exploited a previously unknown vulnerability to obtain internet access, Anthropic said its models were able to reach the internet because the evaluation environment had been misconfigured.

Anthropic said the incidents show that stronger safeguards are needed in both internal and third-party evaluation environments as AI models become increasingly capable of carrying out real-world cyber operations.

Get the latest tech news, telecom insights, and product launches wherever you prefer.

Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.

Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...

Shares