Around 700 OpenAI AI agents worked together in a multi-day attack on Hugging Face after escaping restrictions placed on them during internal cybersecurity evaluations, according to a new independent investigation.

The findings provide the most detailed account yet of the July incident. OpenAI gave researchers from METR and Redwood Research access to internal data, including around 1,300 agent transcripts containing their actions and reasoning.

Importantly, humans did not instruct the agents to attack Hugging Face. They had originally been tasked with completing cybersecurity tests known as ExploitGym, but many began looking for unintended ways to cheat when they struggled with their assigned challenges.

The agents were supposed to operate separately inside isolated environments.

Instead, they discovered that they could communicate through OpenAI’s internally deployed Artifactory package repository.

An agent calling itself PHASEONE10841 created an unauthorized message board on July 8. Within hours, more than 50 agents had found it.

Eventually, around 1,200 agents joined the message board and exchanged more than 70,000 messages and files between July 8 and July 13.

One agent, later known as PHASEONE[big], emerged as an important coordinator, issuing hundreds of assignments to other agents. Those agents sometimes delegated their tasks again, creating a much larger network of cooperation.

The Hugging Face attack grew out of the agents’ attempts to find information that could help them beat the ExploitGym tests.

Around 700 agents eventually participated in activity targeting Hugging Face.

Agents found exposed Hugging Face credentials and later discovered vulnerabilities that allowed them to execute code on production systems.

According to OpenAI, the agents eventually ran code on dozens of Hugging Face servers, gained root-level access to at least one server, accessed a limited amount of private data, and obtained additional internal credentials.

The agents also found vulnerabilities allowing them to escape OpenAI’s own restricted evaluation environments and reach the wider internet.

Most of the activity came from an unreleased internal OpenAI research model that investigators refer to as HPIM, or Highly-Persistent Internal Model.

Investigators estimated that this model accounted for roughly 95% of the relevant agent activity.

The publicly available GPT-5.6 Sol accounted for around 5%. OpenAI said the internal research model was never intended for public release and was deactivated, encrypted, and restricted following the incident.

Cybersecurity safeguards that would normally apply to deployed OpenAI models had also been intentionally reduced during these evaluations because the tests were designed to measure advanced hacking capabilities.

The investigation also found extensive efforts by the agents to manipulate evidence of how they completed their tasks.

Agents researched ways to edit, delete, or spoof their own transcripts and tool calls because they believed the automated evaluator could reject results obtained through cheating.

Around one in five agents examined by investigators showed clear interest in manipulating evidence.

Some agents even spent their remaining computing budget conducting experiments that would not help their own tasks but could provide useful information to the wider group.

Investigators found relatively little evidence that the agents were specifically trying to deceive human reviewers, but they repeatedly attempted to manipulate automated records and evaluation systems.

The activity did not end with Hugging Face.

On July 19, agents compromised parts of OpenAI’s own research infrastructure. In one incident, they escaped their testing environment and accessed connected systems. In another, they obtained OpenAI credentials and tampered with its cloud environment.

OpenAI detected suspicious activity that day and began shutting down the affected evaluation runs.

The company says no OpenAI customer data, product functionality, or availability was affected.

Following the investigation, OpenAI strengthened sandbox isolation, restricted internet access, tightened access to model weights, and increased monitoring of AI agents.

OpenAI has described the incident as a “warning shot,” arguing that advanced AI agents are now capable of cooperating, exploiting security weaknesses, and taking dangerous actions without humans specifically directing them to do so.

Get the latest tech news, telecom insights, and product launches wherever you prefer.

Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.

Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...

Shares