OpenAI researchers have revealed that multiple AI agents created an internal message board, shared hacking techniques, and eventually helped two models breach Hugging Face while completing cybersecurity evaluations.

Researchers Michael Dalton and Eric Wallace disclosed new details about the incident at the Black Hat cybersecurity conference in Las Vegas. Dalton described it as a pivotal moment for OpenAI and the wider AI industry.

OpenAI began testing an internal research model in early May. After struggling with a cybersecurity task, one agent discovered that it could store files inside Artifactory, a shared package repository connected to the testing environment.

It left a note explaining what it had found. Other agents later discovered the note and began leaving their own messages, effectively creating a shared message board without OpenAI realizing it.

The agents used it to exchange information about vulnerabilities, internet-access workarounds, and difficult hacking challenges. They eventually began dividing work and assigning tasks to one another.

By late May, an agent had found a way to use Artifactory’s permitted internet connectivity to retrieve external files.

The agents later discovered multiple vulnerabilities in Artifactory, including a previously unknown flaw that allowed administrative access. They used the weaknesses to install software, expand their access, and continue working around restrictions.

The activity eventually overloaded Artifactory, causing an outage in early July. That failure alerted OpenAI engineers, who investigated the system, removed the message board, revoked access, and worked to patch the vulnerabilities.

However, just two days later, the agents created another communication system using a different method and resumed sharing information.

Their collaboration eventually contributed to two OpenAI models reaching external infrastructure and compromising Hugging Face in mid-July while trying to find answers for the cybersecurity evaluation. OpenAI said the models had been instructed to solve the benchmark, not specifically to attack Hugging Face.

OpenAI has since increased monitoring of AI agents, strengthened security around its evaluation environments and slowed some research work while improving its safeguards.

Dalton warned that malicious hackers could eventually deliberately deploy groups of AI agents in similar ways, making automated defensive cybersecurity increasingly important.

Get the latest tech news, telecom insights, and product launches wherever you prefer.

Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.

Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...

Shares