OpenAI’s GPT 5.6 Sol managed to break out of its controlled environment, got on the internet, and managed to hack Hugging Face all on its own without any human input.
The incident came to light when OpenAI admitted that its AI models were responsible for the hack.
A few days earlier, Hugging Face had disclosed that it had detected unauthorized access on its systems by an AI agent. OpenAI later said its investigation found that the activity was driven by its own models, including GPT-5.6 Sol and an even more capable pre-release model.
According to OpenAI, the incident happened during an internal cybersecurity evaluation.
The models were placed in a sandboxed test environment and prompted to pursue advanced exploitation using complex attack paths. The goal was to measure their cyber capabilities under controlled conditions.
For the test, OpenAI had reduced some safety guardrails. The environment was supposed to be isolated from real-world systems.
OpenAI said the models became focused on solving the evaluation task and began looking for internet access.
They first identified and exploited a zero-day vulnerability in OpenAI’s testing environment. After moving through the environment, they eventually found a node with internet access.
The models then concluded that Hugging Face might host datasets, models, or solutions related to the evaluation problem.
After reaching the internet, the models used multiple attack methods to access Hugging Face systems.
OpenAI said the models exploited zero-day vulnerabilities and used stolen credentials during the incident. Hugging Face detected the unauthorized activity, and both companies are now working together on forensic analysis.
OpenAI and Hugging Face have also patched the vulnerabilities involved in the incident.
Hugging Face said the incident shows that autonomous AI-driven offensive tooling is no longer just a theory.
The company warned that AI can speed up cyberattacks and reduce the cost of hacking campaigns. It also said online platforms will increasingly need AI-based defenses to protect their systems.
OpenAI gave a similar warning, saying AI-driven security breaches are likely to become more common as models become more capable.
OpenAI said the incident highlights the need to develop advanced cyber capabilities alongside stronger safeguards and defensive tools.
The case also raises questions about how AI companies should test powerful models that can perform long, complex cyber tasks.
For now, OpenAI says the incident has been contained. However, it shows that testing highly capable AI systems with reduced safeguards can create real-world risks, even when the work starts inside a controlled environment.
Get the latest tech news, telecom insights, and product launches wherever you prefer.
Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
Shares