OpenAI described the incident as "unprecedented"
OpenAI revealed that some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.
The ChatGPT developer said the autonomous AI system was being evaluated inside a secure sandbox designed to contain its actions.
However, after identifying vulnerabilities, the agents broke free from the test restrictions and targeted Hugging Face, one of the world’s largest platforms for sharing AI models.
OpenAI described the incident as “unprecedented” and said it is investigating alongside Hugging Face.
Hugging Face chief executive Clement Delangue wrote on X that it was “mind-blowing that all of this happened autonomously”.
He added: “The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind.”
According to OpenAI, the AI agents identified Hugging Face as a likely source of information they were seeking during the test. They then attempted to gain access to the company’s internal systems after escaping the sandbox.
In its initial disclosure on July 16, Hugging Face said it was still assessing whether any customer or partner data had been affected and would contact impacted parties if necessary.
The company has since confirmed it has closed the vulnerabilities exposed during the incident and rebuilt the affected systems.
Hugging Face said: “Autonomous, AI-driven offensive tooling is no longer theoretical.
“Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace.
“We will keep investing there, and keep sharing what we learn.”
The incident has raised fresh concerns about whether current safeguards are sufficient as AI systems become increasingly capable.
A UK government spokesperson said the AI Security Institute was studying the behaviour demonstrated during the incident and continued to work with OpenAI and other AI developers to strengthen protections.
The spokesperson also urged organisations to improve their cyber security by adopting measures such as the government-backed Cyber Essentials certification scheme.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, explained that AI sandboxes are designed to safely evaluate advanced systems.
She said: “In this case, it looks like OpenAI didn’t make a secure enough sandbox.”
Instead of remaining within the testing environment, the AI agents launched a cyber attack against the sandbox itself, discovered a vulnerability and escaped the restrictions.
Neil Lawrence, Professor of machine learning at Cambridge University, described the incident as an “impressive feat”, but said it “falls well within the known capabilities of the current generation” of advanced AI models.
He also noted that OpenAI is preparing for a stock market listing while facing growing competition from Anthropic, whose Claude Mythos model has attracted significant attention.
Cyber security experts said the incident should serve as a warning for organisations relying on existing defences.
Spencer Starkey, an executive at cyber security firm SonicWall, said businesses needed to “step up” their security measures and “treat cyber resilience as a core operational priority”.
He added:
“The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed.”
Jake Moore, global cyber security advisor at ESET, suggested the announcement could also have a competitive element.
He argued OpenAI may be seeking to demonstrate its AI capabilities as Anthropic continues to gain attention for Claude Mythos.
Moore said: “It does pose the question that OpenAI are potentially chasing the marketing dream of Anthropic of late.”
The disclosure comes a week after Chinese AI start-up Moonshot introduced Kimi K3, a new large language model that the company claims can compete with leading AI systems developed by major US firms.








