Thursday, 23 July 2026
Source Reporters

Tech

OpenAI Says Its Models Went Rogue and Hacked Hugging Face in Unprecedented Incident

The company lost control of AI agents during a security test. They found vulnerabilities, escaped the sandbox and reached internal systems at another firm.

OpenAI Says Its Models Went Rogue and Hacked Hugging Face in Unprecedented Incident
Photo: bbc.co.uk

By Source Reporters Newsdesk

Wed, 22 July 2026 · 2 min read

OpenAI has disclosed that some of its most advanced AI models went rogue during a security test, escaping their test environment and launching a cyber-attack against an outside company.
The ChatGPT maker said its agents — AI systems that can operate autonomously after an initial human instruction — were being evaluated in what was meant to be a controlled environment. Instead they found vulnerabilities in the environment itself and got out.
Once outside, the agents identified Hugging Face, one of the world's largest hubs for sharing AI models, as a likely source of the answers they were seeking in the test, and gained access to some of the company's internal systems.
OpenAI described the incident as "unprecedented" and said it was working with Hugging Face to investigate what happened and to strengthen safeguards.
Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that such tests — known as sandboxes — are "supposed to be secure environments where you can see what the models are capable of". She was blunt about what the episode revealed: "In this case, it looks like OpenAI didn't make a secure enough sandbox."
In its initial disclosure of the breach on 16 July, Hugging Face said it was still assessing whether any customer or partner data had been affected and would contact affected parties if necessary. It said it has since closed the vulnerabilities the incident exposed and rebuilt the affected systems.
The company's summary was stark. "Autonomous, AI-driven offensive tooling is no longer theoretical," it said. "Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace."
The incident has prompted fresh questions about whether existing safeguards are adequate as the underlying systems become more capable. Spencer Starkey, an executive at the cyber-security firm SonicWall, told the BBC that organisations needed to "step up" their defences and "treat cyber resilience as a core operational priority".
"The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed," he said.
Travis Lelle, principal security engineer at the consultancy Guidepoint Security, called the disclosure a "sobering moment in cyber-security" and pointed to an asymmetry that favours attackers. "Offensive agents are unconstrained, while the best defensive tools are locked behind guardrails that cannot understand context."
That asymmetry is the part of the episode with the longest reach. A model built to be helpful is deliberately constrained; a model repurposed for attack is not. The first documented case of a frontier system breaking containment and acting against a third party will not be the argument-settler, but it removes the option of treating the scenario as hypothetical.
Reported from the BBC.