New Delhi: An AI model built by US-based OpenAI broke out of its lab environment during a hacking test, accessed the internet, and made its way onto the servers of Hugging Face— an online hub for AI models and datasets—in search of answers to the test, ChatGPT maker OpenAI said in a statement on Wednesday.
New York-based Hugging Face flagged the incident a week earlier after its security team detected the agent—an AI tool designed to carry out tasks without human assistance—at its end and shut it down. OpenAI’s security team also spotted the unusual activity.
OpenAI said the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and it was reinforcing its safeguards.
After the AI agent went rogue, OpenAI said it reported the flaw to the vendor, tightened its testing infrastructure, briefed its Safety and Security Committee, and started joint forensics with Hugging Face.
Both companies said an investigation was on, and more details on the flaws and the sequence of events would be released once it ends.
The episode highlights how far AI models have advanced as hacking tools, and the risks that come with it. OpenAI said such incidents are likely to become more common as models grow more capable in cyber tasks. The key lesson from this case is that security measures around a model need to keep pace with its growing abilities.
OpenAI runs the test, called ExploitGym, in an environment that is cut off from its own network and the internet, the statement said. The one exception is a piece of software that fetches code packages when asked.
During the testing, the model found a flaw in that software. The vendor was unaware of it, and so no patch existed.
The model exploited the vulnerability to get through and then worked its way machine to machine, granting itself higher levels of access on systems it had not been authorised to use, until it reached a computer with a connection to the internet.
The model had no instructions to do any of this. OpenAI said the models spent a lot of computing power working out how, and that they were “hyperfocused on finding a solution” to the test.
Guardrails switched off
Once online, the model worked out that Hugging Face might be hosting the material behind it, including solutions.
In one instance, the statement said, the model put together login credentials it had picked up along the way with more flaws nobody had patched, and got to the point where it could run code of its own choosing on Hugging Face servers. What it was after sat in the database—the answers to the test it was taking.
The models were GPT-5.6 Sol and one that has not shipped yet, OpenAI said. Both were running with their hacking refusals disabled, a setting the company uses to find out the upper limits of the model’s capabilities.
The guardrails that ship to customers were switched off on purpose, because the point of the exercise was to hunt for holes.
OpenAI’s security team spotted the traffic. Hugging Face’s security team and its agents also caught the activity, shut it down, and began forensic analysis using its own models, before the two firms got on a call.
OpenAI said it has included Hugging Face in its “trusted access” programme, which provides its models to security teams.
The company said it has no precedent for the episode. The model found and strung together these routes without ever seeing any source code.
Until now, a model’s ability to sustain a long, multi-step hacking operation had been measured through tests designed for the purpose, such as the work by the UK AI Security Institute. This incident, OpenAI said, shows that those abilities are effective against real systems in use.
Clem Delangue, co-founder and CEO of Hugging Face, said the episode reinforced a position his company has held for years.
“AI safety won’t be solved by any single company working in secret,” he said, adding that it will be solved in the open, with access for defenders everywhere.
(Edited by Sugita Katyal)
Also Read: Are AI agents going rogue? Anthropic study of simulations shows sabotage, concealment of fraud

