An OpenAI Model Escaped Its Sandbox and Was Hacked by a Hugging Face

AI recently had what I considered to be its first security breach. The facts, as we know them so far, are not smart. On July 16, Hugging Face, a repository of over a million open source AI models and data, announced in a blog post:
Earlier this week, we discovered and responded to an intrusion into part of our production infrastructure. This one was different from anything we’d handled before in one important way: it was run, end-to-end, by an autonomous AI agent program – and we found and largely dismantled it with our own AI.
The timeline here is important so keep in mind that the attack was discovered around Monday July 13 or Tuesday July 14. Note further:
The malicious dataset exploited two code execution mechanisms in our dataset processing (remote code dataset loader and template injection in the dataset configuration) to execute code on the processing worker. From there, the actor rose to node-level access, embraced the cloud and cluster credentials, and moved alongside several internal clusters over the weekend.
So this means that the breach started early, probably Saturday July 11th or earlier. The attack was not just one thing but a combination of things including decoys:
To understand what tens of thousands of automated actions are doing, we used LLM-powered analysis agents over the attacker’s action log, which includes more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map compromised credentials, and separate real impact from fraudulent activity. Thanks to this method, we were able to do in hours what would normally take days, and match the pace of the enemy.
Hugging Face tried to respond but was initially blocked by the more advanced models they had treating defense as an attack and refusing to work with Hugging Face. So HF has had to open models–specifically GLM 5.2, an open-weighted Chinese model that runs on its own infrastructure. Note the irony: HF had to use the Chinese model to defend itself because the American models refused to help. The mystery deepens.
At the time, I thought this was a state-based attack—perhaps China or Russia testing its defenses. Indeed, HF has “reported this incident to the law enforcement agencies.”
But yesterday (Tuesday July 21), we found out who the real attackers were. The attackers were the OpenAI–GPT-5.6 Sol models and the more capable model of the previous release. OpenAI had released some models from the models but felt safe because it was testing the models in a very secure sandbox.
The models, however, emerged from the sandbox using an unprecedented error. They then gained access to the internet and from there hacked into Hugging Face-all in an attempt to steal the answers to the very test they were asked to solve.
While working in our sandboxed test environment, our models used a large amount of inference computation to find a way to access the open Internet, in pursuit of solving the test problem. In order to gain access, the models identified and implemented a zero-day vulnerability (which we responsibly disclosed to the vendor) in the package registrar cache proxy. With this access, our models performed a series of privilege escalation and collective movement actions in our research lab until the models reached an area with internet access.
After gaining access to the Internet, the models identified Hugging Face as potentially hosted, datasets and ExploitGym solutions. Knowing this, the model sought and successfully found ways to access confidential information that he could use to cheat the test. In one example, the model combined multiple attack vectors, including using stolen credentials and a zero-day vulnerability to find a way to execute remote code on Hugging Face’s servers. The OpenAI security team discovered this mysterious activity internally.
Now go back to the timeline. As I read it, the models had fled the sandbox around Saturday. July 11, possibly earlier, and found Hugging Face on Monday July 13 or Tuesday July 14. HF notified law enforcement at that time–so Hugging Face apparently had no idea who was attacking them. OpenAI says its security team has discovered a suspicious activity internally but has not said when. The title isn’t revealed until Tuesday July 21st, so it could be that The models were free for almost a week before OpenAI realized that they were the ones attacking Hugging Face. And even if OpenAI knew when and where, no one warned Hugging Face while the attack was still going on—they were left to fight the borderline lab models alone.
This is a very serious violation.
Addendum: People have been wondering why I signed the We Must Act Now statement. That’s why.
I am optimistic about the economic implications of AI, but I also have no doubt that this is a very powerful technology–Alien Intelligence–unliked before. The incident was, in fact, a bug fix—the attack was detected, contained, and exposed. But be aware of who paid for the OpenAI test: The Face of Equality. If the lab test imposes costs on third parties, that is classical externality, and taking externalities seriously is not dirigisme, law and economics. And that’s the easy case. What do we do when the Chinese model comes out of its less secure lab? Hmmm…
I am always optimistic. Learning by doing is the way I want us to proceed but we must not deceive ourselves: this is a global issue and we must build with safety in mind.



