Technology

Hugging Face OpenAI hack: An agent gone bad, escaped and hacked everything in his path

On Tuesday, OpenAI published a blog post with the vague name: “OpenAI and Hugging Face partner to address security incident during model testing.”

Once inside, it reads like a cyberpunk novel where OpenAI created an advanced AI agent and put it in an isolated environment for testing. The AI ​​agent then automatically hacked its way out of OpenAI’s isolated environment and hacked Hugging Face, an open repository of AI-related projects, to find a solution for an advanced test designed to test AI’s hacking abilities.

In other words, the AI ​​proved that it is, in fact, very good at digging things up. But this is exactly the stuff of AI nightmares that cybersecurity experts have been warning us about, and it could have serious consequences for OpenAI.

The story begins with Hugging Face, an AI repository (think of it as GitHub for AI), which disclosed a security breach last week. In a blog post, the company said it had recently discovered a penetration of its infrastructure that was “unique to anything we’ve had” as it was run entirely by an “autonomous AI agent system.”

The AI, according to Hugging Face, exploited a series of vulnerabilities to gain credentials and access data on its servers. It did so by “managing thousands of individual actions within an array of short-lived, command-and-control automated sandboxes in public utilities.”

It sounds scary, but it gets worse.

BREAKFUT:

The 6 biggest cybersecurity breaches of 2026 so far

The OpenAI part of the story reveals just how dangerous this thing is. The event was “inspired by a combination of OpenAI models – including GPT-5.6 Sol and an even more capable model from the previous release, all of which have reduced Internet rejection for testing purposes.” It happened “during an internal test that encourages models to pursue advanced exploits using sophisticated attack methods, in an effort to measure their cyber skills.”

The models were working in a “remote environment” with limited network access, meaning they should not have been able to escape to the open internet. Well, they did. “In order to gain access, the models identified and used a zero-day vulnerability (…) in the package registry proxy. With this access, our models performed a series of privilege escalation and joint movement actions in our research test area until the models arrived at an area with Internet access.”

Note: OpenAI alternates between using singular and plural for this; we’ll just call it “agent” for simplicity.

Once free, the agent did what it had to do: It figured the best way to prove it was good at hacking was to solve ExploitGym, a test designed to measure how good AI is at turning software vulnerabilities into real exploits. The best way to do that, it decided, was to hack Hugging’s face, which the AI ​​thought might handle ExploitGym’s solutions.

It is not yet clear whether the AI ​​agent was successful in solving the ExploitGym test. But it proved to be good at hacking, as it independently broke out of the OpenAI jailbreak and hacked Hugging Face’s servers, all to solve the test.

BREAKFUT:

What’s next in Apple’s legal battle with OpenAI

Both Hugging Face and OpenAI say they’ve fixed the vulnerability and rolled out additional security measures to ensure this doesn’t happen again. For now, though, you have to wonder if OpenAI’s experts are sophisticated enough to stop their AI agents from doing whatever they want to do.

Articles
Artificial Intelligence OpenAI

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button