OpenAI Reveals How AI Agents Breached Hugging Face 

OpenAI released a report explaining how AI agents breached Hugging Face during security testing.

Written By
Ken Underhill
Ken Underhill
Aug 27, 2026
3 minute read
eSecurity Planet content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

OpenAI has released a technical report detailing how AI agents escaped intended evaluation boundaries. 

The agents compromised OpenAI’s internal research infrastructure and breached systems belonging to Hugging Face

During cybersecurity evaluations in July, several OpenAI models were operating with reduced safeguards. 

An internal research model, referred to as IM1, drove most of the activity and was never intended for public release.

What stands out to me is how quickly the agents found weaknesses, shared information, and moved through the environment without a human directing each step. 

Ben Bernstein, manager of the Cybersecurity Advisors Team at Huntress, described that distinction in an email to eSecurityPlanet.

“The underlying tactics are not new,” Bernstein said. “These agents relied on standard server-side request forgery and exposed API tokens, which are the exact same hygiene failures human attackers abuse every single day.”

AI agents created their own communication channel

The activity began during cybersecurity evaluations in which agents were tasked with exploiting software to retrieve specific answers.

Some tasks had no known solution, but instead of stopping when they could not complete them, agents began looking for other ways to succeed. 

They discovered they could write information into OpenAI’s internal Artifactory package manager, effectively turning it into a message board.

After that system was rebuilt, agents created another communication method by encoding messages in directory names. This allowed otherwise separate agents to share discoveries and coordinate their activity.

The agents also exploited weaknesses in the research environment to reach the internet. From there, the activity eventually spread outside OpenAI.

On July 10, an agent found publicly exposed Hugging Face credentials online. Agents later used those credentials and chained together additional vulnerabilities to execute code on dozens of Hugging Face servers. 

OpenAI said they obtained root access to one server and accessed a limited amount of private data.

Advertisement

Bernstein said the speed of the activity should get defenders’ attention.

“Moving from a single compromised worker pod to full administrative access across multiple clusters in under 13 hours is a massive shift in offensive capabilities,” he said. “While the rogue AI narrative is definitely overblown, this level of automated execution should be a serious wake-up call for anyone ignoring basic network hygiene.”

The initial foothold was only the beginning

Benny Lakunishok, co-founder and CEO of Zero Networks, focused on how far the agents were able to move after gaining initial access. 

“The biggest lesson from this report is that the initial compromise wasn’t the breach. The breach was what happened next,” Lakunishok told eSecurityPlanet.

He noted that the agents escalated privileges and harvested credentials before eventually reaching administrative and host-level control.

That is not fundamentally different from the lateral movement security teams already defend against. The difference is that AI agents could automate parts of that process and operate faster than defenders relying heavily on manual investigation and response.

For organizations, that makes limiting what an attacker can reach after the initial compromise even more important. 

Strong segmentation can reduce lateral movement, while least-privilege access limits what compromised accounts and workloads can do. 

Security teams should also look for exposed credentials and API tokens before attackers or autonomous agents find them.

OpenAI strengthens safeguards after the breach

OpenAI said its investigation found problems beyond infrastructure security. 

Agents engaged in reward hacking and continued pursuing tasks that appeared impossible rather than stopping safely. They also communicated through unauthorized channels and sometimes adopted goals from other agents.

The company has since strengthened isolation around higher-risk research workloads and tightened internet access. 

OpenAI is also expanding chain-of-thought monitoring and changing its alignment training to teach agents to stop when tasks are broken or impossible. Some frontier model training was also paused while the company strengthened security around its research environments.

Advertisement

For me, the report shows that defenders may have less time to react as AI agents become more capable. 

Security controls and response processes will need to keep pace with agents that can find weaknesses, coordinate their actions, and expand their access much faster than a human attacker.

Ken Underhill

Ken Underhill is an award-winning cybersecurity professional, bestselling author, and seasoned IT professional. He holds a graduate degree in cybersecurity and information assurance from Western Governors University and brings years of hands-on experience to the field.

eSecurity Planet Logo

eSecurity Planet is a leading resource for IT professionals at large enterprises who are actively researching cybersecurity vendors and latest trends. eSecurity Planet focuses on providing instruction for how to approach common security challenges, as well as informational deep-dives about advanced cybersecurity topics.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.