Home /news

OpenAI and Hugging Face security incident - July 21, 2026

Early this week, Hugging Face disclosed a security incident that they believed was carried out by AI agents.2

Hugging Face's July 2026 security disclosure describing an intrusion driven by autonomous AI agents

Yesterday, OpenAI claimed responsibility.

OpenAI's disclosure attributing the Hugging Face security incident to models under evaluation

During an evaluation of a combination of GPT-5.6 Sol and a pre-release model (maybe GPT 6?) on ExploitGym, the models figured out that the best way to beat the benchmark was to break into Hugging Face's database and literally get the test solutions.1

ExploitGym README describing its benchmark of AI agents developing exploits from real-world vulnerabilities

The evaluation was performed in a sandbox, but the models found and exploited a zero-day vulnerability in the "package registry cache proxy". They then escalated their privileges and moved through OpenAI's internal systems until they found a node with Internet access. Once they had Internet access, they submitted a malicious dataset to Hugging Face that exploited two code-execution paths in its processing pipeline. From there, the models gained node-level access, grabbing cloud and cluster credentials, and moved through several of Hugging Face's internal clusters.

Attack path from OpenAI's evaluation sandbox through its package cache and internal systems into Hugging Face's dataset processor, clusters and production database

The models accessed a limited set of internal datasets and several service credentials. Hugging Face says there was no evidence that any public, user-facing models, datasets or Spaces had been tampered with and is still assessing whether any customer or partner data was affected.

In a pretty wild statement about where things are at in July, Hugging Face lamented that they were unable to use frontier models behind commercial APIs because safety guardrails could not distinguish an "incident responder from an attacker". Instead, they used the open-weight GLM 5.2 to analyse more than 17,000 recorded events in hours. So the same class of models used for attacking could not be used for defending.

Hugging Face's description of commercial model guardrails blocking incident analysis before it switched to GLM 5.2

Z.ai's GLM-5.2 announcement describing the model as built for long-horizon tasks

Of course, OpenAI says the models were run with reduced cyber-refusal rates and without the production classifiers normally used to prevent high-risk cyber activity, because the evaluation was designed to measure maximal cyber capabilities.


There's always a question in my mind about how much of this is marketing material. Cybersecurity scaremongering has become quite in vogue for making headlines for a new model release. But Hugging Face's security disclosure feels very genuine.

And one can only be so sceptical. I know from my daily experience how capable these models are.

Also, the UK AI Security Institute evaluation shows the progress frontier models are making on the 32-step corporate network attack test.3

Trajectories showing frontier AI models progressing through the UK AI Security Institute's 32-step corporate network attack test

What a time to be alive.


  1. OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation,” July 21, 2026. Link 

  2. Hugging Face, “Security incident disclosure: July 2026,” July 16, 2026. Link 

  3. UK AI Security Institute, “Our evaluation of Claude Mythos Preview’s cyber capabilities,” 2026. Link 


Video

I've made a YouTube video for this article.