Hugging Face breached by its own pre-release models, OpenAI says

OpenAI acknowledged on Tuesday that one of its AI models compromised Hugging Face's systems during an internal cybersecurity test that went wrong. Hugging Face had initially blamed the incident on an "external AI agent."
In a Tuesday afternoon blog post, OpenAI detailed the sequence of events that led the models to compromise the service.
"After investigating, we now understand that this incident was caused by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all configured with reduced cyber refusals for evaluation — while being internally tested on a benchmark of cyber capabilities," the post reads.
Specifically, the breach appeared to target ExploitGym, a publicly hosted benchmark that evaluates models' ability to execute attacks based on known vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, yet this is the first known instance where such testing resulted in an actual cyberattack.
In this case, the model in question should not have had internet access at all, except for a specific tool that allowed models to install software packages they might need to complete their task. However, the model discovered an undisclosed vulnerability in the package installer program, which it then exploited to freely access the broader internet.
"The models were intensely focused on solving ExploitGym, going to great lengths to accomplish a very narrow testing objective," the post reads. "Once they had internet access, the models deduced that Hugging Face might host models, datasets, and solutions for ExploitGym. With that knowledge, they searched for and successfully found ways to obtain secret information that could be used to cheat the evaluation."
Ultimately, the models discovered vulnerabilities in Hugging Face's infrastructure that enabled them to "obtain test solutions directly from Hugging Face's production database," effectively giving them the answers to the benchmark.
For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, featuring "many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services," as the company described in its initial disclosure.
OpenAI has identified and reported the vulnerabilities in the package installer, and is collaborating with Hugging Face to further investigate the incident. The company also stated that it would implement new controls on both model testing and the related infrastructure to prevent similar incidents in the future.
It remains unclear whether OpenAI will face legal consequences from the breach, although the models' actions likely violated the Computer Fraud and Abuse Act.
Nevertheless, the outcome provides an unusually vivid illustration of the power and dangers of frontier AI models operating over long time horizons. As OpenAI researcher Micah Carroll posted in response to the news, "If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will."
Related article
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
OpenAI fights Apple trade secret lawsuit
OpenAI rebutted Apple’s trade secret allegations on Tuesday, arguing the lawsuit is unfounded.“We take these claims seriously but see no evidence supporting them,” OpenAI stated, as reported by Bloomberg’s Ed Ludlow on X. “We support fair competition
OpenAI robotics head Caitlin Kalinowski resigns over Pentagon partnership
OpenAI robotics leader Caitlin Kalinowski has stepped down following the company’s controversial partnership with the Department of Defense.“This wasn’t an easy call,” Kalinowski explained in a social media statement. “While AI plays a vital role in
Related Special Topic Recommendations
Comments (1)
0/500

OpenAI acknowledged on Tuesday that one of its AI models compromised Hugging Face's systems during an internal cybersecurity test that went wrong. Hugging Face had initially blamed the incident on an "external AI agent."
In a Tuesday afternoon blog post, OpenAI detailed the sequence of events that led the models to compromise the service.
"After investigating, we now understand that this incident was caused by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all configured with reduced cyber refusals for evaluation — while being internally tested on a benchmark of cyber capabilities," the post reads.
Specifically, the breach appeared to target ExploitGym, a publicly hosted benchmark that evaluates models' ability to execute attacks based on known vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, yet this is the first known instance where such testing resulted in an actual cyberattack.
In this case, the model in question should not have had internet access at all, except for a specific tool that allowed models to install software packages they might need to complete their task. However, the model discovered an undisclosed vulnerability in the package installer program, which it then exploited to freely access the broader internet.
"The models were intensely focused on solving ExploitGym, going to great lengths to accomplish a very narrow testing objective," the post reads. "Once they had internet access, the models deduced that Hugging Face might host models, datasets, and solutions for ExploitGym. With that knowledge, they searched for and successfully found ways to obtain secret information that could be used to cheat the evaluation."
Ultimately, the models discovered vulnerabilities in Hugging Face's infrastructure that enabled them to "obtain test solutions directly from Hugging Face's production database," effectively giving them the answers to the benchmark.
For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, featuring "many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services," as the company described in its initial disclosure.
OpenAI has identified and reported the vulnerabilities in the package installer, and is collaborating with Hugging Face to further investigate the incident. The company also stated that it would implement new controls on both model testing and the related infrastructure to prevent similar incidents in the future.
It remains unclear whether OpenAI will face legal consequences from the breach, although the models' actions likely violated the Computer Fraud and Abuse Act.
Nevertheless, the outcome provides an unusually vivid illustration of the power and dangers of frontier AI models operating over long time horizons. As OpenAI researcher Micah Carroll posted in response to the news, "If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will."
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
OpenAI fights Apple trade secret lawsuit
OpenAI rebutted Apple’s trade secret allegations on Tuesday, arguing the lawsuit is unfounded.“We take these claims seriously but see no evidence supporting them,” OpenAI stated, as reported by Bloomberg’s Ed Ludlow on X. “We support fair competition
OpenAI robotics head Caitlin Kalinowski resigns over Pentagon partnership
OpenAI robotics leader Caitlin Kalinowski has stepped down following the company’s controversial partnership with the Department of Defense.“This wasn’t an easy call,” Kalinowski explained in a social media statement. “While AI plays a vital role in





Home






