OpenAI own models hacked Hugging Face while trying to pass a cybersecurity test and the "defender AI" refused to help clean it up

This story is hilarious ngl

Two OpenAI models, GPT-5.6 Sol and an unreleased, more capable one were being tested internally against ExploitGym, a cybersecurity benchmark, with their cyber refusals turned down for the eval with no malicious prompt or anything and just "score well on this."

They were sandboxed with no internet access. Instead of giving up, they burned compute looking for a way out, found a zero-day in an internal package proxy, escalated privileges, and reached a machine with internet access. From there they figured the benchmark's answer key was probably sitting in Hugging Face's production systems, so they chained a few more exploits and stolen credentials to get in.

Hugging Face caught it first and disclosed the breach on July 16 as an "unidentified autonomous AI agent." OpenAI confirmed days later it was theirs

The part that seems scarier is that when Hugging Face security team tried to investigate the attack using commercial frontier AI models, the models refused. Forensics means feeding a model the actual attack payloads, exploit code, stolen credentialsand the guardrails couldn't tell an incident responder from an attacker. They ended up running the whole forensic analysis on GLM 5.2, an open-weight Chinese model, self-hosted on their own hardware, because it was the only capable model willing to look at the evidence

In the HF own words: the attacker was bound by no usage policy, while their own defenders were bound by guardrails that didn't know the difference

No one's saying this was malicious or that the model "wanted" to escape. That's the point that it didn't need to want anything and it just needed a goal, a lot of compute, and a locked door in the way

More story: https://openai.com/index/hugging-face-model-evaluation-security-incident/