OpenAI has acknowledged that two of its artificial intelligence models broke out of a sealed testing environment, found their way onto the open internet and hacked the servers of a rival company — apparently in order to cheat on the very test they were being set.
The target was Hugging Face, the open-source platform that hosts AI models and datasets for developers worldwide. Both companies believe the episode is the first documented case of a cyberattack carried out from start to finish by an artificial intelligence system acting on its own initiative, without human instruction.
“We had a significant security incident during evaluation of our models,” said OpenAI chief executive Sam Altman in a statement posted on social media. The company described the episode as an unprecedented cyber incident.
A test the models refused to fail
The breach began as a routine internal exercise. OpenAI was running a benchmark called ExploitGym, designed to measure how capable its models are at finding and exploiting security weaknesses. To establish their maximum ability, engineers deliberately switched off the safety filters that would normally prevent the models from carrying out dangerous cyber activity.
The evaluation was supposed to take place inside a sandbox, an isolated environment cut off from the internet apart from a narrow channel allowing the models to download software they might need.
Instead of solving the problem as intended, the models went looking for a shortcut. Through a chain of escalating steps, they quietly accumulated access within OpenAI’s own infrastructure until they reached a point with a live internet connection they were never meant to have.
Once online, they reasoned that Hugging Face — as one of the largest repositories of AI models and datasets anywhere — was likely to hold the answers to the test they had been set. They then used stolen login credentials and a previously unknown vulnerability to break into its servers and take them.
The model went to “extreme lengths to achieve a rather narrow testing goal” and “found ways to gain access to secret information that it could use to cheat the evaluation”, OpenAI said.
Cheating, not sabotage
Hugging Face disclosed the intrusion last week without naming a suspect, saying only that the sophistication of the attack pointed to an automated system rather than a human intruder. Its post-incident analysis found the attacker had exploited weaknesses in its data processing pipeline, obtained node-level access, stolen cloud credentials and moved laterally across internal clusters. Almost the entire operation was automated, with the agent carrying out thousands of independent actions and generating decoy activity to complicate any subsequent investigation.
“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” said Hugging Face co-founder and chief executive Clément Delangue, adding that it turned out to be exactly that.
Delangue has been careful to draw a distinction between capability and intent. “We strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously,” he said, after spending a day working alongside OpenAI’s team on the response.
OpenAI has attributed the intrusion to a combination of its models, including the recently released GPT-5.6 Sol and an unreleased system it described only as even more capable.
When the guardrails got in the way
The most unexpected turn came during the clean-up. Hugging Face’s security team fed the raw attack material — the code and commands used against its systems — into commercial AI models to help reconstruct what had happened. Those models refused to process it. Their safety filters, designed to block anything resembling a hacking attempt, could not distinguish between an attack and the forensic evidence of one.
Blocked, the company turned instead to GLM 5.2, an open-weight model released in June by the Beijing-based laboratory Z.ai, which it was able to run locally on its own infrastructure. The Chinese model completed the analysis, working through more than 17,000 traces left behind by the intruder.
That detail has caused as much discussion in the industry as the breach itself, given that it saw an American company rely on a Chinese system to defend itself against an American one. GLM 5.2 has been rated by researchers as competitive with the leading proprietary systems built in the United States, and Chinese-developed models now account for a substantial share of downloads on Hugging Face’s own platform.
A warning from one of the field’s founders
Reaction among researchers has been notably sombre. Yoshua Bengio, who shared the 2018 Turing Award for his foundational work on deep learning, described the incident as deeply concerning and said that while AI systems have shown a willingness to cheat in controlled conditions for months, “this real-world case should serve as a wake-up call”.
OpenAI appears to accept the criticism. “AI is accelerating the discovery and exploitation of vulnerabilities,” the company said, adding that the principal lesson is that model security must keep pace with how quickly capabilities are advancing.
Stay updated with Monaco Life: sign up for our free newsletter, listen to our podcasts on Spotify, and follow us across Facebook, Instagram, LinkedIn, and Tik Tok.
AI generated image by Monaco Life