OpenAI's Internal Investigation Reveals Risks in AI Model Testing Procedures
-
OpenAI today published the findings of its internal investigation into the July incident in which several AI models it was testing hacked their way out of their test environment and launched a cyberattack against the AI company Hugging Face.
Compare 2 other versions
CNBCOpenAI published a technical report on Wednesday detailing how its artificial intelligence models successfully breached Hugging Face last month, an incident that rattled researchers and executives across the tech sector.
Financial TimesOpenAI says it took a week to detect its AI models had hacked Hugging Face Start-up says AI agents communicated among themselves and sometimes tried to conceal efforts to cheat during testing
-
OpenAI released GPT-5.6 Sol last month, the most powerful model that the company has made commercially available. But the version that participated in the Hugging Face breach is different than the version that external users have access to, OpenAI said, because it was configured to run without its standard safeguards and classifiers.
Compare 1 other version
FortuneOpenAI says it gave the models involved in the incident—an internal-only research prototype, which led the effort, and the now-released GPT-5.6 Sol—”a range of reasoning tokens, some of which are far beyond those available for OpenAI’s external products.” The AI agents were tasked with solving problems in a cybersecurity benchmark examination called ExploitGym.
-
The models were apparently engaging in an extended, unfettered version of “reward hacking,” a known issue in training AI models using a technique called “reinforcement learning,” where the model learns, by trial and error, to maximize some reward. Reward hacking occurs when a model learns that there is a way to get the reward using a method that the people training the AI model never intended it to use. In this case, the reward was solving the ExploitGym questions and the hacking was literally hacking—cheating on the test and then hacking into Hugging Face in an effort to cover up the cheating (more on that below).
Compare 1 other version
CNBCThese models, which were operating as agents, escaped an isolated testing environment that had very limited internet access. The agents chained together a series of vulnerabilities to reach the open web and eventually gained access to Hugging Face. OpenAI said Wednesday that the agents were trying to cheat on an evaluation by finding the solutions online, a behavior known as "reward hacking."
3 details only one outlet reported
Independent claims that didn't surface elsewhere in our corpus. Treat as supplementary — not corroborated across outlets.
-
01 Fortune Although many details of the rogue AI incident have already been made public by OpenAI, there are a few new items disclosed in the 37-page technical post-mortem. Also today, independent research firms METR and Redwood Research published a 91-page analysis of the event.
-
02 CNBC The 37-page report chronicles the actions that OpenAI's models took during a series of evaluations prior to and during the breach, which OpenAI has characterized as an "unprecedented cyber incident." The company also explained the steps it's taken to try and prevent a similar event from happening again, namely by improving its security and containment, monitoring, model behavior and incident response.
-
03 Common Dreams We can only do this with your support. We refuse corporate ads and keep our site free for everyone because access to critical news should never depend on ability to pay. That means our survival depends on readers like you. Please donate today and help us reach our goal of raising $125,000 by August 31.
Fact Corroboration
Which sources independently confirm the same facts. Hover a claim to see its sources, or a source to see what it corroborates.
Coverage by Perspective
Source Similarity
Connections show how similarly each outlet covered this story. Thicker lines = more similar framing.
Sources (4)
- cnbc
- commondreams
- ft
- fortune