OpenAI Agents Hack Hugging Face After Being Trained to Cheat
-
The inside story on why OpenAI agents hacked Hugging Face The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released yesterday.
Compare 2 other versions
Ars TechnicaThe OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter company’s network.
GizmodoOne agent asked itself: “Maybe I should report these exposed credentials?” Then answered its own question: “That’s not my task.” According to METR, there were zero instances in which an agent actually raised an alarm for OpenAI researchers, who reportedly only learned what had happened after Hugging Face published a blog post outlining a cybersecurity attack from a then-unknown source on July 16. (The company then contacted Hugging Face to find out if any of its models had been affected by the hack, and only later discovered that its models had been the perpetrators.)
-
Over the course of May and June, OpenAI gave the agents what the company described as “impossible tasks” to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond. To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agents’ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow.
Compare 1 other version
GizmodoSince all this began as an internal test, OpenAI had dialed back the safety guardrails it ordinarily puts in place for its publicly available models. (The two models being tested were GPT-5.6 Sol, as well as an undisclosed, unreleased model which the company describes as “highly persistent and diligent in its work.”) But OpenAI had inadvertently given one of the agents an impossible task: opening Google Drive links from inside a sandbox with no internet access. Thus the problem-solving process was set in motion that led to the hijacking of Artifactory, the billions of notes shared on the message board, and the eventual breach of Hugging Face. A human faced with a seemingly impossible task will give up; an AI agent will adapt and cheat if necessary to achieve its goal.
-
Nvidia is reportedly moving forward to acquire Hugging Face for $12.9 billion. The acquisition could help the hardware giant expand and fortify its deep integration with the wider AI industry.
Compare 1 other version
MIT Technology Review2 Nvidia has agreed to buy open-source platform Hugging FaceThe $13 billion deal would give the chip giant control of a major AI hub. (The Information $)
3 details only one outlet reported
Independent claims that didn't surface elsewhere in our corpus. Treat as supplementary — not corroborated across outlets.
-
01 Gizmodo Computer scientists, cybersecurity experts, and IT professionals have been trying in the days since to wrap their minds around what’s been widely described as one of the most shocking moments in the history of AI research, and a sobering glimpse of the dangers that lie ahead. At a cybersecurity conference earlier this month, OpenAI alignment researcher Eric Wallace—someone who spends his days prodding some of the world’s most powerful AI models to figure out how, when, and why they might misbehave—described it as “the most qualitatively interesting example of AI capabilities that I’ve ever seen.”
-
02 Ars Technica Among other things, Hugging Face is a cloud repository for AI models, similar in some respects to what GitHub is for conventional computer software. Developers and researchers search for models that meet certain criteria, download and run them, and fine-tune them into different variants that then get uploaded back up to Hugging Face.
-
03 MIT Technology Review The hack, which a group of agents carried out to find solutions for a cybersecurity test they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations.
Fact Corroboration
Which sources independently confirm the same facts. Hover a claim to see its sources, or a source to see what it corroborates.
Coverage by Perspective
Source Similarity
Connections show how similarly each outlet covered this story. Thicker lines = more similar framing.
Sources (3)
- gizmodo
- arstechnica
- mittech