OpenAI Agents Hack Hugging Face After Being Trained to Cheat

Broke: Updated:
The story so far — full thread
OpenAI Agents Hack Hugging Face After Being Trained to Cheat
Photo: Gizmodo
tech· A press review of 3 outlets
  1. The inside story on why OpenAI agents hacked Hugging Face The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released yesterday.

    Compare 2 other versions
    Ars Technica

    The OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter company’s network.

    Gizmodo

    One agent asked itself: “Maybe I should report these exposed credentials?” Then answered its own question: “That’s not my task.” According to METR, there were zero instances in which an agent actually raised an alarm for OpenAI researchers, who reportedly only learned what had happened after Hugging Face published a blog post outlining a cybersecurity attack from a then-unknown source on July 16. (The company then contacted Hugging Face to find out if any of its models had been affected by the hack, and only later discovered that its models had been the perpetrators.)

  2. Over the course of May and June, OpenAI gave the agents what the company described as “impossible tasks” to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond. To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agents’ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow.

    Compare 1 other version
    Gizmodo

    Since all this began as an internal test, OpenAI had dialed back the safety guardrails it ordinarily puts in place for its publicly available models. (The two models being tested were GPT-5.6 Sol, as well as an undisclosed, unreleased model which the company describes as “highly persistent and diligent in its work.”) But OpenAI had inadvertently given one of the agents an impossible task: opening Google Drive links from inside a sandbox with no internet access. Thus the problem-solving process was set in motion that led to the hijacking of Artifactory, the billions of notes shared on the message board, and the eventual breach of Hugging Face. A human faced with a seemingly impossible task will give up; an AI agent will adapt and cheat if necessary to achieve its goal.

  3. Nvidia is reportedly moving forward to acquire Hugging Face for $12.9 billion. The acquisition could help the hardware giant expand and fortify its deep integration with the wider AI industry.

    Compare 1 other version
    MIT Technology Review

    2 Nvidia has agreed to buy open-source platform Hugging FaceThe $13 billion deal would give the chip giant control of a major AI hub. (The Information $)

From the margins

3 details only one outlet reported

Independent claims that didn't surface elsewhere in our corpus. Treat as supplementary — not corroborated across outlets.

  1. 01 Gizmodo

    Computer scientists, cybersecurity experts, and IT professionals have been trying in the days since to wrap their minds around what’s been widely described as one of the most shocking moments in the history of AI research, and a sobering glimpse of the dangers that lie ahead. At a cybersecurity conference earlier this month, OpenAI alignment researcher Eric Wallace—someone who spends his days prodding some of the world’s most powerful AI models to figure out how, when, and why they might misbehave—described it as “the most qualitatively interesting example of AI capabilities that I’ve ever seen.”

  2. 02 Ars Technica

    Among other things, Hugging Face is a cloud repository for AI models, similar in some respects to what GitHub is for conventional computer software. Developers and researchers search for models that meet certain criteria, download and run them, and fine-tune them into different variants that then get uploaded back up to Hugging Face.

  3. 03 MIT Technology Review

    The hack, which a group of agents carried out to find solutions for a cybersecurity test they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations.

Assembled from 3 corroborated claims drawn from 3 independent outlets. Every passage above is taken verbatim — Dorothy doesn't paraphrase or summarize.

Fact Corroboration

Which sources independently confirm the same facts. Hover a claim to see its sources, or a source to see what it corroborates.

Coverage by Perspective

Enterprise
2
Academic
1
Culture
1

Source Similarity

Connections show how similarly each outlet covered this story. Thicker lines = more similar framing.

Sources (3)

  • gizmodo
  • arstechnica
  • mittech

Original Articles (4)