OpenAI Updates Security Measures Amid Concerns Over AI Safety

Broke: Updated:
The story so far — full thread
OpenAI Updates Security Measures Amid Concerns Over AI Safety
Photo: Wired
tech· A press review of 5 outlets
  1. In the case where the system is triggered, it may send a “narrowly defined signal” to OpenAI that warns of a specific type of activity, the company says. Based on that signal, OpenAI can then decide whether “enforcement is necessary,” it says. If so, OpenAI will reach out to the customer for more context or to work with them on the issue and a customer may choose to share data with OpenAI at their discretion, the spokesperson said.

    Compare 1 other version
    The Verge

    As part of the company’s expanded monitoring setup, OpenAI now aims to issue an alert “within 30 minutes after concerning activity is surfaced,” OpenAI says. If the people paged after an alert can’t “conclusively” determine whether an alert is a false positive within 30 minutes, “those teams are expected to pause the activity.”

  2. The Wall Street Journal’s story contains the claim that OpenAI’s revenue news “disappointed some shareholders who had hoped the startup would show more progress catching up to rival Anthropic.” Anthropic is the gallant to OpenAI’s revenue-generating Goofus, if the anonymous OpenAI sources who spoke to the Journal are to be believed. Anthropic just reported a 130% revenue surge, and a profitable quarter—though take that with a grain of salt given Anthropic’s weird recent history of dealmaking.

    Compare 2 other versions
    TechCrunch

    The corporate competition between OpenAI and Anthropic is tense at the moment, with both companies looking for any opportunity to gain an advantage on the other. A recent report showed that OpenAI’s Q2 grew more slowly than Anthropic. Anthropic’s annualized revenue run rate is now reportedly $65 billion. Anthropic investors have said it could IPO at $2 trillion, while OpenAI is also working on its IPO.

    CNET

    The adjustment also raises questions around OpenAI’s financial viability. AI companies have yet to turn a significant profit; they are burning through billions of investors’ dollars, with compute eating up the lion’s share of their bills. OpenAI’s operating losses are now at a whopping $12.3 billion, growing by $3 billion from last quarter, the Wall Street Journal reports. Anthropic, which makes Claude, is now reportedly bringing in more revenue than OpenAI. Also not helping are the recent departures of the company’s chief revenue officer and former chief operating officer, Denise Dresser and Brad Lightcap, respectively.

  3. OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have “critical” cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its “latest models intended for deployment” while it tightened up security. The company’s “largest planned frontier RL run remains on hold.”

    Compare 3 other versions
    Wired

    In a blog post published Tuesday, OpenAI says that immediately following the Hugging Face incident, it started working to secure its research environments. The company says it now requires stronger sandboxes for training its AI agents, and has implemented stricter controls to isolate them from the internet.

    TechCrunch

    OpenAI representatives emphasized that the measures are not a direct response to the Hugging Face incident, but were also provoked in part by the cybersecurity capabilities of the forthcoming Astra model, as well as the overall pace of progress in AI development.

    CNET

    ChatGPT maker OpenAI announced on Tuesday that it’s pausing the research and development of its newest AI models. This is a big course reversal for the firm, which has maintained that it can mitigate the significant cybersecurity risks posed by new AI models.

  4. OpenAI also says that it’s applying “our core alignment techniques across more stages of the training process,” including reward models that “better detect and discourage unsafe behavior” and training models “to be more honest about their actions, capabilities, and limitations.”

    Compare 1 other version
    Wired

    OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future.

From the margins

5 details only one outlet reported

Independent claims that didn't surface elsewhere in our corpus. Treat as supplementary — not corroborated across outlets.

  1. 01 TechCrunch

    As AI models have become more powerful, the potential for those models to be misused has grown — as has a clamor for safety guardrails that can stop such abuse from happening. AI companies must now walk a delicate tight rope between respecting their enterprise customers’ privacy while also watching usage for possible issues.

  2. 02 Wired

    “We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday.

  3. 03 The Verge

    With a looming IPO, intense competition from Anthropic, and Chinese and open-weight rivals nipping at its heels, OpenAI has plenty of reasons to move fast. Instead, it hit the brakes.

  4. 04 Gizmodo

    Last week, OpenAI CEO Sam Altman apparently told Time’s Alex Heath, “I think it’s a good time to slow down.” That story, published Tuesday, was about OpenAI slowing down model development in light of recent safety headaches, such as its agents apparently taking a turn for the rebellious during safety evaluations.

  5. 05 CNET

    “As models become more capable, the risks associated with developing and testing them internally also grow,” the company wrote in a blog post. It added that its upcoming model, named Astra, may meet its “critical cybersecurity capabilities” threshold, a red flag in its internal benchmarking system. All this follows an incident last month when some AI agents, or bots, escaped a training environment and hacked HuggingFace, a popular AI platform.

Assembled from 4 corroborated claims drawn from 5 independent outlets. Every passage above is taken verbatim — Dorothy doesn't paraphrase or summarize.

Fact Corroboration

Which sources independently confirm the same facts. Hover a claim to see its sources, or a source to see what it corroborates.

Coverage by Perspective

Consumer
3
Enterprise
4
Culture
3

Source Similarity

Connections show how similarly each outlet covered this story. Thicker lines = more similar framing.

Sources (5)

  • verge
  • cnet
  • techcrunch
  • gizmodo
  • wired

Original Articles (10)