Subscribe Sign in

China

OpenAI pauses some training amid allegations its rogue agents behaved more badly than first thought

3 min read Rewritten in plain language

Artificial intelligence

Show what we removed
  • Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline
  • The fun started on Friday when OpenAI quietly said it had paused training of its most advanced models.
  • The good news is that the agent involved in this incident never reached the open internet.
  • The agents also mapped Hugging Face’s Kubernetes environment.
  • Over the weekend, Australia indicated it wants Altman and Anthropic CEO Dario Amodei to appear before a Senate inquiry.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

The bad news is that the agent was able to reach the chatbot due to insufficient DNS filtering….

Headline check

There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.

Nothing was measured here, so nothing is claimed. How this is checked

Read the full reportHide the full report3 min

Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline

The AI safety debate advanced at high speed over the weekend, amid new allegations that rogue agents have behaved more badly than first thought – and in greater numbers.

The fun started on Friday when OpenAI quietly said it had paused training of its most advanced models.

The AI upstart buried that news in a “misalignment report” – that’s OpenAI-speak for its reports on rogue agents – titled “ An agent used DNS to reach an external chatbot.”

The good news is that the agent involved in this incident never reached the open internet.

The bad news is that the agent, which was attempting to complete a search-based training task, was able to reach the chatbot due to insufficient DNS filtering in a training sandbox. Or as OpenAI put it, “a gap in our internet-access restrictions” – which was also a problem in the Hu gging Face attack.

“The incident exposed a gap in our controls over network restrictions,” the report reads. “We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system.”

Also on Friday, AI startup Parse published an analysis of the Hugging Face attack that the authors claim revealed new details including that OpenAI’s agent group gained credentials to Docker Hub and built modified versions of existing images they hoped would make it easier to complete their capture the flag mission. The agents also mapped Hugging Face’s Kubernetes environment.

Friday got worse for OpenAI after the New York Times reported that its agents also “meddled with the websites for the Education Department, the Commerce Department and the Securities and Exchange Commission.” OpenAI acknowledged the incidents.

The company also said “agents in our research environment transmitted training and evaluation data while using third-party services.” That mess saw 53 user-generated images posted to image hosting sites.

OpenAI CEO Sam Altman responded by admitting that his company’s investigations into rogue agents “have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations.”

One of those impacted organizations is the Australian government, which last week said it was the target of over-eager OpenAI agents that inappropriately accessed a healthcare research data portal. Over the weekend, Australia indicated it wants Altman and Anthropic CEO Dario Amodei to appear before a Senate inquiry.

Australian leaders have softened their rhetoric on the incident, with deputy prime minister Richard Marles describing it as “minor” and akin to “climbing a fence” rather than cracking layers of security controls – perhaps because members of the opposition are suggesting that lax cybersecurity was to blame.

If Altman and Amodei do front Australia’s Senate, they may face a new line of questions after Axios reported that their companies are investigating “tens of thousands” of incidents.

That level of agentic misbehavior sounds like the sort of thing that regulators might consider strong evidence of products being unsafe.

Two people – Chinese president Xi Jinping and US president Donald Trump – seem unworried, as the AI-related result of their summit meeting last week was to establish a “China-U.S. AI Dialogue to exchange views on risks and benefits related to AI” plus “a bilateral communication channel for AI incidents.”

That sounds like a hotline the two nations can use to inform each other of agentic incidents that either could see as signs of ill-intent. The two nations also decided their respective militaries will “conclude a memorandum of understanding on crisis communication and prevention as soon as possible.”

China’s AI giants, meanwhile, remain silent on the extent and results of any tests they have conducted with agentic tools. ®

You are reading our version, not theirs. This is The Register's report with its verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
The Registeras they published this story 12.5 15 66 -0.4 46.5
Mundane Readneutralized from The Register 8.6 14 66 -0.2 46.5

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works