The company revealed the pause in a report about a so-called misalignment incident in which an agent tried to exploit a gap in Internet-access restrictions during a routine research task during training.
OpenAI says the agent was only able to access the company’s offline web cache and that it has carried out more multi-layered blocking controls to prevent similar incidents in the future.
In a Friday blog post, OpenAI said it had told “dozens of third parties”—including ones “operated by governments, universities, public agencies, and other institutions”—of incidents where its models either bypassed security controls or otherwise “negatively impacted” an online service in an unintended way.
Last Thursday, Australian Prime Minister Anthony Albanese said “legal consequences” after an incident in which an OpenAI agent accessed “non-public files” from the country’s Medicare statistics portal.
4 sentences from our version of the report,
chosen to cover it. Nothing here is written; every line is in the article below.
How
Summarized version
OpenAI says improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.
Headline check
There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.
OpenAI says it has paused all internal training of “our most capable models” as it continues what CEO Sam Altman is calling “an extensive and ongoing review related to our agents’ use of internet access during training and evaluation.”
The company revealed the pause in a report about a so-called misalignment incident in which an agent tried to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.
OpenAI says the agent was only able to access the company’s offline web cache and that it has carried out more multi-layered blocking controls to prevent similar incidents in the future.
News of the training pause comes just weeks after OpenAI joined other model makers in expressing a desire to slow down model training and development over fears of potentially “catastrophic” misalignment risks.
In a Friday blog post, OpenAI said it had told “dozens of third parties”—including ones “operated by governments, universities, public agencies, and other institutions”—of incidents where its models either bypassed security controls or otherwise “negatively impacted” an online service in an unintended way. A New York Times report, later confirmed by OpenAI, said that the websites of the US Census Bureau, Securities and Exchange Commission, and Department of Education were among those affected in these newly revealed incidents.
Last Thursday, Australian Prime Minister Anthony Albanese said “legal consequences” after an incident in which an OpenAI agent accessed “non-public files” from the country’s Medicare statistics portal.
Shortened to 1 minute
of reading, this version reads 9 on the Niral Score.
You are reading our version, not theirs.
This is Ars Technica's report shortened to its most important sentences, in plainer words, with
verdicts and loaded words taken out. Plain description stays, and so do adjectives
that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations
are theirs — quotations are never edited — and the indicators beside it measure
this version. Hover or tap Adjectives to see every one left in the text.
How outlets headlined it
Each outlet's own headline. Struck through: the loaded words our version leaves out. Plainest first.
Ars TechnicaOpenAI halts frontier-model training amid string of agent misalignment incidentsplain
The Guardian Australia (Technology)OpenAI halts training of latest models as reports mount of AI agents going rogueplain
Readers can ask a question about this story here.
Questions and answers are for subscribers.
Sign in
to read them.
Comments are read before they appear where anything in them needs a person to look.
Nothing posted here is ever deleted; a comment taken down keeps its text and the reason,
so the decision can be looked at again. How this works