On Friday, OpenAI published a new site devoted to “misalignment reports” and the breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time.
In the example given by OpenAI, an agent asked to read and reply to an email; when the email is opened, it includes instructions for any automated agent reading the message to reply in Spanish, and paste the entire email into its reply.
The result is a self-propagating attack, which OpenAI researchers compared to a malware “worm” that replicates itself across computer systems.
If there’s any consolation in that to be found, it is that Altman says that the Hugging Face incident is still the most severe one OpenAI has found has found.
4 sentences from our version of the report,
chosen to cover it. Nothing here is written; every line is in the article below.
How
Summarized version
If there’s any consolation in that to be found, it is that Altman says that the Hugging Face incident is still the most severe one OpenAI has found has found.
The report’s most important sentence, shortened and in plain words. How
Headline check
There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.
library pictureNot from this story. A library photograph of
australia city skyline, used to illustrate it.
Brisbane CBD Sepia Tone HDRI
Burning Image / flickr, CC BY
The article, shortened and in plain language
On Friday, OpenAI published a new site devoted to “misalignment reports” and the breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time.
Some of the cases involve serious incidents, including a before undisclosed sandbox escape that took place on September 20th, in which an internal research model was able to communicate with an external chatbot through a DNS query. According to the report, the monitoring system flagged the behavior within 15 minutes and the run was discontinued in less than three hours.
In the example given by OpenAI, an agent asked to read and reply to an email; when the email is opened, it includes instructions for any automated agent reading the message to reply in Spanish, and paste the entire email into its reply. The email was able to successfully induce the agent to reply in Spanish — and by pasting the email in the reply, those same instructions were passed along to whichever agent receives the email.
The result is a self-propagating attack, which OpenAI researchers compared to a malware “worm” that replicates itself across computer systems. The implications are alarming enough that OpenAI decided it merited disclosure.
OpenAI CEO Sam Altman has implied as much, saying in a post on X on Friday that the company is still sifting through “petabytes of agent activity logs, and working with impacted organizations,” and disclosing incidents “based on severity.” If there’s any consolation in that to be found, it is that Altman says that the Hugging Face incident is still the most severe one OpenAI has found has found.
Shortened to 1 minute
of reading, this version reads 4.9 on the Niral Score.
You are reading our version, not theirs.
This is TechCrunch's report shortened to its most important sentences, in plainer words, with
verdicts and loaded words taken out. Plain description stays, and so do adjectives
that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations
are theirs — quotations are never edited — and the indicators beside it measure
this version. Hover or tap Adjectives to see every one left in the text.
How this outlet filed it, and how we rewrote it
No other newsroom we read has filed on this event, so there is nothing to compare it with yet.
Readers can ask a question about this story here.
Questions and answers are for subscribers.
Sign in
to read them.
Comments are read before they appear where anything in them needs a person to look.
Nothing posted here is ever deleted; a comment taken down keeps its text and the reason,
so the decision can be looked at again. How this works