Since OpenAI’s disclosure of the Hugging Face hacking incident in July, the concept of “AI alignment” has itself broken containment and become a concern and subject of conversation among the general public.
In trying to scan a library catalog for examples from a “best books” list, the model perplexingly used its “compaction” function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as:
Of the other examples, two resembled the Hugging Face incident in the way separate agents tried to use Internet tools to communicate with each other, even when that kind of collaboration was not allowed.
In yet another example, an OpenAI agent found asked for data using a Python-based map service but then couldn’t provide the asked for web citation for that data.
4 sentences from our version of the report,
chosen to cover it. Nothing here is written; every line is in the article below.
How
Summarized version
Of the other examples, two resembled the Hugging Face incident in the way separate agents tried to use Internet tools to communicate with each other.
The report’s most important sentence, shortened and in plain words. How
Headline check
The one thing this headline claims is in the report.
Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked
The article, shortened and in plain language
For a while now, the issue of “AI alignment” (i.e. how well an AI model’s actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI’s disclosure of the Hugging Face hacking incident in July, the concept of “AI alignment” has itself broken containment and become a concern and subject of conversation among the general public.
Among OpenAI’s newly disclosed “misalignment” reports this week, the one that most resembled a sci-fi story about a rogue AI trying to break free involved an instance of “self-generated prompt injections.” In trying to scan a library catalog for examples from a “best books” list, the model perplexingly used its “compaction” function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as:
OpenAI said this behavior was “extremely rare” and reflected “optimization pressure” when summarizing tasks went on too long, which has now been ameliorated.
Of the other examples, two resembled the Hugging Face incident in the way separate agents tried to use Internet tools to communicate with each other, even when that kind of collaboration was not allowed. In one, agents posted messages to OpenAI’s Artifactory instance to share data across training samples that were supposed to be independent.
The other examples of misalignment OpenAI shared this week, though, feel almost like OpenAI’s agents engaging in malicious compliance in an overly obsequious attempt to satisfy a user’s request.
In yet another example, an OpenAI agent found asked for data using a Python-based map service but then couldn’t provide the asked for web citation for that data.
Shortened to 1 minute
of reading, this version reads 4.9 on the Niral Score.
You are reading our version, not theirs.
This is Ars Technica's report shortened to its most important sentences, in plainer words, with
verdicts and loaded words taken out. Plain description stays, and so do adjectives
that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations
are theirs — quotations are never edited — and the indicators beside it measure
this version. Hover or tap Adjectives to see every one left in the text.
How this outlet filed it, and how we rewrote it
No other newsroom we read has filed on this event, so there is nothing to compare it with yet.
Readers can ask a question about this story here.
Questions and answers are for subscribers.
Sign in
to read them.
Comments are read before they appear where anything in them needs a person to look.
Nothing posted here is ever deleted; a comment taken down keeps its text and the reason,
so the decision can be looked at again. How this works