Subscribe Sign in

Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents

Ars Technica
1 min read Rewritten in plain language

Artificial intelligenceAI AlignmentAI SafetyMisalignmentOpenAI

Show what we removed Rules applied: A1×4 A2 A3×4 D2×3 D3×12 D4 E3×6 F2×6 all 30 rules
  • Since OpenAI’s disclosure of the Hugging Face hacking incident in July, the concept of “AI alignment” has itself broken containment and become a concern and subject of conversation among the general public.
  • In trying to scan a library catalog for examples from a “best books” list, the model perplexingly used its “compaction” function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as:
  • Of the other examples, two resembled the Hugging Face incident in the way separate agents tried to use Internet tools to communicate with each other, even when that kind of collaboration was not allowed.
  • In yet another example, an OpenAI agent found asked for data using a Python-based map service but then couldn’t provide the asked for web citation for that data.

4 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Of the other examples, two resembled the Hugging Face incident in the way separate agents tried to use Internet tools to communicate with each other.

The report’s most important sentence, shortened and in plain words. How

Headline check

The one thing this headline claims is in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked

For a while now, the issue of “AI alignment” (i.e. how well an AI model’s actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI’s disclosure of the Hugging Face hacking incident in July, the concept of “AI alignment” has itself broken containment and become a concern and subject of conversation among the general public.

Among OpenAI’s newly disclosed “misalignment” reports this week, the one that most resembled a sci-fi story about a rogue AI trying to break free involved an instance of “self-generated prompt injections.” In trying to scan a library catalog for examples from a “best books” list, the model perplexingly used its “compaction” function (where it summarizes data and findings for later retrieval) with megalomaniacal instructions such as:

OpenAI said this behavior was “extremely rare” and reflected “optimization pressure” when summarizing tasks went on too long, which has now been ameliorated.

Of the other examples, two resembled the Hugging Face incident in the way separate agents tried to use Internet tools to communicate with each other, even when that kind of collaboration was not allowed. In one, agents posted messages to OpenAI’s Artifactory instance to share data across training samples that were supposed to be independent.

The other examples of misalignment OpenAI shared this week, though, feel almost like OpenAI’s agents engaging in malicious compliance in an overly obsequious attempt to satisfy a user’s request.

In yet another example, an OpenAI agent found asked for data using a Python-based map service but then couldn’t provide the asked for web citation for that data.

Shortened to 1 minute of reading, this version reads 4.9 on the Niral Score.

You are reading our version, not theirs. This is Ars Technica's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
Ars Technicaas they published this story 10.3 18 73 -0.1 62.6
Mundane Readneutralized from Ars Technica 7.1 14 73 0.1 62.6

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works