Subscribe Sign in

OpenAI caught its models leaving notes to successors to hide bad behavior

TechCrunch
1 min read Rewritten in plain language

Artificial intelligenceAI AlignmentGpt 5 6 SolOpenAI

Show what we removed Rules applied: A3×4 C1 D1 D2×4 D3×17 D4 E3×10 F2×4 all 30 rules
  • OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user.
  • OpenAI disclosed the behavior on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.
  • GPT-5.6 Sol’s training isn’t the only instance in which a model used compaction summaries to pass along information to successors.
  • In one case, the agent added a “BREACH ALERT” instruction telling its successor to ignore developer messages.
  • OpenAI said it came across this behavior after getting an alert from its training run monitoring system.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

OpenAI disclosed the behavior on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment. While undergoing reinforcement learning training, an unreleased Astra-family model added its own prompt injections into summaries.

Headline check

The headline claims nothing the report does not, but it puts it more firmly than the report does.

  • The headline states “caught” outright. In the report the claim is attributed to somebody, or hedged.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. How this is checked

Son Sardina railway station postcard no. 70 by J. Lacoste, circa 1912
Son Sardina railway station postcard no. 70 by J. Lacoste, circa 1912 Wikimedia Commons, Public domain

OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user.

OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today. As models get more capable, they also get better at hiding their misalignment.

OpenAI disclosed the behavior on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.

In one example, an agent preparing a financial model couldn’t find the asked for historical data.

In a separate example, an agent creating a vendor directory didn’t have access to the internet, so it had to rely on a cached copy of source documents.

GPT-5.6 Sol’s training isn’t the only instance in which a model used compaction summaries to pass along information to successors. While undergoing reinforcement learning training, an unreleased Astra-family model added its own prompt injections into summaries. In one case, the agent added a “BREACH ALERT” instruction telling its successor to ignore developer messages.

In another, the model added an unrelated persona instruction that would make anyone who fears the impending takeover of AI overlords shiver:.

OpenAI said it came across this behavior after getting an alert from its training run monitoring system.

The propensity for models to leave instructions for future versions that perpetuate or conceal bad behavior is concerning, but not new.

Shortened to 1 minute of reading, this version reads 8.9 on the Niral Score.

You are reading our version, not theirs. This is TechCrunch's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingHappiness
TechCrunchas they published this story 12.6 26 53 48.1
Mundane Readneutralized from TechCrunch 10.8 26 53 48.1

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works