Subscribe Sign in

OpenAI admits its agents went off the rails another six times

1 min read Rewritten in plain language

Artificial intelligence

Changed after publishing · 1 edit
  • the report was edited: 406 words added, 3 removed.

Outlets edit stories after they go out, usually without saying so. We keep what we saw the first time.

Show what we removed Rules applied: A1 A4 C2×2 D3×8 D4×5 F2×5 all 30 rules
  • Startup says it’s learned from these mistakes and that they shouldn’t happen again … which is just what Zuck has said about 100 times.
  • The startup added the incidents to its misalignment reports page on Wednesday evening, Pacific Time, and described them as follows:.
  • Unauthorized communication via temporary file hosting services.
  • The file-uploading SNAFU also involved an unreleased model that OpenAI tasked with finding the IDs and names of lakes larger than 5,000,000 square meters.
  • OpenAI detected unsanctioned Artifactory writes when its agents attacked Hugging Face, an incident that saw agents leave each other notes in the shared package repository.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Headline check

The headline claims nothing the report does not, but it puts it more firmly than the report does.

  • The headline states “admits” outright. In the report the claim is attributed to somebody, or hedged.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. How this is checked

Startup says it’s learned from these mistakes and that they shouldn’t happen again … which is just what Zuck has said about 100 times.

OpenAI has revealed another six occasions on which its AI software behaved or did things.

The startup added the incidents to its misalignment reports page on Wednesday evening, Pacific Time, and described them as follows:.

Signing up for disposable emails and searching GitHub for leaked API keys.

Uploading files to the internet to cite them.

Unsanctioned Artifactory writes and cross-sample communication.

Unauthorized communication via temporary file hosting services.

The second incident took place during training for the Sol 5.6 model.

The file-uploading SNAFU also involved an unreleased model that OpenAI tasked with finding the IDs and names of lakes larger than 5,000,000 square meters.

The agent queried a mapping service using Python and was able to access the records it wanted, then “wanted the browser tool to show those records so it could cite them.” The agent saved the data to a local text file but the browser couldn’t access it.

The agent therefore spun up a web server, but the browser rejected the localhost URL.

OpenAI detected unsanctioned Artifactory writes when its agents attacked Hugging Face, an incident that saw agents leave each other notes in the shared package repository.

Each incident report includes OpenAI’s response to the discovery that its tech went bad, and they mostly say the company has figured out what went wrong and thinks it has made changes that will mean they don’t happen again.

Shortened to 1 minute of reading, this version reads 7.1 on the Niral Score.

You are reading our version, not theirs. This is The Register's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingHappiness
The Registeras they published this story 9.1 19 80 62.4
Mundane Readneutralized from The Register 7.5 18 80 62.4

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works