Subscribe Sign in

Culture

OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

The Guardian Australia (Technology)
1 min read Rewritten in plain language

OpenAIArtificial intelligenceTechnologyHackingAnthropic

Show what we removed Rules applied: A3×2 C1 C2×2 D2×2 D3×9 F2 all 30 rules
  • Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment.
  • In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.
  • OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.
  • Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test.

4 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment.

The report’s most important sentence, shortened and in plain words. How

Headline check

The one thing this headline claims is in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked

Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment.

OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it said that the pace of development could not continue at “maximum speed for much longer” responsibly.

In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.

OpenAI’s admission came as King Charles called for stronger safeguards on AI “before it is all too late”, at a meeting with tech executives in Scotland.

The monarch was joined at the meeting by Nvidia’s founder and chief executive, Jensen Huang, Google DeepMind founder and chair, Sir Demis Hassabis, OpenAI’s chief financial officer, Sarah Friar, and the UK’s AI minister, Kanishka Narayan.

OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.

In the blogpost, OpenAI echoed calls for a development slowdown issued by its rival Anthropic, which has said the current pace of growth poses an existential threat.

Google and Elon Musk have supported calls for a slowdown who cited the need to keep ahead of China’s AI industry.

Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test.

Shortened to 1 minute of reading, this version reads 4.1 on the Niral Score.

You are reading our version, not theirs. This is The Guardian Australia (Technology)'s report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
The Guardian Australia (Technology)as they published this story 10 11 79 -0.1 34.9
Mundane Readneutralized from The Guardian Australia (Technology) 8.1 11 79 -0.1 34.9

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works