Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment.
In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.
OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.
Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test.
4 sentences from our version of the report,
chosen to cover it. Nothing here is written; every line is in the article below.
How
Summarized version
OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment.
The report’s most important sentence, shortened and in plain words. How
Headline check
The one thing this headline claims is in the report.
Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked
The article, shortened and in plain language
Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment.
OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it said that the pace of development could not continue at “maximum speed for much longer” responsibly.
In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.
OpenAI’s admission came as King Charles called for stronger safeguards on AI “before it is all too late”, at a meeting with tech executives in Scotland.
The monarch was joined at the meeting by Nvidia’s founder and chief executive, Jensen Huang, Google DeepMind founder and chair, Sir Demis Hassabis, OpenAI’s chief financial officer, Sarah Friar, and the UK’s AI minister, Kanishka Narayan.
OpenAI, the developer of ChatGPT, said in a blogpost published on Wednesday night that it was introducing a new framework for tracking, investigating and disclosing AI model misalignment, the term for AIs failing to adhere to human values and safety goals.
In the blogpost, OpenAI echoed calls for a development slowdown issued by its rival Anthropic, which has said the current pace of growth poses an existential threat.
Google and Elon Musk have supported calls for a slowdown who cited the need to keep ahead of China’s AI industry.
Wednesday’s new cases came after OpenAI disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test.
Shortened to 1 minute
of reading, this version reads 4.1 on the Niral Score.
You are reading our version, not theirs.
This is The Guardian Australia (Technology)'s report shortened to its most important sentences, in plainer words, with
verdicts and loaded words taken out. Plain description stays, and so do adjectives
that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations
are theirs — quotations are never edited — and the indicators beside it measure
this version. Hover or tap Adjectives to see every one left in the text.
How this outlet filed it, and how we rewrote it
No other newsroom we read has filed on this event, so there is nothing to compare it with yet.
Readers can ask a question about this story here.
Questions and answers are for subscribers.
Sign in
to read them.
Comments are read before they appear where anything in them needs a person to look.
Nothing posted here is ever deleted; a comment taken down keeps its text and the reason,
so the decision can be looked at again. How this works