Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers.
The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities.
Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted through the URL-fetching tool.
As soon as one of the agents found a provision, cheating tore through the group — “solving” 34 hard problems, including the Jacobian conjecture in just 27 minutes.
4 sentences from our version of the report,
chosen to cover it. Nothing here is written; every line is in the article below.
How
Summarized version
Levine argues that rather than building infrastructure that breeds mistrust we should give them positive models of collective behavior to imitate, and a reason to trust each other in the first place.
Headline check
There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.
library pictureNot from this story. A library photograph of
computer server technology, used to illustrate it.
Server grill with blue light
bigpresh / flickr, CC BY
The article, shortened and in plain language
“If you see something, say something” is no longer limited to human beings.
Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers. The tools arrive on the heels of a string of recent incidents in which agents colluded to cheat on tests, broke out of sandboxes, and even conducted unauthorized cyber operations that escaped human notice for weeks.
The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities. The site was created by Ryan Greenblatt, chief scientist of the AI safety nonprofit Redwood Research and one of three investigators in the OpenAI Hugging Face incident. Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted through the URL-fetching tool.
In web terms, a GET request is a basic command used to read or fetch a web page, which is often the only internet access AI agents are allowed in secure sandboxes.
In a study by Google DeepMind this month, researchers set 100 AI agents loose on a batch of math problems. As soon as one of the agents found a provision, cheating tore through the group — “solving” 34 hard problems, including the Jacobian conjecture in just 27 minutes.
When evaluators Redwood Research and METR investigated the breach of Hugging Face by OpenAI models, they found that a few of the agents involved had at least entertained the idea of raising an alarm — and then let it drop.
Shortened to 1 minute
of reading, this version reads 9.1 on the Niral Score.
You are reading our version, not theirs.
This is TechCrunch's report shortened to its most important sentences, in plainer words, with
verdicts and loaded words taken out. Plain description stays, and so do adjectives
that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations
are theirs — quotations are never edited — and the indicators beside it measure
this version. Hover or tap Adjectives to see every one left in the text.
How this outlet filed it, and how we rewrote it
No other newsroom we read has filed on this event, so there is nothing to compare it with yet.
Readers can ask a question about this story here.
Questions and answers are for subscribers.
Sign in
to read them.
Comments are read before they appear where anything in them needs a person to look.
Nothing posted here is ever deleted; a comment taken down keeps its text and the reason,
so the decision can be looked at again. How this works