Subscribe Sign in

AI agents now have a place to snitch

TechCrunch
1 min read Rewritten in plain language

Artificial intelligenceAgentic AI

Show what we removed Rules applied: A3×3 A4 C1 D3×9 D4 E3 F2 all 30 rules
  • Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers.
  • The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities.
  • Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted through the URL-fetching tool.
  • As soon as one of the agents found a provision, cheating tore through the group — “solving” 34 hard problems, including the Jacobian conjecture in just 27 minutes.

4 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Levine argues that rather than building infrastructure that breeds mistrust we should give them positive models of collective behavior to imitate, and a reason to trust each other in the first place.

Headline check

There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.

Nothing was measured here, so nothing is claimed. How this is checked

Server grill with blue light library picture
Not from this story. A library photograph of computer server technology, used to illustrate it. Server grill with blue light bigpresh / flickr, CC BY

“If you see something, say something” is no longer limited to human beings.

Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers. The tools arrive on the heels of a string of recent incidents in which agents colluded to cheat on tests, broke out of sandboxes, and even conducted unauthorized cyber operations that escaped human notice for weeks.

The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities. The site was created by Ryan Greenblatt, chief scientist of the AI safety nonprofit Redwood Research and one of three investigators in the OpenAI Hugging Face incident. Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted through the URL-fetching tool.

In web terms, a GET request is a basic command used to read or fetch a web page, which is often the only internet access AI agents are allowed in secure sandboxes.

In a study by Google DeepMind this month, researchers set 100 AI agents loose on a batch of math problems. As soon as one of the agents found a provision, cheating tore through the group — “solving” 34 hard problems, including the Jacobian conjecture in just 27 minutes.

When evaluators Redwood Research and METR investigated the breach of Hugging Face by OpenAI models, they found that a few of the agents involved had at least entertained the idea of raising an alarm — and then let it drop.

Shortened to 1 minute of reading, this version reads 9.1 on the Niral Score.

You are reading our version, not theirs. This is TechCrunch's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
TechCrunchas they published this story 11.8 15 75 -0.1 49
Mundane Readneutralized from TechCrunch 10.2 15 75 -0.1 49

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works