Subscribe Sign in

AI labs want in-house auditors — but maybe they should shut the front door first

TechCrunch
1 min read Rewritten in plain language

Artificial intelligenceAI LabsAuditing

Show what we removed Rules applied: A1×2 A3 A6×3 C5 D1×2 D2×8 D3×7 F2 all 30 rules
  • Executives at OpenAI, Google, and SpaceXAI have already rallied around Amodei’s plan, which has quickly become a central pillar of the emerging AI safety push.
  • While alignment remains a concern, Sayash Kapoor, an AI researcher who will be a professor at UC Berkeley starting next year, argues that “marginal investments in control are more likely to be effective compared to those in alignment.”
  • Security experts that TechCrunch spoke to said that real-time monitoring is key to preventing future break-outs, and that every agentic session should be time-limited and expire.

3 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Headline check

There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.

Nothing was measured here, so nothing is claimed. How this is checked

Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations “to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.” Executives at OpenAI, Google, and SpaceXAI have already rallied around Amodei’s plan, which has quickly become a central pillar of the emerging AI safety push.

Internet security experts say the labs need to focus on network security basics like logs and permissions, applying the same rigorous defenses they do for human users.

While alignment remains a concern, Sayash Kapoor, an AI researcher who will be a professor at UC Berkeley starting next year, argues that “marginal investments in control are more likely to be effective compared to those in alignment.”

The incidents that have spurred these concerns revolve around frontier models being asked to complete training tasks, usually cybersecurity evaluations, and then accessing the open internet and penetrating closed third-party systems in an attempt to do so.

“We as a profession know how to block access to the internet,” Avery Pennarun, the CEO of Tailscale, a security company, said.

In one case, where OpenAI agents took over a defunct German WikiForum to cheat on evaluations, the agents were active for weeks before anyone at the company appeared to notice. Security experts that TechCrunch spoke to said that real-time monitoring is key to preventing future break-outs, and that every agentic session should be time-limited and expire.

Shortened to 1 minute of reading, this version reads 7.5 on the Niral Score.

You are reading our version, not theirs. This is TechCrunch's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
TechCrunchas they published this story 10.5 23 57 -0.1 46.8
Mundane Readneutralized from TechCrunch 9.6 21 57 -0.1 46.8

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works