Subscribe Sign in

Anthropic and OpenAI want to embed safety evaluators. Will they be independent?

TechCrunch
1 min read Rewritten in plain language

Artificial intelligenceOpenAIAnthropicDario AmodeiAI Safety

Show what we removed Rules applied: A1×3 A3×8 C2 C5 D2×4 D3×14 D4×2 E3×3 F2×5 all 30 rules
  • Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research access to the company’s systems.
  • Alexander Meinke, head of research at Apollo Research, told TechCrunch.
  • Adam Gleave, CEO of FAR.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed.
  • When investigating the Hugging Face incident, OpenAI gave METR and Redwood about a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations.

4 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out if they’re to know whether they will function as independent watchdogs or vendors operating on the AI companies’ terms.

Headline check

Headline as published: Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

The one thing this headline claims is in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked

In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are aligned, and share their unvarnished findings with the world.

Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research access to the company’s systems. CEO Sam Altman said OpenAI also would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups.

Alexander Meinke, head of research at Apollo Research, told TechCrunch.

Now, evaluators that TechCrunch spoke to propose giving them access not just to the final model, but to intermediate versions, or “checkpoints,” from its lifetime of training. Adam Gleave, CEO of FAR.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed.

Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear.

“It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark,” Steidley said, comparing it to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.

When investigating the Hugging Face incident, OpenAI gave METR and Redwood about a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations.

Shortened to 1 minute of reading, this version reads 5.9 on the Niral Score.

You are reading our version, not theirs. This is TechCrunch's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
TechCrunchas they published this story 11 29 68 -0.1 49.5
Mundane Readneutralized from TechCrunch 8.7 26 68 -0.1 49.5

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works