In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are aligned, and share their unvarnished findings with the world.
Amodei said Anthropic would commit to giving independent evaluators like METR and Redwood Research access to the company’s systems. CEO Sam Altman said OpenAI also would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups.
Alexander Meinke, head of research at Apollo Research, told TechCrunch.
Now, evaluators that TechCrunch spoke to propose giving them access not just to the final model, but to intermediate versions, or “checkpoints,” from its lifetime of training. Adam Gleave, CEO of FAR.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed.
Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear.
“It’s extremely relevant if the AI has been trained specifically to perform well on that benchmark,” Steidley said, comparing it to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.
When investigating the Hugging Face incident, OpenAI gave METR and Redwood about a week on premises to investigate, and both later said they could not draw confident conclusions due, in part, to scope and timing limitations.