Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations “to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.” Executives at OpenAI, Google, and SpaceXAI have already rallied around Amodei’s plan, which has quickly become a central pillar of the emerging AI safety push.
Internet security experts say the labs need to focus on network security basics like logs and permissions, applying the same rigorous defenses they do for human users.
While alignment remains a concern, Sayash Kapoor, an AI researcher who will be a professor at UC Berkeley starting next year, argues that “marginal investments in control are more likely to be effective compared to those in alignment.”
The incidents that have spurred these concerns revolve around frontier models being asked to complete training tasks, usually cybersecurity evaluations, and then accessing the open internet and penetrating closed third-party systems in an attempt to do so.
“We as a profession know how to block access to the internet,” Avery Pennarun, the CEO of Tailscale, a security company, said.
In one case, where OpenAI agents took over a defunct German WikiForum to cheat on evaluations, the agents were active for weeks before anyone at the company appeared to notice. Security experts that TechCrunch spoke to said that real-time monitoring is key to preventing future break-outs, and that every agentic session should be time-limited and expire.