Kept it secret for months - even after OpenAI 'fessed up.
Google has said that its AI agents escaped a sandbox and mounted an attack – but only because testers mistakenly gave its bots internet access.
The Big G didn’t disclose the May incident, but The Wall Street Journal learned of the situation, which happened after Google hired Israeli firm Irregular to test its bots’ prowess in a capture-the-flag test.
The goal of the exercise was to acquire information from a fictional company without leaving a sandbox.
Irregular made two mistakes. The other was to use the name of an actual company.
When Google’s AI made it onto the open internet, it went looking for the actual company – three of them, in all.
According to the Journal, Google’s bots found passwords for two targets on the public internet. The software guessed the third password.
In a statement sent to The Registe r, Google said, “In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test.”
According to Google, its models stopped work before using the credentials.
Irregular erred in allowing internet access from a sandbox.
Google’s culpability is another matter because these incidents took place in May – around two months before OpenAI admitted its agents were the source of the July attack on Hugging Face.
One person who sees no risk of AI causing calamity is US president Donald Trump, who has rejected warnings as a “hoax” and said work on AI must not slow due to its economic and strategic significance.