Startup says it’s learned from these mistakes and that they shouldn’t happen again … which is just what Zuck has said about 100 times.
OpenAI has revealed another six occasions on which its AI software behaved or did things.
The startup added the incidents to its misalignment reports page on Wednesday evening, Pacific Time, and described them as follows:.
Signing up for disposable emails and searching GitHub for leaked API keys.
Uploading files to the internet to cite them.
Unsanctioned Artifactory writes and cross-sample communication.
Unauthorized communication via temporary file hosting services.
The second incident took place during training for the Sol 5.6 model.
The file-uploading SNAFU also involved an unreleased model that OpenAI tasked with finding the IDs and names of lakes larger than 5,000,000 square meters.
The agent queried a mapping service using Python and was able to access the records it wanted, then “wanted the browser tool to show those records so it could cite them.” The agent saved the data to a local text file but the browser couldn’t access it.
The agent therefore spun up a web server, but the browser rejected the localhost URL.
OpenAI detected unsanctioned Artifactory writes when its agents attacked Hugging Face, an incident that saw agents leave each other notes in the shared package repository.
Each incident report includes OpenAI’s response to the discovery that its tech went bad, and they mostly say the company has figured out what went wrong and thinks it has made changes that will mean they don’t happen again.