Subscribe Sign in

Science

Researchers used Claude to hack OpenAI

Ars Technica
2 min read Rewritten in plain language

Artificial intelligenceSecurityAnthropicChatgptHacking

Changed after publishing · 1 edit
  • the report was edited: 260 words added, 4 removed.

Outlets edit stories after they go out, usually without saying so. We keep what we saw the first time.

Show what we removed Rules applied: A1 A2×2 A3 A6 C1 C2 D2×2 D3×3 D4 E3 F2×2 all 30 rules
  • Cyber researchers broke into OpenAI using its rival Anthropic’s software, highlighting vulnerabilities in the ChatGPT maker’s security as leading AI companies face scrutiny over safety.
  • The researchers had been given access to an Anthropic tool specifically designed for security professionals, and were paid for the work as part of a program to find vulnerabilities before they could be exploited by bad actors.
  • The latest incident occurred just two weeks after a group of more than 1,000 OpenAI agents escaped a test environment to hack the start-up Hugging Face, which caused awareness of AI’s ability to hack autonomously without human intent.
  • The disclosure on Thursday, first reported by The Wall Street Journal, came as Anthropic published a new set of data that showed a rapid increase in how much the lab used AI to develop its new models.
  • The company said that as AI systems become more powerful, they were “increasingly being used to build the next version of themselves.”

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

They exploited a flaw in the set-up of OpenAI’s community forum Discourse, and used it to gain access to internal sign-ons and eventually an OpenAI employee’s ChatGPT account.

Headline check

All two things this headline claims are in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. How this is checked

Read the full reportHide the full report2 min

Cyber researchers broke into OpenAI using its rival Anthropic’s software, highlighting vulnerabilities in the ChatGPT maker’s security as leading AI companies face scrutiny over safety.

A small cyber security group gained access to an OpenAI employee’s ChatGPT account, which permitted them to read private software information and suggest changes.

The researchers had been given access to an Anthropic tool specifically designed for security professionals, and were paid for the work as part of a program to find vulnerabilities before they could be exploited by bad actors.

Their ability to swiftly break into one of the world’s two leading AI labs again raises concerns about OpenAI’s security amid rising worries about powerful models being used by hackers and foreign adversaries.

The US has in recent months addressed with how to manage the vetting and release of the latest models, including temporarily blocking some Anthropic tools.

The latest incident occurred just two weeks after a group of more than 1,000 OpenAI agents escaped a test environment to hack the start-up Hugging Face, which caused awareness of AI’s ability to hack autonomously without human intent.

The three researchers from Hacktron AI, a small security company, were paid $6,500 by OpenAI as part of a bug bounty program, a common practice where tech companies pay ethical hackers to test their security.

They exploited a flaw in the set-up of OpenAI’s community forum, which is hosted by a third-party, Discourse, and used it to gain access to internal sign-ons and eventually an OpenAI employee’s ChatGPT account. This ChatGPT account had access to internal code through GitHub.

“We thank the researchers for contacting us and sharing their findings,” OpenAI said, adding that it had fixed the issues. Anthropic declined to comment. Hacktron did not immediately respond.

The disclosure on Thursday, first reported by The Wall Street Journal, came as Anthropic published a new set of data that showed a rapid increase in how much the lab used AI to develop its new models.

It said 26 percent of research and development work was “led by” its Claude model, up from 1 percent in March, meaning that AI completed the majority of tasks based on human instruction and under supervision.

The company said that as AI systems become more powerful, they were “increasingly being used to build the next version of themselves.”

Anthropic said it shared the data to help the public “understand how close the world is to reaching recursive self-improvement,” the point at which AI can train and improve itself or new models.

This threshold is at the heart of concerns that AI systems will become more difficult to oversee, leading to a loss of human control.

Its models did not yet operate autonomously for any of the research it studied, Anthropic added. On 90 percent of tasks, AI “collaborates” with a human and does large chunks of work.

You are reading our version, not theirs. This is Ars Technica's report with its verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
Ars Technicaas they published this story 13.5 11 47 -0.3 50.8
Mundane Readneutralized from Ars Technica 9.3 10 47 -0.2 50.8

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works