Subscribe Sign in

World / China

AI hallucination of Chinese nuclear components almost led to US military attack

Ars Technica
3 min read Rewritten in plain language

Artificial intelligenceChinaDepartment of DefenseDodHegseth

Show what we removed Rules applied: A1 A3×3 C1 D3×3 D4×2 E3×7 F2×4 all 30 rules
  • The US avoided boarding a Chinese ship based on an “entirely false” US intelligence report generated with the help of AI tools, according to a CNN report.
  • The US military was preparing to intercept and board the ship, with air support, before officials discovered a chatbot used in generating the report had “inaccurately identified the material the ship was carrying.”
  • CNN’s report said the analyst in question used a chatbot to analyze intelligence reports regarding the Chinese ship’s manifest, leading to the near-disastrous result.
  • Last December, the Department of Defense announced it would use Google’s Gemini for Government as the basis for its bespoke “GenAI.mil” platform.
  • In June, a Pentagon representative bragged to Congress that they use generative AI to help create congressionally mandated reports, and that 1.5 million active DoD personnel have used the military’s generative AI tools.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

CNN’s report said the analyst in question used a chatbot to analyze intelligence reports regarding the Chinese ship’s manifest, leading to the near-disastrous result.

The report’s most important sentence, shortened and in plain words. How

Headline check

The one thing this headline claims is in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked

Read the full reportHide the full report3 min

The US avoided boarding a Chinese ship based on an “entirely false” US intelligence report generated with the help of AI tools, according to a CNN report.

That erroneous intelligence, submitted by a US Special Operations Command analyst, suggested the Chinese ship was transporting nuclear arms program components through the Middle East, according to “four sources familiar with the episode” cited by CNN. The US military was preparing to intercept and board the ship, with air support, before officials discovered a chatbot used in generating the report had “inaccurately identified the material the ship was carrying.”

One source told CNN the AI-powered outcome “almost started a war.”

CNN’s report said the analyst in question used a chatbot to analyze intelligence reports regarding the Chinese ship’s manifest, leading to the near-disastrous result. That chatbot “fused together open-source intelligence with secret signals intelligence in government holdings,” and that information was packaged into an intelligence report that almost set off a chain of events, according to CNN.

The near-miss is one of the more potent and consequential instances of a hallucinating AI ruining the reliability of a professional report. Since “hallucinating” became the Cambridge Dictionary’s word of the year in 2023, we’ve seen prominent examples of non-fiction authors, journalists, academic researchers, judges, doctors, police departments, corporate call centers, and more getting taken in by AI tools that make something up when their training data doesn’t provide sufficient context. And despite some adorable attempts at “do not hallucinate” prompts, some researchers suggest that it may be impossible to prevent LLMs from hallucinating altogether.

One would hope the US military would be aware of these kinds of problems when relying on AI for analysis of intelligence reports. But the Department of Defense in January rolled out an “AI acceleration strategy” that sought to “make all appropriate data available across federated IT systems for AI exploitation, including mission systems across every service and component.”

“AI is only as good as the data that it receives, and we’re going to make sure that it’s there,” Defense Secretary Pete Hegseth said in rolling out that initiative.

Last December, the Department of Defense announced it would use Google’s Gemini for Government as the basis for its bespoke “GenAI.mil” platform. Last month, the department added Grok for Government as an option on the platform. Anthropic also offers a customized version of Claude for US spy work.

In June, a Pentagon representative bragged to Congress that they use generative AI to help create congressionally mandated reports, and that 1.5 million active DoD personnel have used the military’s generative AI tools.

In 2023, a State Department “Declaration on Responsible Military Use of Artificial Intelligence and Autonomy” stressed that “principled” use of AI by armed forces “should include careful consideration of risks and benefits, and it should also minimize unintended bias and accidents.” That report also urged that “accountable” use of AI systems must always involve “a human in the loop, a responsible human chain of command and control.”

In the years since then, though, we’ve seen autonomous attack drones used in the Russian conflict in Ukraine and tested by NATO-backed military contractors. In March, the Department of Defense blacklisted Anthropic over the company’s opposition to its models’ use in autonomous weapons systems, a move that a federal judge said last month was “unlawful retaliation in violation of the First Amendment.”

The reported near miss comes as extinction-level warnings from AI researchers have led to a newly prominent national conversation on AI safety, including calls for regulation and coordinated research “pacing” from leading frontier AI labs.

You are reading our version, not theirs. This is Ars Technica's report with its verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
Ars Technicaas they published this story 8.8 13 86 -0.3 41.7
Mundane Readneutralized from Ars Technica 6.1 12 86 -0.1 41.7

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works