Subscribe Sign in

Politics / United States

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

TechCrunch
1 min read Rewritten in plain language

Artificial intelligenceGovernment and PolicyCopyrightMicrosoftNew York Times

Show what we removed Rules applied: A1×3 A3×2 D1×5 D2 D3×6 D4×3 F2 all 30 rules
  • New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a threat to publications.
  • The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times at first alleged the firms violated copyright law by training generative AI models on its content.
  • An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

3 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Headline check

The one thing this headline claims is in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. One claim in this headline could be checked, so this is a narrow pass and not a thorough one. How this is checked

016
016 MTAPhotos / flickr, CC BY

New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a threat to publications.

Per the lawsuit, a top Microsoft executive privately described the companies’ AI training practices as “theft,” and OpenAI’s own leadership said its AI models posed an “existential threat” to the publishers and journalists whose work trained them.

The unsealed material also details how the companies allegedly got and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data.

The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times at first alleged the firms violated copyright law by training generative AI models on its content.

For example, Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

Other admissions cut against different pillars of the fair-use test: OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”

Shortened to 1 minute of reading, this version reads 6.7 on the Niral Score.

You are reading our version, not theirs. This is TechCrunch's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
TechCrunchas they published this story 10.8 20 53 -0.1 48
Mundane Readneutralized from TechCrunch 9.4 17 53 -0.1 48

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works