Subscribe Sign in

Politics / United States

Microsoft executive called OpenAI's web scraping the 'largest theft of labor in human history'

Engadget
1 min read Rewritten in plain language

Artificial intelligence

Show what we removed Rules applied: A3 D1×2 D3×2 F2 all 30 rules
  • Newly released court documents show that multiple Microsoft execs thought AI training was an 'existential threat' to journalism.
  • Microsoft's director of Applied Science, Dr. Brent Hecht, said OpenAI's work was akin to the "largest theft of labor in human history" that could create a "doom loop," while OpenAI exec Nick Turley said it represented an "existential threat to publishers."
  • Along with the quotes, the unredacted materials also show how OpenAI and its partners got training content by bypassing paywalls, and built training datasets by scrapping millions of documents and erasing copyright notices from training data.
  • In another, an OpenAI software engineers told colleagues in 2023 that "no matter how prominently we show the links, users won't click."

4 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

That makes the comments from Microsoft and OpenAI executives interesting.

Headline check

All two things this headline claims are in the report.

Figures, names and quoted words in the headline, looked for in the report itself — not in the summary above. How this is checked

Newly released court documents show that multiple Microsoft execs thought AI training was an 'existential threat' to journalism.

Executives from OpenAI and Microsoft were reportedly worried about ChatGPT training that scraped millions of news articles, The New York Times reported. Microsoft's director of Applied Science, Dr. Brent Hecht, said OpenAI's work was akin to the "largest theft of labor in human history" that could create a "doom loop," while OpenAI exec Nick Turley said it represented an "existential threat to publishers." Only snippets from the documents were made public without any context around them.

Along with the quotes, the unredacted materials also show how OpenAI and its partners got training content by bypassing paywalls, and built training datasets by scrapping millions of documents and erasing copyright notices from training data.

Several such lawsuits have already swung in favor of AI companies, but judges have pointedly stated that their rulings were made because the law around AI use has yet to be settled.

In another, an OpenAI software engineers told colleagues in 2023 that "no matter how prominently we show the links, users won't click." OpenAI's Turley confirmed that, saying AI products are "largely substitutive" to journalism.

Shortened to 1 minute of reading, this version reads 3.7 on the Niral Score.

You are reading our version, not theirs. This is Engadget's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
Engadgetas they published this story 10.5 6 58 -0.6 42.3
Mundane Readneutralized from Engadget 10.2 6 58 -0.6 42.3

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works