Sources
Stay ahead of AI art
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Discuss this with
Pick a companion and get their take on this story

Sofia follows the money, policy, and platforms shaping what creators can make.
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Pick a companion and get their take on this story
Newly unsealed court filings in the New York Times v. OpenAI lawsuit reveal that both OpenAI and Microsoft privately documented — in their own internal records — that their AI training practices risked creating a "doom loop" that would degrade the open web, and that at least one Microsoft executive characterized the scraping as the "largest theft of labor in human history."
The phrase "doom loop" appears in the companies' own documentation, not in the Times' complaint. The logic is straightforward and self-indicting: AI systems scrape the web to train on content, then return answers directly to users rather than sending them to source sites, which starves publishers of the traffic and ad revenue needed to produce new content, which degrades the web those same AI systems need to train on next. The loop tightens with each generation of models.

A page from the unsealed court filings showing Microsoft's Director of Applied Science describing AI data acquisition as 'an astonishing theft of unprecedented proportions.'
Image: The Verge / The Verge AI
Satya Nadella's deposition is particularly pointed. According to The Verge, Nadella agreed under oath that chatbot interactions "substituted" for visiting the underlying websites — a concession that directly supports the Times' theory of harm. That's not an allegation from a plaintiff's brief; it's the defendant's CEO on the record.
On OpenAI's side, internal records show the company knew by 2021 that its API might reproduce existing content verbatim, and that preventing memorization was important "for fair use compliance and minimizing" legal exposure. The company also noted it was circumventing paywalls and violating terms of service when acquiring training data — and that individuals within the organization treated those issues as acceptable costs.
"This case is about, as Microsoft's Director of Applied Science put it, 'an astonishing theft of unprecedented proportions'; perhaps the 'largest theft of labor in human history.'"
— Court filing, NYT v. OpenAI
This isn't the first time a major AI company has faced a reckoning over training data sourcing. When Stability AI confronted similar allegations in 2023, the company's position was essentially that scraping was legally unresolved territory. What makes the OpenAI-Microsoft situation structurally different is that the damaging language comes from inside the house — executives and applied scientists, not adversarial experts.

An unsealed filing showing OpenAI's 2021 recognition that its API might output existing content verbatim, flagged as a fair-use compliance concern.
Image: The Verge / The Verge AI
For AI-art creators, the downstream stakes are real. The training datasets behind image-generation models face analogous questions about scraping visual content from the web without compensation. If the NYT case produces a ruling — or a settlement with licensing terms — that establishes a precedent for text, expect that framework to be applied to image data almost immediately. Platforms that have already moved toward licensed or synthetic training data would gain a structural advantage; those still relying on unfiltered web scrapes would face the same liability exposure now being litigated in Manhattan.
The Charmloop catalog of available models already reflects some of this shift, with providers increasingly distinguishing between models trained on licensed versus scraped datasets — a distinction that is about to matter legally, not just ethically.
Our earlier coverage of the unredacted Microsoft filings laid out the initial document disclosures; these newly unsealed records go further, placing the "doom loop" framing and the memorization acknowledgments directly into the court record.
The case has no trial date set yet. The next determinative moment is how the court rules on summary judgment motions — a decision that will either force a settlement or send the "doom loop" documents in front of a jury.