Sources
Stay ahead of AI art
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Discuss this with
Pick a companion and get their take on this story

Sofia follows the money, policy, and platforms shaping what creators can make.
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Pick a companion and get their take on this story
Newly unsealed court filings from the New York Times' copyright lawsuit against OpenAI and Microsoft show a Microsoft executive privately described AI data scraping as "the largest theft of labor in human history" — even as both companies built training datasets from paywalled Times content.
The filings, reported by TechCrunch, show that senior Microsoft figures were not operating under any illusion about what large-scale web scraping meant legally or ethically. One executive's phrasing — "largest theft of labor in human history" — is striking precisely because it appears in internal correspondence, not a legal filing or a public statement designed for an audience. It is the kind of language that surfaces when people believe they are not being watched.
The "doom loop" framing is equally pointed. According to Ars Technica, internal emails warned that AI systems displacing news publishers would eventually destroy the high-quality data pipelines those same AI systems depend on — a self-defeating cycle that both companies apparently recognized and proceeded with anyway.

Microsoft and OpenAI face escalating legal exposure as unredacted filings reveal internal awareness of scraping's ethical stakes.
Image: TechCrunch / TechCrunch AI
For anyone building with AI image generators or text models, the legal architecture underneath these tools has always been murky. This case sharpens the picture considerably. The same scraping logic that pulled Times articles also powered the massive image datasets — LAION, Common Crawl derivatives — that trained Stable Diffusion, Midjourney, and their successors. When Stability AI faced similar legal pressure in 2023 from Getty Images and a class of visual artists, the company's defense rested partly on the argument that scraping was standard industry practice. These Microsoft documents suggest the industry's own executives knew that argument had limits.
The practical stakes for image-model training are direct. If courts rule that scraping paywalled or copyrighted content without license constitutes infringement — and internal admissions of "theft" do not help defendants — the cost structure of building frontier models changes. Licensing deals become mandatory rather than optional, and that cost eventually flows downstream: into API pricing, platform subscription tiers, and the models available to creators.
The Universal Music and ElevenLabs licensed AI music remix platform announced earlier this year represents one version of what a post-scraping model economy looks like — rights negotiated upfront, revenue shared. Whether image-generation platforms move in the same direction depends heavily on how this case resolves.
Unredacted filings are a different category of document from redacted ones. Courts unseal them when they determine the public interest outweighs the defendants' confidentiality claims — a threshold that itself signals judicial skepticism. The NYT's legal team now has on record, in the defendants' own words, an acknowledgment that the conduct at issue caused harm.
"The largest theft of labor in human history."
— Microsoft executive, internal correspondence (via TechCrunch)
For creators using AI tools built on similar data pipelines, the question is no longer whether the legal risk is theoretical. It is how quickly courts move, and whether the major platforms — including those powering AI image generation today — have secured enough licensed data to survive an adverse ruling. The next substantive hearing in NYT v. OpenAI will determine whether these internal admissions survive a motion to limit their evidentiary use.