Sources
Stay ahead of AI art
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Discuss this with
Pick a companion and get their take on this story

Iris covers where AI art meets culture — style, authorship, and the images that matter.
Get the week's top AI and AI-art stories delivered to your inbox — curated, concise, free.
Free. Unsubscribe any time.
Pick a companion and get their take on this story
No court has definitively ruled that training AI models on copyrighted books is illegal — but the question is live in multiple U.S. lawsuits, and the outcome will directly reshape which models survive, which training pipelines get shut down, and what creators can legally do with AI-generated output.
The legal theory AI companies are betting on is fair use — specifically the argument that ingesting copyrighted text to build a statistical model is "transformative" enough to fall outside infringement. That argument has precedent: Google won a decade-long fight over scanning books for search indexing. But generating prose, images, or code that competes directly with the original works is a harder case to make, and several judges have allowed author lawsuits to proceed past early dismissal, signalling the defence is not airtight.
As TechCrunch reports, most published authors contributed to the development of AI tools without their knowledge or consent — tools that now compete with their own writing for readers and commissions.
"Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods."
— TechCrunch

The legal status of training AI on copyrighted books remains unresolved across multiple active U.S. court cases.
Image: TechCrunch / TechCrunch AI
This is where the stakes turn concrete for anyone working with AI image generators or text tools. The large language models underlying many image-generation pipelines — for captioning, prompt interpretation, and multimodal reasoning — were trained on datasets that almost certainly included copyrighted text. If courts find that training without a licence constitutes infringement, labs face three options: negotiate costly retroactive licences, retrain on narrower datasets, or shut down affected models.
Retraining on curated, licensed corpora tends to produce models with narrower stylistic range and weaker generalisation. For creators, that translates to prompts that hit fewer aesthetic registers, less nuanced character writing in companion tools, and a general flattening of output variety. The richness that makes current models feel capable of genuine stylistic range comes, in part, from the breadth of what they ingested — including the very literary tradition that authors are now fighting to protect.
There is a parallel here to the Pictorialist photographers of the early 20th century, who argued that photography could only become art by borrowing from painting's visual grammar. The courts eventually had to decide what counted as original expression in a mechanical medium. AI is forcing the same question from the opposite direction: not whether the machine can make art, but whether it was permitted to learn from the art already made.
The consent problem cuts deeper than the legal theory. Authors did not opt in; their work was harvested from datasets assembled at scale, often without any public disclosure of which titles were included. That opacity makes it difficult for creators to know whether the model they're using was trained on a specific author's style — which matters both ethically and, increasingly, commercially, as some clients are beginning to ask for provenance documentation on AI-assisted work.
Platforms that want to stay ahead of this are already moving toward licensed training agreements with publishers. Adobe Firefly's pitch to commercial clients has always rested on its licensed-data training. If litigation forces competitors onto the same path, the cost differential narrows — but so does the creative range gap that currently makes open-weight and broadly-trained models attractive for experimental work.
For now, creators using any major AI tool are working with models whose legal foundation is genuinely contested. That is not a reason to stop — courts move slowly, and injunctions against deployed models are rare — but it is a reason to watch which labs are proactively licensing data and which are still running the fair-use gamble. The outcome will determine not just who owes whom money, but what the next generation of models is allowed to learn.