The writer's problem
A novelist wondering whether a chatbot trained on published fiction owes anything to the authors it learned from is asking a question that had, until 2025, no single official government answer. On 9 May 2025 the U.S. Copyright Office released the pre-publication version of Part 3 of its Report on Copyright and Artificial Intelligence, addressing whether training generative AI models on copyrighted works is fair use. The Office says it released this version early in response to congressional inquiries, and states a final version is expected without substantive changes, a status the Office's AI initiative page confirms alongside the report's other installments.
What the documents show
The report does not issue a blanket ruling. It states that several stages of building a generative AI system, including copying works to build a training corpus, implicate a copyright owner's exclusive rights, making fair use the decisive question. Applying the statute's four factors, the Office concludes that training uses 'are likely to be transformative,' but that transformativeness alone does not settle fairness: the outcome 'will depend on what works were used, from what source, for what purpose, and with what controls on the outputs.' Its own example draws a line at deployment: a model used for research or analysis is unlikely to substitute for the works it learned from, but using copyrighted works to generate expressive content that competes with them in the market, especially through illegal access, 'goes beyond established fair use boundaries.' The Office also considers government-mandated licensing premature given the growth of voluntary licensing, pointing instead to extended collective licensing for remaining gaps.
The editorial choice
For a publisher or agent negotiating a licensing or opt-out clause with an AI vendor, the report gives no shortcut past that clause. This is an editorial reading of the document: the fact-specific test it describes means a vendor's blanket assurance that its training is fair use should be checked against the report's own factors, namely what was used, how it was obtained, and what the output competes with, rather than accepted at face value.
What stays with the author
The report is a policy analysis advising Congress and the courts, not a ruling binding any company's practice; it says as much of its own earlier reports and the same caveat applies here. Whether a specific model's training on a specific author's books was fair use is left to litigation and to the licensing markets the report describes as still forming. An author's own leverage, such as registering copyrights and tracking licensing offers, remains theirs regardless of where the general analysis lands.
- Was the source of the training data an authorized purchase, a licensed dataset, or something the report would treat as illegally accessed?
- Does the AI system's output compete with the market for the works it trained on, or serve a different purpose such as analysis?
- Is a vendor's fair-use claim resting on the transformative nature of training alone, without addressing market effect?
Because this is a pre-publication text pending a final version, any specific page citation should be checked against the report the Office eventually finalizes.
Follow the source.
Establishes the report's fact-specific fair-use analysis of AI training and its pre-publication status.
Source date: 9 May 2025 · Retrieved: 16 Sept 2026
Confirms the 9 May 2025 release date of the pre-publication version and its place among the report's other parts.
Source date: Not established · Retrieved: 16 Sept 2026
Site publication is not established by an event date. Original record ID: 0030-bf-012. This local design review does not change its editorial status.