“Even putting aside the fact that human beings use LLMs to create, the Kadrey court’s application of the fourth fair-use factor is deeply flawed.” – DOJ statement
On September 1, the U.S. Department of Justice (DOJ) filed a Statement of Interest urging the U.S. District Court for the Southern District of New York to hold that using copyrighted written works to train large language models (LLMs) is fair use, arguing that a contrary result would distort copyright law, suppress innovation and weaken U.S. competitiveness and national security.
The statement was filed in reference to the multidistrict copyright litigation against OpenAI, although it specifically addressed claims by The New York Times and said its reasoning also applies to the related cases involving book authors and publishers. The government said the generative AI training process is different from collecting works and generating outputs, and that courts must evaluate each alleged use separately.
The statement argues that the cases could have consequences well beyond the parties, and that domestic AI development is important to economic growth and national-security functions. Requiring licenses for model training could create entry barriers that only the largest technology companies could absorb and could concentrate the market for LLMs and disproportionately benefit legacy media organizations with large archives, it argued.
At the same time, the government stressed that it was addressing only the use of works during training. It did not contend that the government authorized any challenged conduct, and it left open separate copyright questions involving the acquisition and storage of training data or outputs that reproduce protected expression.
The DOJ’s brief centered on the first and fourth statutory fair-use factors: the purpose and character of the challenged use and its effect on the market for the copyrighted work.
Training is “exceedingly transformative,” the government argued, because a model does not use an article to entertain or inform readers in the manner the author intended. Instead, it converts training data into numerical representations and learns statistical relationships involving vocabulary, syntax and knowledge, enabling the model to predict text and perform tasks ranging from editing and translation to generating new material.
Citing Google v. Oracle and the Second Circuit’s Authors Guild v. Google ruling, the DOJ said copying can be fair when it enables a new technological function, even if entire works are reproduced during an intermediate step. OpenAI’s commercial purpose should carry little weight, it added, because commercial uses can qualify as fair and the importance of commercialism diminishes as a use becomes more transformative.
The government acknowledged that particular outputs could present different issues if an LLM reconstructs and distributes copyrighted material. But those outputs should not determine whether the earlier training use is transformative, it said. Likewise, a small number of anomalous, reconstructive outputs could not justify relief that broadly restricts model training or exposes all LLM uses to massive liability.
With respect to market harm, the filing argued that copyright recognizes harm from significant substitutive competition, not competition in the abstract. Training alone makes no protected expression available to the public and therefore does not replace an article or impair its market in the legally relevant sense, the government said. Outputs that compete with articles but do not reproduce substantially similar protected expression also are not copyright substitutes merely because they occupy the same genre.
The government also sharply criticized the 2025 ruling in Kadrey v. Meta Platforms, in which a judge suggested that the copying of works for training LLMs will usually be found to be infringing, although in the case at hand, the plaintiffs’ arguments had missed the mark. “Even putting aside the fact that human beings use LLMs to create, the Kadrey court’s application of the fourth fair-use factor is deeply flawed,” said the DOJ. The court’s analysis improperly combined training and output into one continuous use and treated generalized competition as cognizable market harm, it added.
To illustrate its objection, the filing noted that Joan Didion typed Ernest Hemingway’s stories to study how his sentences worked. Under the Kadrey reasoning, it argued, Didion could have owed Hemingway whenever her later writing competed in the literary market—an outcome incompatible with the principle that copyright protects expression rather than ideas, methods, styles or influence.
The fourth fair use factor also requires consideration of public benefits, the government said. LLMs can assist writers, researchers, journalists and national-security officials, and can help independent publishers compete with established outlets. The filing also noted that Times writers reportedly use AI to conceptualize and edit articles.
Ultimately, said the DOJ, “it would be problematic—and legally incorrect—to impose broad copyright liability that would generally render training of AI models impermissible without licensing.”


Join the Discussion
No comments yet. Add my comment.
Add Comment