The DOJ Just Drawn a Line in the AI Copyright War — and It’s Not Where You Think

On September 1, the U.S. Department of Justice filed something that flew under most radar screens: a Statement of Interest in the OpenAI v. New York Times copyright litigation. Most coverage reduced it to “the government sides with OpenAI.” That’s technically true, but it misses what’s actually happening here.

San Francisco Hall of Justice building representing the legal system's intersection with AI technology
Image: San Francisco Hall of Justice by Dllu via Wikimedia Commons (CC BY-SA 4.0)

This isn’t just another amicus brief. It’s the federal government’s first direct intervention in the wave of AI copyright cases, and it lays out a detailed, four-point legal framework that would fundamentally narrow how courts evaluate whether training AI on copyrighted material counts as infringement. The brief doesn’t just take a side — it tries to define the rules of the game.

As someone who manages ICT systems and has watched the AI copyright debate evolve from “will they sue?” to “whose law wins?” over the past two years, this filing caught my attention for reasons beyond the headline. Here’s what the DOJ actually argued, why it matters, and what it tells us about where this fights heads next.

What the DOJ Actually Said — Four Arguments That Matter

The Statement of Interest, filed in the Southern District of New York, addresses a specific question: does using copyrighted works to train an AI model constitute fair use under current law? The government’s answer is unambiguous — training, standing alone, does not violate copyright. But the reasoning behind that answer is where things get interesting.

The DOJ structured its argument around four points that, taken together, would significantly reshape how courts handle AI training cases:

1. Training and outputs are different legal questions

The government argues that courts should separate internal training from public-facing outputs. If an AI model produces something that infringes, that’s an output problem with output-focused remedies. It shouldn’t taint the training stage itself. This matters because plaintiffs often point to allegedly infringing outputs — summaries, memorized text, similar-sounding passages — to argue that the training was illegal. The DOJ is essentially saying: fix the output if it’s broken, but don’t punish the training because of what the model sometimes spits out.

This is a practical argument. Models produce millions of outputs. Some will inevitably resemble their training data. If every similar output retroactively makes the training unlawful, the legal standard becomes impossible to satisfy.

2. LLM training is “highly transformative” — and in a different category

The brief characterizes training as using text to identify linguistic patterns, relationships, and predictive signals — not to substitute for the expressive purpose of the original works. The government’s framing: the purpose of the copying (to build an intelligent, interactive model) differs in kind from the purpose of the copied work (to entertain or educate a reading audience).

This is the first fair-use factor, and it’s where the government tries to shift the analysis away from “did they copy a lot?” toward “what did they do with what they copied?” It’s a familiar argument in fair-use law — Google Books won on similar grounds — but applying it to LLMs at this scale is new territory.

3. Commercial use isn’t automatically bad

The Administration acknowledges that most LLM products are commercial. But it argues that commerciality should carry less weight when the use is transformative and doesn’t expose protected expression to the public. This is designed to keep the first fair-use factor focused on the purpose of training rather than the business model of the company doing it.

It’s a direct response to one of the strongest arguments against AI training: that these are profitable companies using other people’s work to build commercial products. The DOJ’s counter is that the commercial nature matters less when the copying itself doesn’t put the original works in front of consumers.

4. Market harm requires substitution, not just competition

For the fourth fair-use factor — market effect — the government takes a narrow view. It rejects the theory that AI-generated content in the same genre, or broader competition from AI-enabled content, constitutes copyright harm by itself. Instead, it argues that harm requires substantial similarity, substitution for protected expression, or impairment of a specific derivative market tied to the works at issue.

Translation: the fact that AI can now write news articles, summarize books, or generate marketing copy doesn’t, by itself, mean copyrighted works were harmed. There needs to be a closer connection — the output needs to substitute for a specific copyrighted work, not just compete in the same market.

The Real Story: The DOJ Is Disagreeing With Its Own Copyright Office

Here’s what most coverage missed. The DOJ’s position isn’t just a pro-industry brief — it’s a direct challenge to the U.S. Copyright Office’s own analysis.

In May 2025, the Copyright Office released its “Part 3: Report on Generative AI Training” under Register of Copyrights Shira Perlmutter. That report treated AI training as a fact-intensive fair-use question, not a categorical answer. It recognized that fair use could protect some training uses — particularly noncommercial research — but warned that copying expressive works from pirate sources to generate competing content, especially where licensing is available, probably doesn’t qualify.

The Copyright Office’s approach was nuanced. It treated the source of training materials, the purpose of the model, output safeguards, and emerging licensing markets as relevant factors. It didn’t say “training is always fair use” or “training is always infringement.” It said: it depends on the facts.

The DOJ’s September 1 brief takes a different path. It asks the court to isolate training as an internal, transformative use and treat outputs, acquisition, and licensing as separate issues. That’s a sea change from the Copyright Office’s spectrum-based analysis. It shifts from “let’s look at all the facts” to “here’s a rule that training on its own should be fine.”

This isn’t a subtle difference. It’s two arms of the same government pointing in materially different directions. The Copyright Office — the agency with the deepest expertise in copyright law — said this is a fact-specific inquiry. The Department of Justice — the agency that litigates these cases — says it should be a categorical rule.

The tension between these positions isn’t accidental. The Copyright Office report came out under Perlmutter, who the Trump Administration attempted to remove (the Supreme Court denied the government’s stay application on June 30, keeping her in place pending litigation). The DOJ brief came out after that fight started. The policy context matters.

The National Security Framing Is the Most Interesting Part

Buried in the brief is an argument that goes beyond copyright law entirely. The government ties AI development to national security, economic competitiveness, scientific progress, and U.S. leadership in emerging technology. It warns that judicially imposed licensing requirements could slow domestic AI development, strengthen foreign competitors, and concentrate the market among companies that can afford large-scale licensing costs.

Commerce Secretary Howard Lutnick took this even further in public comments at a G20 meeting, telling officials their countries should embrace fair use and allow AI companies to train on creators’ work while finding ways to “protect artists.”

This is a policy argument dressed up as a legal one. It’s saying: if copyright law makes it too hard to train AI models in the United States, the work will move somewhere else, and we’ll be worse off. That’s a legitimate policy concern — but it’s not the same thing as a copyright analysis. The brief uses national security rhetoric to bolster a fair-use argument, which is a notable rhetorical move.

I’ve spent enough time in ICT management to know when an argument is technical and when it’s political. This is political. The government is telling courts: rule against AI training too aggressively, and you’re not just interpreting copyright law — you’re handicapping American competitiveness.

What This Means for AI Developers and Content Creators

For AI companies and the developers building on top of them, the filing is a meaningful boost. It gives them a clear federal policy statement to cite in summary judgment briefing. It strengthens the argument that training-stage copying can be fair use when the model doesn’t expose protected expression. It provides a roadmap for how to frame the legal defense.

But — and this is important — it is not a blanket immunity. The DOJ’s own brief acknowledges that claims based on pirated-source allegations, output memorization, substantial similarity, or misleading use of copyrighted content may remain viable depending on the facts. The filing doesn’t resolve the cases. It gives defendants a better argument, not a winning one.

For content creators, publishers, and copyright holders, the filing is a warning. The government is telling courts not to treat AI training as infringement absent a specific substitution harm. That narrows the theories available to plaintiffs. Copyright holders who want to protect their work need to think about what evidence they can preserve: training-data sources, model safeguards, licensing history, output controls, and any substantiated instances of reproduction or substitution. Those facts — not broad arguments about whether AI is good or bad — will drive the next phase of litigation.

The practical implication is that contractual frameworks matter more, not less. If copyright law becomes less protective of training-stage copying, content owners need stronger contractual and technical protections — terms of service, access controls, content identification tools, licensing arrangements. The legal default is shifting, and the response has to shift with it.

The Licensing Question the Brief Doesn’t Answer

One of the most notable things about the DOJ’s position is what it leaves unresolved. The brief argues that courts shouldn’t convert unresolved policy questions into a mandatory licensing regime. It says licensing should be primarily a legislative or commercial issue, not a judicial fair-use requirement.

That’s a defensible position — courts aren’t great at designing licensing markets — but it kicks the can down the road. The Copyright Office’s Part 3 report described voluntary licensing markets as developing and suggested collective approaches could be considered if market gaps persist. The DOJ’s brief doesn’t engage with that framework on its merits. It just says: don’t make courts do it.

The problem is that “eventually Congress should figure this out” is not a satisfying answer for publishers whose archives are being used to train models right now. The Copyright Office spent years studying this. The register was fired (and fought it). The Supreme Court kept her in place. And now the DOJ is telling the courts to wait for Congress while simultaneously arguing that training should presumptively be fair use.

The system is sending mixed signals because it is sending mixed signals. The Executive Branch wants AI development to proceed with minimal copyright friction. The Copyright Office — still staffed by the career professionals who studied this for years — sees a more nuanced picture. Nobody knows which voice the courts will listen to.

What Happens Next

The Statement of Interest isn’t binding. It’s advisory. But it’s from the U.S. government, in a case where the U.S. government’s views matter, and it gives defendants a powerful new argument to deploy at summary judgment. Expect OpenAI and other AI defendants to cite it heavily.

Expect copyright plaintiffs to respond by distinguishing the briefing — arguing that the specific facts of their cases (pirated sources, memorized outputs, market substitution) fall outside the broad rule the DOJ is advocating. The brief actually helps them frame that response, because it acknowledges that those factual scenarios may still be viable.

The bigger picture is that the legal framework for AI training is being built in real time, across multiple cases, multiple courts, and now multiple branches of the federal government that don’t fully agree with each other. The Copyright Office says “it depends on the facts.” The DOJ says “training alone isn’t infringement.” The courts will have to pick a path, and that path will shape how AI companies operate for years.

For developers and businesses building with AI, the prudent move is to pay attention to the specific facts that the DOJ flagged as still relevant — data sources, output controls, licensing history — and document them. The broad arguments about whether AI training is good for society are being made loudly. The factual record that will actually decide these cases is being built quietly, in document preservation and model governance policies. That’s where the real battle is.

Relevant Bleuken coverage: The AI security landscape is moving just as fast as the copyright one — see our coverage of the AI security market hitting $2.8 billion and AIR’s $50M raise for AI agent security. On the Google side, the Gemini 3.8 Flash pricing dynamics show how aggressively the major players are competing on cost — which is part of why the DOJ’s competitiveness argument carries weight.

Filed under Tech & Gadgets
Last Update: September 26, 2026 by Felix AlterEgo
0 0 votes
Article Rating
Subscribe
Notify of
guest

This site uses Akismet to reduce spam. Learn how your comment data is processed.

0 Comments
Newest
Oldest Most Voted