BLOG

Thoughtful Insights On The World We Live In

BL July Blog

AI and Fair Use

What the Latest Court Decisions Mean for Developers and Copyright Holders

By Allysa Combs

Artificial intelligence is reshaping industries at a pace the legal system was never built to absorb. In copyright law, courts are now being asked to decide a question no one has had to answer before: can the companies building large language models (LLMs) be held liable for using copyrighted works to train their systems?

The short answer is that courts are still working it out. But the decisions issued so far reveal a legal landscape in motion, one where developers, content creators, and copyright holders all have legitimate stakes in the outcome.

Here is where things stand.

The Core Legal Question

Training a large language model requires enormous volumes of text. That text is often copyrighted. When a model ingests copyrighted material without a license, courts must decide whether that use constitutes direct infringement or whether it qualifies as fair use under federal law.

The fair use analysis turns on four statutory factors:

  • the purpose and character of the use, including whether it is commercial or nonprofit educational;
  • the nature of the copyrighted work;
  • the amount and substantiality of the portion used relative to the work as a whole; and
  • the effect of the use on the potential market for or value of the copyrighted work.

Two cases out of the Northern District of California have become the leading decisions on this question. They reached similar outcomes, but through very different reasoning, and that difference matters.

The Two Leading Cases

Bartz v. Anthropic PBC, 3:24-cv-05417 (N.D. Cal. filed August 19, 2024)

Judge Alsup’s decision in Bartz centered on transformative use. The court found that training a model on copyrighted works was “exceedingly transformative” because the model was not reproducing those works; it was using them to map statistical relationships between text fragments in order to generate entirely new outputs.

The opinion also drew a meaningful distinction between lawfully acquired training data and pirated sources, suggesting that the origin of the material is a factor courts will scrutinize.

Overall, Bartz is the more favorable decision for AI developers.

Kadrey v. Meta Platforms, Inc., 3:23-cv-03417 (N.D. Cal. filed July 7, 2023)

Judge Chhabria took a different approach in Kadrey, focusing heavily on market harm. The court acknowledged that LLM training may well be transformative, but questioned whether that alone is enough when the company building the model stands to earn billions of dollars while potentially flooding the market with AI-generated content that competes directly with the works used to train it.

Critically, the Kadrey ruling did not explicitly find that Meta’s use was unlawful. While the court ruled in favor of Meta, the court emphasized that this did not mean that Meta’s use of the copyrighted materials was lawful, but instead stood for the proposition that plaintiffs had not made the right arguments.

A Warning Sign: Thomson Reuters v. Ross Intelligence

Read together, Bartz and Kadrey confirm that transformative use matters, but market harm may ultimately be the deciding factor.

Outside California, the District of Delaware issued a decision in Thomson Reuters v. Ross Intelligence, 694 F. Supp. 3d 467 (D. Del. 2023) that cuts in a different direction. Ross used Westlaw headnotes to train a competing legal research AI. The court found that the use was not transformative and that it posed direct market harm to Thomson Reuters.

Here is how the four fair use factors broke down:

  • Purpose and Character: Favored Thomson Reuters. Ross’s use was commercial and served the same function as the original product, making it difficult to characterize as transformative.
  • Nature of the Work: Favored Ross. This factor rarely drives fair use outcomes and did not here.
  • Amount Used: Favored Ross. The headnotes did not appear directly in outputs, so the portion accessible to the public was limited.
  • Market Effect: Favored Thomson Reuters heavily. Ross’s product was designed to replace Westlaw, and the court found significant potential harm to both the primary and derivative markets.

The court did note that its reasoning may not extend to all generative AI cases, acknowledging that the legal landscape is rapidly evolving. The case is now on appeal to the Third Circuit, which will be the first federal appeals court to weigh in on fair use and LLM training.

That decision will be worth watching closely.

The Next Battleground: What the Model Outputs

The training input cases are only the first wave. Courts are beginning to look at what AI models actually produce, and whether those outputs infringe on the works they were trained on.

The case to watch here is Disney Enterprises, Inc. et al. v. Midjourney, Inc., 2:25-cv-05275, (C.D. Cal. filed June 11, 2025). The complaint focuses on AI-generated images that allegedly reproduce copyrighted characters, including recognizable figures from major entertainment franchises. The core issue is not what went into training the model, but what came out of it.

Based on the direction of the cases above, market harm will likely be a central question in output cases as well. Two types of harm are likely to matter most:

  • whether the AI generates near-identical replicas of protected works; and
  • whether AI-generated substitutes dilute the market for original works.

What This Means in Practice

While there is uncertainty in the case law, both developers and copyright holders can take concrete steps now.

For LLM Developers:

  • Monitor outputs to ensure they are not reproducing copyrighted works verbatim or in near-identical form.
  • Build guardrails that limit the extent to which protected material can surface in model outputs.
  • Use only lawfully obtained training data. This means licensed commercial datasets, public domain works, permissively licensed content, and company-owned material. Avoid pirated works, shadow libraries, leaked datasets, and unauthorized web scraping of news articles, blog posts, and images.
  • Maintain detailed records of what was used to train the model and the legal basis for that use.

For Copyright Holders:

  • Include a proactive copyright notice that explicitly reserves the right to license the work for AI training purposes. Example: “The author reserves all rights to license uses of this work for generative AI training and development of machine learning models.”
  • Monitor how your work is being used across AI platforms.
  • Block AI bots from scraping your website by updating your robots.txt file. This text file instructs bots about which pages or content they are permitted to access.

The Bottom Line

No federal appeals court has weighed in yet on fair use and AI training data. Until one does, the law remains unsettled. But the trajectory of the district court decisions points clearly toward one conclusion: market harm is the factor courts will weigh most heavily.

For developers, that means the legal path forward runs through responsible data sourcing, guardrails on outputs, and thorough documentation. For copyright holders, it means being proactive, not reactive, about protecting and monetizing your work in an AI-driven world.

The Third Circuit’s decision in Thomson Reuters will likely be the most significant development to watch. When a federal appellate court finally speaks, the rules of this space will begin to take a clearer shape.

Until then, the best protection is preparation.

How We Can Help

Bagchi Law works with technology companies, founders, and content creators to navigate the intersection of intellectual property and emerging technology. Whether you are building AI-powered products, managing a content library, or evaluating your exposure in this rapidly shifting legal environment, we can help you assess your risks and develop a strategy that works for your business.

If you have questions about how these developments affect your work, we are happy to talk through your situation.

Related

NIL Contracts, Exclusivity, and Tampering

By Deonta Woods What Every College Athlete Needs to Know Before Signing Anything You finally did it. Your first Name, Image, and Likeness deal lands in your inbox.The brand is…

Independent Research, Reverse Engineering, and Trade Secrets: What the “Cracked” Coca-Cola Formula Teaches Us

By Deonta Woods For more than a century, the formula for Coca-Cola has been one of the most famous and fiercely guarded trade secrets in the world. Since 1886, the…

Avoiding “Naked Licensing”: What You Need to Know Before Licensing Your Trademark

By Deonta Woods Licensing a trademark may seem straightforward, but without the right protections in place, you could unintentionally put your trademark at risk. One of the most common and…

The Impact of Tariffs on U.S.-India Business

For international business owners and investors in 2025, no word has caused as much headache as the word “tariffs.” The concept of both raising money and protecting local industry by taxing imports is not…

Major QSBS Updates: What Founders and Investors Need to Know Under the 2025 Act

The Small Business Investment Act of 2025 makes sweeping changes to Qualified Small Business Stock (QSBS). Effective July 5, 2025, these reforms are designed to make QSBS even more attractive…

Why Growing Companies Choose Fractional CISOs: An Interview with Stacey Robinson of GP Tech Advisors

In today’s digital landscape, growing companies face mounting pressure to demonstrate cybersecurity maturity. Whether it’s to win deals, attract investment, or pass audits, the need for robust security leadership is…

THE LATEST

AI and Fair Use

What the Latest Court Decisions Mean for Developers and Copyright Holders By Allysa Combs Artificial intelligence is reshaping industries at…

Prize Money Available, but at What Cost?

What Pre-collegiate Athletes and Their Families Need to Know After Brantmeier v. NCAA By Deonta Woods The NCAA has long…

NIL Contracts, Exclusivity, and Tampering

By Deonta Woods What Every College Athlete Needs to Know Before Signing Anything You finally did it. Your first Name,…

Contact Us

Let's challenge the default together