Ingestion or Infringement? How Legacy Copyright Law Is Unraveling in the Age of Artificial Intelligence

Share
Ingestion or Infringement? How Legacy Copyright Law Is Unraveling in the Age of Artificial Intelligence

Executive Overview

The global rise of generative artificial intelligence has relied on an unprecedented accumulation of human intellect. Large language models (LLMs)—including the systems powering platforms such as ChatGPT, Gemini, and Claude—have been trained on vast repositories of published literature. Hundreds of millions of copyrighted books, digital articles, peer-reviewed academic papers, and creative writing datasets have been ingested to construct the neural networks driving contemporary AI.

For the creative community, this technical practice presents a profound existential crisis. Authors, journalists, and researchers discover that their life’s work has been absorbed without explicit knowledge, opt-out mechanisms, or financial compensation into tools that could ultimately dilute their professional value. To the lay observer, extracting proprietary intellectual property to build multi-billion-dollar commercial software appears to be a clear breach of copyright law.

However, judicial interpretations reveal a far more complex legal reality. United States intellectual property frameworks do not explicitly prohibit software systems from "learning" or "reading" protected material. Instead, federal judges are forced to apply a decades-old statute—the Copyright Act of 1976—to software architectures that were unimaginable when the legislation was drafted.

As judicial opinions begin to trickle through federal dockets, a nuanced precedent is taking shape: while unauthorized data scraping from illegal piracy repositories remains subject to severe penalties, the core process of machine learning itself is increasingly framed by courts as a permissible form of transformative study. This dynamic has sparked a fierce legal battle among artificial intelligence developers, published creators, legal scholars, and corporate interests over fair use, market competition, and the boundary of human authorship.


Detailed Chronology of Landmark Precedents

   [1976] US Copyright Act Enacted (Baseline Statutory Framework)
     │
     ▼
   [Early AI Era] Unsanctioned Mass Ingestion of Web Data & Shadow Libraries
     │
     ▼
   [Thomson Reuters v. Ross Intelligence]
   Courts test market substitution; training on proprietary legal data to build a competing tool ruled NOT fair use.
     │
     ▼
   [Thaler v. Perlmutter]
   D.C. Circuit confirms 100% AI-generated output cannot receive copyright protection.
     │
     ▼
   [Anthropic Litigation / Judge Alsup Settlement]
   Judge orders $1.5B penalty for pirated shadow library acquisition, BUT validates LLM training as lawful "reading."

Unsanctioned Ingestion and the Shadow Library Crisis

In the foundational stages of LLM development, tech firms swept up expansive tranches of text from across the internet. To achieve human-level language fluency, these systems required trillions of tokens—units of text drawn from novels, technical documentation, news archives, and online forums. In their pursuit of data density, several organizations turned not only to public websites but also to illicit online "shadow libraries"—unauthorized repositories containing millions of pirated books.

This methodology established the groundwork for a wave of class-action copyright lawsuits brought by novelists, journalists, and publisher advocacy groups seeking to challenge the structural mechanics of generative AI training.

Anthropic Rulings and the "Reading vs. Copying" Dichotomy

In a landmark judicial action, U.S. District Judge William Alsup presided over a high-profile copyright suit filed by a collective of authors against AI developer Anthropic. The case yielded a headline-grabbing outcome: Judge Alsup ordered Anthropic to pay a massive $1.5 billion settlement to the aggrieved writers.

However, a closer look at the judicial reasoning reveals a critical distinction. Judge Alsup did not penalize Anthropic for the act of training its LLMs on copyrighted literature. Rather, the $1.5 billion sanction was applied specifically because Anthropic had obtained those literary works through illegal online shadow libraries—a clear act of digital piracy.

On the underlying question of whether training an AI on text constitutes copyright infringement, Judge Alsup ruled in favor of the developer. Drawling an explicit parallel between human learning and machine ingestion, the judge observed:

"Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different."

This ruling created a major legal benchmark: while the unlawful acquisition of source material remains punishable, the computational process of analyzing and learning from that text falls closer to non-infringing consumption than illegal duplication.

Commercial Substitution: Thomson Reuters v. Ross Intelligence

The judicial consensus shifts dramatically when an AI system is trained explicitly to replicate or supplant the primary market of the target data. This dynamic was tested in Thomson Reuters v. Ross Intelligence, where legal publisher Thomson Reuters sued AI platform Ross Intelligence for scraping its proprietary legal database (Westlaw) to build a competing, AI-driven research platform.

Presiding over the case, Judge Stephanos Bibas rejected the defendant’s fair use defense, emphasizing that Ross Intelligence ingested the structured legal summaries specifically to build a direct market replacement. Judge Bibas noted:

"Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s."

The decision established an important legal baseline: training an AI system on proprietary data to build a product that directly competes with the original rights-holder crosses the line from transformative learning into market substitution.

The Ownership Dilemma: Thaler v. Perlmutter

While cases against Anthropic and Ross Intelligence focus on the inputs of AI development, the judicial system has simultaneously evaluated the legal status of AI outputs.

In Thaler v. Perlmutter, federal courts evaluated whether a visual work generated entirely by an autonomous AI system without human involvement could receive federal copyright protection. The court delivered an unambiguous answer: purely AI-generated output is ineligible for copyright protection, as human authorship remains an indispensable requirement under U.S. law.

This decision introduced a complex operational challenge for creative industries: if a work created entirely by an AI model cannot be copyrighted, where is the legal line drawn when human creators use AI tools for assistance?


Supporting Context & Economic Metrics

To assess the broader landscape of AI copyright litigation, it is essential to examine the underlying financial realities, statutory frameworks, and legal doctrines driving court decisions.

+-------------------------------------------------------------------------+
|                  AI COPYRIGHT PRECEDENT FRAMEWORK                       |
+------------------------------------+------------------------------------+
|            INPUT LEVEL             |            OUTPUT LEVEL            |
|       (Training & Ingestion)       |      (Synthetic Generation)        |
+------------------------------------+------------------------------------+
| Anthropic Precedent:               | Thaler v. Perlmutter Precedent:    |
| Ingesting text to "learn" patterns | 100% AI-generated content cannot   |
| is permissible transformative use; | receive statutory copyright        |
| piracy in data gathering remains   | protection; requires human         |
| illegal.                           | authorship.                        |
+------------------------------------+------------------------------------+
| Ross Intelligence Precedent:       | Spectrum Uncertainty:              |
| Ingestion to directly supplant     | How much human prompt engineering  |
| or compete with the original tool  | or editing is required to establish|
| forfeits fair use protections.     | copyrightability?                  |
+------------------------------------+------------------------------------+

Financial Realities of AI Settlements

While a $1.5 billion copyright settlement would devastate a conventional software enterprise, the hyper-capitalized nature of the modern AI sector recalibrates those economics.

For a company like Anthropic—which is projecting annual revenues to reach approximately $200 billion by 2028—a $1.5 billion regulatory or legal penalty represents a manageable cost of doing business. Because Judge Alsup’s ruling essentially validated the legality of the underlying LLM training process, the financial penalty was a favorable trade-off for establishing a key legal precedent for the industry.

AI Revenue Forecast vs. Judicial Penalties (USD Billions)
│
├── Anthropic Projected Annual Revenue (2028):  $200.0 B
│   ██████████████████████████████████████████████████████████████████
│
└── Anthropic Piracy Settlement Fine (Alsup Ruling): $1.5 B
    █▌

The 1976 Copyright Act’s Structural Obsolescence

The primary legislative engine governing United States intellectual property is the Copyright Act of 1976. Drafted nearly five decades ago, the statute was enacted during an era dominated by analog publishing, physical print runs, broadcast television, and magnetic tape.

Because Congress could not anticipate neural network training, statistical text synthesis, or algorithmic deep learning, federal judges are tasked with adapting 20th-century statutory text to novel high-tech practices. This statutory disconnect forces courts to rely heavily on the judicially created doctrine of Fair Use.

The Four Pillars of Fair Use Analysis

Under Section 107 of the Copyright Act of 1976, judges evaluate fair use defenses using a four-factor balancing test:

  1. The Purpose and Character of the Use: Is the work being used for commercial benefit, and critically, is the new use transformative—adding new meaning, message, or utility?
  2. The Nature of the Copyrighted Work: Is the source material creative fiction or factual text?
  3. The Amount and Substantiality Used: How much of the original work was taken, and was the core essence ("the heart") extracted?
  4. The Effect Upon the Potential Market: Does the new application serve as a market substitute that hurts the economic value of the original work?

Expert Insights & Official Judicial Statements

Legal analysts highlight that the clash over generative AI centers on a fundamental question: does algorithmic processing constitute "copying" material or "consuming" it?

Intellectual Property Perspective: "Consuming" vs. "Replicating"

Cathy Gellis, a prominent attorney specializing in intellectual property, copyright, and technology law, emphasizes the vast gap between how the public views AI processing and how copyright law actually operates.

"I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on," Gellis noted. "It’s very complex and there are a lot of raw feelings about what is happening, both for and against."

Commenting on the structural implications of Judge Alsup’s ruling against Anthropic, Gellis explained that the decision fundamentally favors AI development by isolating the legal act of reading from illegal duplication:

"I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work," Gellis stated. "Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work."

Gellis also points out the grey area that emerges when applying Thaler v. Perlmutter to human-AI collaboration. If tools like spellcheck do not strip a human author of their copyright, at what point does an generative AI tool cross the line and undermine ownership?

"If you write your novel in Microsoft Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel," Gellis said. "[AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while."

Market Competition and Judicial Reasoning

Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, points out that judicial outcomes hinge primarily on market impact rather than technical scraping mechanics.

"Everybody is very worried right now because the law is all over the place, and it’s because of this question," Henderson observed. "They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question."

Henderson notes that courts are increasingly using direct commercial competition as the primary line of demarcation in fair use defenses:

"Copyright is always about protecting and growing the market," Henderson stated. "The courts are kind of all over the place in their reasoning [in AI cases]. What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay."


Future Outlook: A Fractured Precedent Landscape

The legal landscape governing artificial intelligence remains in flux. While early rulings offer a baseline framework for AI developers, the lack of federal legislative updates means legal standards will continue to be established case-by-case across different judicial circuits.

AI developers face a complex operational reality:

  1. Scraping Protocols & Data Sourcing: Developers are aggressively distancing themselves from shadow libraries and illicit data repositories. To mitigate liability, firms are shifting toward licensed content ecosystems, public-domain datasets, and clear opt-out web-crawling standards.
  2. The Threat of Circuit Splits: Different federal district and appellate courts may arrive at contradictory conclusions regarding fair use and transformative ingestion. A ruling from a conservative circuit that rejects the "reading" analogy could trigger a conflict that ultimately reaches the U.S. Supreme Court.
  3. The Ownership Dilemma in Enterprise Output: Corporations utilizing AI systems to produce code, marketing copy, graphics, and legal documentation must deal with uncertain IP protections. If output generated primarily by AI cannot be copyrighted, enterprise users risk creating assets that competitors can copy with impunity.

As Cathy Gellis points out, these early judicial opinions are establishing the ground rules for a multi-trillion-dollar industry, even as higher courts prepare to review them down the line:

"What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later stages of litigation to figure out which one will prevail," Gellis concluded. "But in the meantime, all these decisions are shaping everything that’s happening. It would be kind of foolish for the AI companies to ignore them."

Until Congress enacts comprehensive, technology-focused copyright reform, the AI ecosystem will operate under a patchwork of judicial rulings. As a result, tech companies and content creators alike must navigate an uncertain legal environment where the line between learning and infringement remains constantly in flux.

Did you find this story helpful?

Share it with your friends and colleagues on social media.

Share

Leave a Comment

Your email address will not be published. Required fields are marked *