in ,

Is it legal to train AI models on copyrighted books? The answer is complicated

Is it legal to train AI models on copyrighted books?

AI models such as ChatGPT, Gemini, Claude, and other chatbots are trained on huge amounts of data. This data can include books, news articles, research papers, websites, and other published material.

Many of these works were created by human authors. In many cases, those authors did not know their work could be used to develop AI systems. They also did not give direct permission for that use.

Hosting 75% off

That has created a major question. Can AI companies legally train their models on copyrighted material?

At first, the answer may seem obvious. If someone uses a copyrighted book without permission, it may seem like copyright infringement. But AI training has created a much more complicated legal debate.

The Law Is Still Developing

Cathy Gellis is an attorney who focuses on intellectual property, copyright, and technology. She says this debate involves many different issues.

The technology is moving quickly. Copyright law is not. That gap has made the issue difficult for courts, authors, and AI companies.

The debate also brings strong opinions from both sides. Authors are worried about their work and income. AI companies argue that training models is not the same as simply copying and republishing books.

Courts are now being asked to decide where that line should be drawn.

Anthropic Case Raises Important Questions

One of the biggest cases so far involved Anthropic. The company was sued by authors over the use of their books in AI training.

In 2025, Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to a group of writers. The case initially looked like a major victory for authors.

But the ruling was more complicated than that.

Judge Alsup found that Anthropic’s AI training itself was lawful. The problem was how the company obtained some of the books used for that training.

Anthropic had obtained books from illegal online sources. Those sources are often described as shadow libraries. The judge treated that conduct differently from the actual process of training an AI model.

This distinction is important.

The ruling suggests that using copyrighted works to train an AI model may not automatically violate copyright law. Obtaining those works through illegal methods can create a separate legal problem.

Is AI Training Similar to Reading?

Judge Alsup also compared AI training with the way people learn from books.

A writer can read thousands of books. They can study different writing styles and ideas. They can then create something new.

The judge viewed AI training in a somewhat similar way. An AI model processes enormous amounts of text. It then uses what it has learned to generate new material.

That does not mean every form of AI training is automatically legal.

However, it does show why the issue is difficult. Copyright law generally focuses on copying protected work. It does not necessarily prevent someone from reading, studying, or learning from that work.

Gellis believes this part of the ruling could benefit AI companies. The decision separates the act of training from the act of copying protected material.

Why Copyright Law Is Struggling With AI

Copyright law in the United States was not designed for modern AI systems.

Many of today’s copyright rules were created decades before generative AI existed. Courts are now applying those rules to technology that can process billions or even trillions of pieces of information.

That creates a difficult situation.

AI companies are building models using massive datasets. Authors are asking whether their books should be part of those datasets without permission.

Courts must decide how existing copyright rules apply.

Jason Henderson, a senior attorney and founder of the IP & Media Practice at JWL International, has also pointed to this problem. The technology has moved ahead much faster than the law.

That means different courts may reach different conclusions.

Fair Use Could Be the Key

Much of the AI copyright debate comes down to fair use.

Fair use allows certain uses of copyrighted material without permission. It can apply in areas such as criticism, commentary, education, research, and parody.

But fair use does not provide an automatic exemption.

Courts usually consider several factors. They may look at why the material was used. They may consider how much of the work was used. They may also examine whether the use affects the original work’s market.

The purpose of the new use can be especially important.

If a company uses copyrighted material to create something that directly competes with the original, a court may view that use more critically.

If the new use has a different purpose, the argument for fair use may be stronger.

Competition Can Change the Picture

This issue became clearer in a case involving Thomson Reuters and Ross Intelligence.

Thomson Reuters accused Ross Intelligence of using its copyrighted legal content to build an AI-powered legal research platform.

Ross was developing a product that would compete with Thomson Reuters.

Judge Stephanos Bibas ruled that Ross’s use of the material was not transformative enough to qualify as fair use.

The court looked closely at the purpose of the use. Ross was using the material to build a competing legal research product.

That made the case different from some arguments surrounding general-purpose AI models.

The ruling does not mean that every AI training system using copyrighted material is illegal. It does show that commercial competition can play an important role in a fair use analysis.

Could Chatbots Compete With Authors?

This creates another difficult question for writers.

AI chatbots can generate stories, articles, scripts, and other forms of written content. Those systems may have learned from books written by human authors.

Some authors could argue that AI-generated writing competes directly with their work.

That argument could become more important as AI-generated content improves.

However, courts have not established a simple rule saying that AI training becomes illegal whenever the resulting system can compete with human creators.

The legal debate is still developing.

AI Training and AI-Generated Content Are Different

There is another important distinction.

Copyright questions surrounding AI training are not the same as questions about copyright protection for AI-generated work.

Training asks whether copyrighted material can be used to develop an AI model.

AI-generated content asks whether the resulting work can receive copyright protection.

These are separate legal issues.

A major case involving AI-generated content is Thaler v. Perlmutter.

The court ruled that a work created entirely by AI cannot receive copyright protection under current U.S. law.

That decision raises new questions.

What happens when a person creates something with AI assistance?

What if a person writes most of a book but uses AI to rewrite certain sections?

What if AI creates an image but a human makes significant edits?

There is no simple answer for every situation.

Human Involvement Could Matter

The role of the human creator may become increasingly important.

Consider a simple example. Someone writes a novel using Microsoft Word. They use spell-check to fix mistakes. Nobody would normally argue that Microsoft owns the novel.

AI tools make the situation more complicated.

An author may use AI to brainstorm ideas. They may ask it to rewrite a paragraph. They may use it to edit grammar. They may also generate entire sections with AI.

At some point, the question becomes harder.

How much human creativity is needed for copyright protection?

And how can courts determine how much of a work was created by a person?

These questions are still being worked out.

AI Companies Still Face Legal Challenges

Most major AI companies remain involved in copyright disputes.

The lawsuits cover different issues. Some focus on training data. Others focus on how copyrighted material was obtained. Some involve AI-generated content.

Because the cases are still moving through the courts, there is no single rule that answers every question.

One court may reach a different conclusion from another. An early decision can also be changed later in the appeals process.

That makes the legal landscape difficult to predict.

What Happens Next?

The current cases could shape the future of AI.

Courts are creating early precedents. Those decisions can influence future lawsuits and business practices.

AI companies will need to pay close attention to these rulings. Authors and publishers will also be watching closely.

The outcome could affect how future AI models are trained. It could also influence licensing deals between AI companies and content owners.

Some companies may choose to license copyrighted material directly. Others may continue arguing that their training methods fall under fair use.

New laws could also change the situation.

For now, there is no universal answer to the question of whether AI companies can legally train models on copyrighted books.

The answer depends on the facts. It can depend on how the material was obtained. It can depend on how it was used. It can also depend on whether the use is considered transformative.

The legal battle is far from over.

What happens in the next few years could determine how AI companies train their models and how creators protect their work.

Hosting 75% off

Written by Hajra Naz

Nevada Approves Up to 8,000 Robotaxis for Tesla, Waymo and Uber in Las Vegas

OpenAI Calls for Stronger AI Safety Rules in California

OpenAI Calls for Stronger AI Safety Rules in California