Expert Explains | Why ‘lawful access’ may not be required for AI training in India

In recent months, news reports have covered European booksellers receiving bulk orders for second-hand books, presumably to train large language models (LLMs).
In June last year, US district courts in two different cases had broadly held that training LLMs on lawfully obtained books, even if protected by copyright, is fair use. Nearly a year later, the Delhi High Court, in a significant interim ruling on July 24, held that using copyrighted works to train LLMs can fall within India’s fair dealing exception and does not, by itself, amount to copyright infringement.
The order came in a suit by news agency Asian News International, which alleged copyright infringement by ChatGPT-maker OpenAI for training its LLMs on ANI’s copyrighted works.
Though subject to appeal, the ruling gives AI developers such as OpenAI a measure of legal cover to train on Indian-origin content without licensing agreements.
Arul George Scaria, law professor at National Law School of India University, Bengaluru, and an expert in intellectual property and competition law, assisted the Delhi High Court as an amicus curiae in the case. In an interview with , he analysed the Delhi High Court verdict in the broader global context of how courts have interpreted copyright protections surrounding LLM training. Edited excerpts:
How does the Delhi High Court’s interim verdict bode for AI development and LLM training in the Indian jurisdiction?
In its interim order, the Delhi High Court has held that it is permissible for a machine to learn by training on copyrighted works, provided the final output is not an instance of memorisation and regurgitation of the original work. In doing so, it has taken the position that “data is the oil for LLMs to work efficiently”.
We can now say that we are following the US courts’ approach to training large language models (LLMs). Overall, this is a very balanced and forward-looking judgement. It is good for innovation because people can now engage in training with more confidence. But it isn’t a blanket check because there could still be liability on AI companies for copyright infringement on the output side of LLMs, depending on the specific facts of the case.
Story continues below this ad
Imagine the reverse scenario: if the judgment ruled that using copyrighted materials to train LLMs was infringement, it would shut down all indigenous generative AI development.
How does the Delhi High Court’s fair dealing exception compare with the fair use exception cited in two key US district court verdicts on LLM training and copyright infringement—involving Anthropic and Meta—last year?
One thing we must always remember is that different jurisdictions have very different approaches to protecting exclusive rights and crafting exceptions to those exclusive rights.
In the US context, courts have the “fair use” exception, a broad, open-ended exception based on a set of four factors to determine “fair use”: the purpose and character of use, nature of the copyrighted work, amount and substantiality of the portion taken, and the effect of the use upon the potential market. In the two US cases, both found that using copyrighted works for LLM training constitutes “fair use”, though there are nuances.
Story continues below this ad
The framework in India is slightly different. We have the “fair dealing” exception, in which the analysis happens in two stages. First, the court looks at whether the use was for one of the specific purposes mentioned in the fair dealing exceptions. Second, the court engages in a “fairness analysis.”
The Delhi High Court has rejected the blind adoption of the US’ four-factor test for fair use because our fair dealing provision is much narrower. It would be terribly wrong to import those factors directly into the Indian analysis.
When ANI vs OpenAI was underway, many people doubted whether India could follow a fair use approach given our narrower “fair dealing” exception. Here, the ultimate purpose is learning, which is now covered under the exception of “private use, including research.” The Delhi High Court has taken a liberal and dynamic approach to defining these terms, because unless the first stage is cleared, we can’t move on to the next.
Equally important, the Delhi High Court has laid down a unique fairness test. It asks us to look at three factors, including whether there will be market harm for the plaintiff — in this case, ANI, a syndicating agency—and whether ANI’s customers will substitute their services with LLM responses—where the answer is clearly no.
Story continues below this ad
Most importantly, the Delhi High Court noted that LLMs have a public interest dimension. Overall, the court concluded that training activity can fall within the ambit of the fair dealing exception.
How has the Delhi High Court interpreted copyright infringement on the output side—the responses generated after training? What rights do copyright-owners exercise here?
ANI argued that ChatGPT users could reproduce their content verbatim, which would be copyright infringement. But the court took a very pragmatic approach. It showed that all the training happened before the publication of the articles which they had cited as examples of infringement. Most examples cited were from 2024, whereas ChatGPT’s training happened before 2022.
The core issue was whether the outputs were “substantially similar” to the copyrighted content. News materials receive relatively lesser copyright protection because you cannot own facts, and there are only so many ways a fact can be presented. The net result is that the amount of copyrightable subject matter is much less than what the plaintiff is trying to project.
Story continues below this ad
Additionally, most outputs in ANI’s case were not substantially similar because the machine makes some changes to the inputs while generating the output. This means it didn’t meet the threshold for copyright infringement. The court said there can’t be liability on the output side here.
This is the outcome of this case, but this doesn’t mean that tomorrow anyone can reproduce copyright-infringing outputs. In a different case with different facts, there might be a finding of infringement on the output side. Delhi High Court has taken a pragmatic position while still looking at the output side for potential future liability for copyright infringement.
If paragraphs are being reproduced verbatim, the publisher/copyright owner can sue for output liability. As we move into specialised domains, LLMs will inevitably reproduce underlying content, making licensing by AI companies inevitable. This is why AI companies are entering licensing agreements with big publishers—to prevent future trouble.
OpenAI says it does not train on paywalled content or material blocked by web crawlers. Since ANI’s works were not behind a paywall, does bypassing such barriers to train LLMs amount to copyright infringement?
Story continues below this ad
Definitely yes. If an LLM circumvents technological protection measures, then there might be liability coming your way. In this case, the court noted that ANI didn’t use the “opt-out” options provided by OpenAI, like “robots.txt.”
The Delhi High Court has a pending suit regarding “shadow libraries” (like Sci-Hub) despite a preliminary observation that such libraries infringe on copyright. Since LLMs are often trained on these libraries, does this present a legal grey area in India?
In Elsevier vs Alexandra Elbakyan, I feel the court did not have the opportunity to hear the perspectives of the academic and research community. However, ANI vs OpenAI clarifies that “lawful access” is not a general requirement under our law, and under most laws. If you can prove as a user that your use comes under the fair dealing exception, the source shouldn’t matter. Wherever “lawful access” is intended, it is specifically mentioned in the statute.
Most jurisdictions don’t demand “lawful access” across the board because it would end fair use and its purpose. The best example is Google LLC vs. Oracle America, Inc: Oracle claimed that Google had copied copyrighted application programming interface code from the Java programming language to train its Android operating system. The US Supreme Court, in 2021, ruled in a 6–2 majority that Google’s use of the Java APIs was within fair use, and that Google copied everything without permission.




Leave a Reply