India’s first judgment on artificial intelligence and copyright raises a deeper question than whether training an AI model on someone’s work infringes on their copyright: how would a copyright owner ever prove the infringement?
In Asian News International v OpenAI, ANI alleged that ChatGPT was reproducing its copyrighted news articles. Much to the news agency’s dismay, the Delhi High Court was not persuaded by its claims and handed a double win to OpenAI. First, it held that ChatGPT was not reproducing enough of ANI’s actual writing to count as copying its work.
Second, and more strikingly, the court held that OpenAI did not require ANI’s permission to train on its articles, since doing so falls within activities exempt from India’s copyright law. An inevitable question arises from this preliminary finding from the court: how can copyright holders ever prove, to a court’s satisfaction, that their work was used to train an AI model?
Hard to establish similarity
One of the ways Indian courts assess copyright infringement is through the principle of “substantial similarity”. Copying facts or ideas alone is not enough; it is the style of expressing an idea that determines the case.
ANI argued that ChatGPT replicated its news articles word-for-word and in the same style. In order to establish verbatim reproduction, the burden of proof was on ANI to show that ChatGPT had memorised its works during training. Copyright jurisprudence also dictates that when two works are being compared for infringement, they should be compared in their entirety and not just the selected parts.
Applying these principles, the court looked at a specific example. When ANI asked ChatGPT what Olympian Neeraj Chopra’s mother had said in an interview, ChatGPT reproduced only one line from ANI’s article and added its own commentary. As ChatGPT had provided the answer with its own distinct flavour, ANI could not convince the court of any substantial similarity with its works. ANI had also failed to show that ChatGPT was reproducing entire articles—making their case weaker.
OpenAI’s models were trained before some of the works ANI presented as evidence were even published. The court explained that ChatGPT can pull current information from live website links when it responds to a query, which is distinct from what the model learned during training. What we are left with is the court’s understanding that a chatbot’s answer can resemble a news article without the model ever having “learned” it. This take is fundamentally at odds with how copyright law has evolved around static works such as books or movies, which are not dynamic like AI-generated answers tend to be.
Also read: Why Prashant Kishor’s Bankipur victory in Bihar should ring alarm bells for Modi-Shah’s BJP
Direct prompting may weaken publishers’ case
ANI’s attempts to get ChatGPT to reproduce its work only weakened its case. The court noted that ANI manipulated ChatGPT by asking it to quote “exactly” what Neeraj Chopra’s mother said. It likened ANI’s technique to ‘adversarial prompting’, which refers to manipulating the prompts provided to a large language model (LLM) such as ChatGPT to make it produce harmful or undesirable outputs. The court contrasted this with a recent German case, where ChatGPT reproduced substantial song lyrics verbatim just from simple prompts.
If a copyright holder uses carefully designed prompts to test whether an LLM has used their works, should a court reject that simply because of how the prompts were framed? Does this mean that rights owners now also have to be responsible for prompting the model only in a particular permissible way? Considering the fact that it is difficult for copyright holders to access training data or chat logs of LLMs, it may become near-impossible for them to discharge their burden of proof and show substantial similarity with their works.
OpenAI submits in the ANI case that ChatGPT “is not designed to reproduce” extracts of its training content. If that’s true, shouldn’t users be unable to bypass such safeguards? And doesn’t the fact that ANI could extract the content through deliberate effort suggest it remains extractable to anyone willing to try (the question of live links aside)? This push and pull between rigorously testing a model and being accused of manufacturing results may gain even more significance as the ANI trial progresses.
Also read: Behind every Supreme Court petition is a waiting game
Looking forward
The ANI case demonstrates the evidentiary shortcomings likely to arise in AI-related copyright infringement cases.
However, it also demonstrates that innovation cannot wait until the end of lengthy trials. In this scenario, policymakers and the judiciary must carefully analyse how licensing can provide AI companies not only with good-quality training data but, more importantly, certainty against lawsuits. ANI itself had proposed that OpenAI license their content, and the fact that hundreds of such deals have been signed globally is evidence that the approach works in favour of all. Government-backed efforts such as India’s AIKosh platform also help by giving AI developers a pool of datasets contributed on agreed terms—but this may work only for data that’s been opted in. It does not solve for or tell a court what to do when a rightsholder, such as ANI, has not opted in.
Before Indian courts settle how far copyright’s exceptions stretch to cover AI training, they may need to settle a more basic question: how can a copyright holder ever prove their work was used by a model? On the evidence produced so far, that may be the harder problem to solve.
Soumya S is an associate at Koan Advisory Group, a Delhi-based consultancy firm. Views are personal.
This article is part of ThePrint-Koan Advisory series that analyses emerging policies, laws and regulations in India’s technology sector. Read all the articles here.
(Edited by Prasanna Bachchhav)

