New Delhi: The software that schools, journals and newspapers use to catch artificial intelligence (AI)-written text can be fooled with a simple trick, a study by researchers at Duke University, North Carolina in the US, and the Indian Institute of Science (IISc), Bengaluru has found. Instead of writing an answer itself, an AI system can borrow most of its words from a second AI model that detectors do not recognise.
The paper, by Bhuwan Dhingra of Duke and Danish Pruthi of IISc was posted online on 25 September, and is yet to be peer-reviewed. It arrives weeks after Europe began requiring AI companies to mark machine-made content, and months after a major AI conference rejected research papers based on a detector’s verdict.
The trick rests on a gap researchers at Carnegie Mellon University spotted in May. Chatbots, such as Claude or ChatGPT, begin life as ‘base’ models, trained only to predict the next word by reading huge amounts of human writing. They are later trained further to follow instructions and hold conversations, and it is this second stage that leaves the telltale patterns detectors look for.
Text from the raw base model, the Carnegie Mellon team found, mostly passes as human. Pangram, the company behind one of the most widely used detectors, has acknowledged that such text falls outside what its tool is built to catch.
Dhingra and Pruthi gave Anthropic’s Claude Opus 5 access to one such raw model, released by the Allen Institute for AI, a Seattle-based non-profit research body founded by Microsoft co-founder Paul Allen in 2014, and told it to use that model’s words instead of its own.
This is how it works. Since a raw model cannot take instructions, Claude has to coax it. It writes the opening of a fake document, say a Reddit post, office meeting notes or a newspaper’s health advice column, and lets the raw model carry on writing. It does this dozens of times, picks the sentences and phrases that fit, and stitches them into an answer. Claude adds only the few words needed to join the pieces, so up to 90 per cent of the final text is not its own.
The researchers tested the method on 400 tasks: short stories, questions about a given document, medical queries, and writing tasks with strict rules, such as using only lowercase letters. Pangram flagged 77 percent of the text when Claude wrote on its own. With the stitching method, that dropped to 24 percent. For short stories, Claude’s own writing was caught every time, while the stitched versions were scored as almost entirely human. Three other detectors caught less than 10 percent of it.
The answers held up well. The stitched responses followed formatting rules almost as well as Claude’s own, and stayed accurate to the documents they drew on. Stories were slightly less coherent, and medical answers suffered the most, losing about a tenth of their score. Medical answers were also the hardest to disguise, as Claude was told to put accuracy ahead of borrowing words. Older ‘humaniser’ tools, which repeatedly reword AI text, fared worse on both counts and often mangled the meaning. Simply asking Claude to sound human did not work, and neither did letting it test its drafts against Pangram.
The stakes are rising because detectors are now deciding outcomes for people. In June, NeurIPS, one of the world’s largest AI research conferences, rejected 178 papers, nearly one in five submissions in one section, after Pangram flagged them, with no chance to appeal. A study using the same tool estimated that about 9 percent of articles in American newspapers were partly or fully AI-written.
The paper also looks at watermarks, invisible signatures that AI companies hide in the pattern of words their models produce. Under the European Union’s (EU) AI Act, companies have had to mark AI-generated content since 2 August, and Anthropic announced that month it would watermark Claude’s writing worldwide. But a watermark can only sit in words Claude actually writes.
In a simulation, the researchers estimated that stitched text would be caught less than two percent of the time in most tasks, and about 40 percent of the time for medical answers. They caution these are estimates, and that Anthropic’s system works differently from the one they modelled. Anthropic has itself said that rewriting every word would remove its watermark.
India’s own rules, which require platforms to label AI-generated images, audio and video from February this year, do not cover text. Pruthi has questioned detectors before. A paper he co-wrote this year found they cannot reliably tell apart text that AI merely polished from text it wrote outright, and a 2024 IISc committee he served on advised caution in using such tools against students.
The method is not cheap. Because Claude reads through so many samples, each answer cost up to 30 times more than normal at Anthropic’s listed prices, about $1 to $3 each, and the second model needed powerful computer chips to run. The authors note that people on monthly subscriptions would simply use up their limits faster.
They suggest detector companies train their tools on text from raw models, and say enforcing labelling laws may need more than marks hidden in the text. The paper also discloses that Claude ran most of the experiments.
(Edited by Nardeep Singh Dahiya)
Also Read: Scholars are using AI for research work, but worry about ‘AI shame’, says new study
