New Delhi: OpenAI will start adding an invisible watermark to text generated by ChatGPT and Codex for users in the European Union (EU), the company announced on Monday. The step is aimed at complying with the EU Artificial Intelligence (AI) Act, which requires AI companies to make generated content identifiable by machines.
According to OpenAI, the watermark will be rolled out to eligible ChatGPT and Codex users across all plans in the EU “over the coming weeks”. Developers using OpenAI’s API anywhere in the world, including India, can turn on watermarking for select models from Monday, though the feature will remain off by default. Access to the tool that detects the watermark will, for now, be restricted to approved researchers and expert organisations.
ChatGPT users outside the EU will not see any change. “We are not making text watermarking a global default at launch,” the company said in a blog post, adding that the regional rollout would allow it to learn from real-world use and feedback.
What EU law says
Article 50 of the EU AI Act, which came into effect on 2 August, requires providers of generative AI systems to mark text, images, audio and video in a machine-readable format so that they can be detected as AI-generated. Violations of the Act’s transparency obligations can attract fines of up to €15 million or 3 per cent of a company’s global annual turnover, whichever is higher.
In June, the European Commission published a voluntary Code of Practice laying down how companies can meet these requirements. Signatories are presumed to be in compliance with the law. Anthropic has said that around 190 organisations have signed the code. The rules also require deepfakes and AI-generated text published on matters of public interest to be labelled, and users to be informed when they are interacting with a chatbot.
OpenAI said access to its detector would be granted on a case-by-case basis “in accordance with the Code of Practice”.
How the watermark works
OpenAI’s system, called textGrain, works by embedding a statistical pattern in the words a model chooses while writing a response. Language models generate text one word at a time, and at many points, more than one word fits the sentence equally well. Normally, the choice between them is random. A watermark replaces this randomness with a pattern linked to a secret key, which a detector can later look for.
The company said the watermark does not add any visible marks or special characters to the text. It also does not identify the user, organisation, account, prompt or conversation.
In its post, OpenAI framed the purpose narrowly, saying the watermark helps answer a single question: “Was this likely generated by an OpenAI model?”
Where it falls short
OpenAI also published data on the system’s limitations. At a target false positive rate of 1 per cent, its detector identified the watermark in about 80 per cent of 200-token passages and about 95 per cent of 400-token passages on psychology-related content. Detection rates were lower for mathematics, where the model has fewer options in word choice.
Editing the text weakened the signal. In 400-token passages, replacing 10 per cent of the words with synonyms brought detection down from about 92 per cent to 66 per cent. When 25 per cent of the words were replaced, it fell to 17 per cent.
“Rewriting or translating text can completely remove the watermark,” the company said.
OpenAI claimed textGrain “matched or exceeded” the other methods it tested, including Google’s SynthID for text, but cautioned that results under test conditions may not hold in everyday use. On output quality, the company said watermarking did not affect the performance of Astra, its latest frontier model. On the GPQA Diamond benchmark, Astra scored 94.44 per cent without the watermark and 93.94 per cent with it. On the Artificial Analysis Intelligence Index, the scores were 49.57 and 49.76 points respectively.
OpenAI also listed what a detection result cannot establish. It does not indicate how much a human edited the text, who owns it, who is responsible for it, or whether it is accurate. Equally, the absence of a watermark does not prove a human wrote the text, since it may be too short, edited, translated, produced by an older model or generated by another company’s tool.
Also Read: What is Griffin AI? World’s first ‘human interaction model’ passes Turing Test
Anthropic’s approach
Anthropic took a different route. In a post published on 14 August, the company said all Claude models released after 2 August embed a watermark in the text they generate, and that the change applies worldwide with no option for users to switch it off. This means Claude users in India are already receiving watermarked text from newer models.
Anthropic said it applied the watermark globally because it does “not yet have a durable way to scope it by region”, and that it would continue to evaluate other approaches.
Its method is based on SynthID-Text, a technique Google DeepMind described in a paper published in Nature in 2024. Anthropic said the watermark does not affect the quality of Claude’s responses, does not add hidden characters, does not use extra tokens and carries no information that could be traced to a user or a conversation.
The company acknowledged similar limitations. The watermark is weaker in code, factual passages and light proofreading, where most of the words belong to the user, and a complete rewrite removes it. Translations done by Claude, however, carry the watermark since every word is chosen by the model.
Anthropic has made a detection tool available in private preview to regulators, law enforcement agencies, media organisations, fact-checkers, researchers, educational institutions and EU civil society groups. Image files generated by Claude now carry a signed credential in their metadata under the C2PA industry standard. The company said older Claude models will get the watermark over the coming months.
Google DeepMind added SynthID-Text to Gemini in 2024 and released the code as open source for developers. According to Anthropic, Google tested the watermark on a section of Gemini users and found no difference in how they rated the responses.
A shift for OpenAI
The announcement marks a change in OpenAI’s position. In 2024, the company said it had developed a text watermarking method but was not releasing it, citing concerns that it could be bypassed through translation or rewording, and that it could unfairly affect non-native English speakers who use AI tools to improve their writing.
On Monday, OpenAI said its approach “must reflect the practical limits of current technology”. The company said it will extend watermarking to its models offered through cloud partners in the coming weeks, publish more technical details on textGrain and eventually release it as open source.
“We expect to revisit each part of this approach as the technology, standards, and evidence evolve,” OpenAI said.
(Edited by Nardeep Singh Dahiya)
Also Read: No training for AI models without risk assessment & senior leadership’s nod, says OpenAI
