How developers are trying to remove Anthropic’s AI text watermarks

AI companies are looking to implement watermarks for machine-generated content. But developers are also building tools to remove them. As regulators push for greater transparency, proving what was created by AI tools is fast becoming a cat-and-mouse game.
Earlier this month, Anthropic announced its decision to embed invisible, machine-readable watermarks in anything generated by future Claude models. Within days, several developers took to social media to share watermark removal tools.
One such override by Guillaume Meyer has gone viral on GitHub, drawing more than 100 contributors to improve the tool and leading to several more downloads, according to a report by Wired.
Story continues below this ad
Similar tools that claim to remove imperceptible watermarks in AI-generated content, especially text, have popped up online since Anthropic’s announcement. These attempts to evade watermarking come amid concerns that the blanket watermarking of AI-generated content increases the risk of false positives.
Anthropic itself has acknowledged that its watermarks only indicate that Claude was ‘likely’ involved in processing the content at some point. It cannot distinguish ‘Claude wrote this’ from ‘Claude heavily edited this’. It also does not confirm if the text was human-written or whether it was generated by an AI model from a different company.
If watermarking and detection practices are adopted widely, it could lead to employers unfairly rejecting candidates or overblown accusations of researchers and students using AI tools for their work just because the detector flags it.
“We’re adding marking to Claude’s output to comply with the EU AI Act, and other labs are taking similar steps. It’s hard to identify AI-generated text, and this gives people better tools for identification. Text from supported Claude models, including output from Claude Code, will carry an invisible watermark, and it doesn’t change the meaning, quality, or readability of Claude’s responses. We also plan to ship a text-detection API so users can do more of this themselves,” Anthropic was quoted as saying by Wired.
Story continues below this ad
To understand the workarounds that have emerged, let’s first take a look at Anthropic’s watermarking method.
How Anthropic will watermark Claude outputs
Anthropic’s watermarking efforts are mainly aimed at complying with transparency obligations under Article 50(2) of the European Union’s AI Act, which went into effect on August 2, 2026.
The rules require AI model providers to label synthetic audio, image, video, or text so that this material can be detected by a machine as AI-generated. They also state that model providers cannot market circumvention tools. Non-compliance could lead to fines of up to 3 per cent of annual turnover.
Anthropic also joined over 190 organisations, including OpenAI, Microsoft, and Meta, to voluntarily sign the EU’s transparency code of practice.
Story continues below this ad
Anthropic’s watermarking method is based on an approach (SynthID-Text) by Google DeepMind. The search giant has been using it to watermark its AI-generated content since 2023. It broadly involves watermarking text invisibly by leaving a pattern in Claude’s choice of words and phrases that is indiscernible to a human reader but would be detectable by a machine that knows how to look for it.
Some users have raised concerns that the watermarking method could degrade the quality of Claude’s responses, though Anthtropic insists that this will not be the case.
Early workarounds: Are they effective?
The tool designed by Guillaume Meyer involves using a non-watermarking AI model to generate multiple rewrites, swapping in synonyms, and slightly re-organising content. But this workaround is likely to work only when using other AI models that do not insert watermarks.
There is also no certainty that this tool can remove the watermarks until Anthropic releases the software it uses to detect a watermark. Meanwhile, other developers have also built watermarking removal tools. One such tool is said to be capable of removing invisible and look-alike characters, reorder sentences within paragraphs, and swap several words for synonyms.
Story continues below this ad
The watermarks can also be removed by condensing Claude’s response, translating it into a dialect like Arabic, and then translating it back, according to Leon Chlon, a visiting fellow at the University of Oxford.




Leave a Reply