Further to my earlier piece on AI transparency The Great AI Labelling Illusion; When Compliance Becomes More Important Than Trust, AI watermarking sounds like the answer to a defining problem of generative AI, how do we know whether something was created by a machine? These new transparency regulations come as part of the European Unions (EU) Artificial Intelligence Act, which came into effect on 2nd August’26. The Act is intended to reduce deception and manipulation and help people make informed choices within the wider aim of fostering trustworthy AI in Europe.
How will this be achieved? Let’s start with Anthropic’s Claude watermarking, which is based on Google’s SynthID. It subtly influences word choices during generation to create a statistically detectable pattern. Clever but it isn’t a digital fingerprint and as I wrote earlier, we know this pattern. Tick the box. Pass the audit. Celebrate compliance. Meanwhile, nobody has asked whether the watermarking actually delivers.
Watermarking can provide evidence of AI involvement BUT it cannot provide durable proof of AI authorship and that creates a potentially governance problem. If organisations start treating watermark detection as a binary test ie: watermark = AI; no watermark = human, they will be placing considerably more evidential weight on the technology than it can support. (Watermark under Fire)
The weakness is transformation. Light editing may preserve the signal but substantial paraphrasing, translation or regeneration can degrade it. Anthropic itself acknowledges that heavily edited, paraphrased or translated content may lose the watermark.
More importantly, a market already exists for what might reasonably be called ‘AI laundering’ services, these are tools designed to take AI generated text and rewrite it while preserving meaning but reducing its detectability. Examples include:
- Undetected.ai claims to rewrite AI text to clear major detectors.
- BypassAI markets undetectable rewriting and a ‘Stealth Writer’.
- AI Undetect explicitly supports rewriting ChatGPT, Claude and other AI output to bypass detectors.
- BypassAI.ai markets itself as an AI detection remover.
- GPTHuman.ai explicitly markets humanisation against multiple AI detectors.
- Unwatermarker, strip Anthropic text watermarks by sanitising characters, reordering sentences, swapping synonyms and flagging stock phrasing.
The Claude specific arms race has already started. WIRED reported that within 4 hours of Anthropic confirming its watermarking plans, developer Guillaume Meyer had published an open-source approach intended to remove the watermark by having an unwatermarked LLM generate multiple rewrites. Its effectiveness against Anthropic’s eventual production detector remains uncertain but the direction is unmistakable.
This is Goodhart’s Law in action, once a measure becomes a target, behaviour adapts around it. Detector improves. Laundering model improves. Repeat.
Microsoft 365 Copilot illustrates another approach. Microsoft now supports AI watermarking for images, video and audio and adds provenance metadata to supported media. Its current documentation, however, does not describe an equivalent statistical watermark embedded in ordinary Copilot-generated text from Word, Outlook or Copilot Chat.
That distinction matters. Microsoft is effectively placing greater emphasis on provenance surrounding the artefact, including metadata aligned with C2PA, rather than attempting to make every generated sentence intrinsically detectable.
Then there is OpenAI which takes a similarly different approach. ChatGPT generated images gained SynthID watermarking in May 2026, complementing C2PA Content Credentials and OpenAI extended SynthID to supported audio in July 2026. Yet OpenAI’s current public verification service supports images and audio, not ordinary ChatGPT text.
All these approaches however, expose the same fundamental limitation that provenance becomes harder to establish once content leaves the controlled environment and is transformed. Imagine three apparently identical paragraphs:
- Claude = potentially watermarked
- ChatGPT = no documented text watermark
- Microsoft 365 Copilot = no documented text watermark
Now ask ChatGPT or Copilot to substantially rewrite the Claude paragraph. What exactly does a subsequent failure to detect Claude’s watermark prove? Potentially very little.
That makes the proposition, no watermark = human, dangerously flawed. Different AI platforms apply different provenance technologies to different media and AI generated material can subsequently pass through multiple models.
The deeper issue is therefore not watermarking technology but the question it’s trying to answer. As AI becomes embedded in research, drafting, analysis, editing and everyday productivity, Was this written by AI? becomes increasingly meaningless.
A human may develop the argument, Claude draft it, ChatGPT restructure it, Copilot polish it and a human approve the final version. Trying to identify a single author from the statistical characteristics of the finished text misses how AI assisted work is actually evolving. I suggest perhaps the better questions are:
- What role did AI play?
- What is the provenance of the work?
- Who remains accountable for the result?
Watermarking remains a useful signal but treating it as forensic proof risks turning an imperfect and inconsistently deployed technical measure into a dangerously convenient judgement. You can watermark the words. You cannot watermark the ideas and increasingly, you may not even know which machine or human wrote the final words.
There is also a useful irony here; we are potentially using AI to identify AI, while simultaneously using AI to make AI harder for AI to identify.
Posted on August 21, 2026
0