24 August 2026
AI Watermarks Won't Solve AI Detection
I've been following the new AI watermarking developments quite closely.
Anthropic recently announced invisible watermarks for text generated by supported Claude models. Google, OpenAI, Meta and Microsoft have signed the same commitment: the Code of Practice on Transparency of AI-generated Content, backed by around 190 signatories in July 2026 and aligned with Article 50 of the EU AI Act, which became applicable on 2 August 2026.
Sounds useful. But there's a catch.
A watermark shows involvement, not authorship
A watermark can show that an AI model was involved. It can't reliably show what the model actually did.
AI can edit someone's original writing, and the watermark can remain. On the other hand, simple paraphrasing can remove it.
A 2026 evaluation found that paraphrasing removed the watermark from 98β100% of detected texts.
This matters in education. Using AI to fix a few sentences is not the same as generating an entire essay. A watermark doesn't tell you the difference.
The marks get missed before anyone tries to remove them
The same evaluation is worth reading for a second number, which I find more striking than the paraphrasing one.
On short passages, and before any attempt to strip the mark, the schemes missed watermarked text most of the time: false-negative rates of 70% for KGW, 83% for Unigram and 80% for a research implementation of SynthID.
That is ordinary generated text, submitted as it came out. No trickery involved. If a mark is missed that often when nobody is hiding anything, it cannot carry an academic misconduct case on its own.
Nobody outside the AI companies can check a document yet
This is the part that surprised me most, and it rarely comes up in the discussion.
There is currently no tool that lets a university, a supervisor or a student check a document for a watermark. Anthropic says third-party detection is still being developed, with details to follow in its technical documentation.
So even where a mark is present and intact, no one outside the model provider can read it today. A policy written around watermark detection right now is a policy with nothing to run it on.
It only points forwards, and only one way
Watermarking applies to models released from 2 August 2026 onwards. Everything written before that date carries nothing to find.
And Anthropic is explicit that the absence of a mark does not mean the content wasn't AI-generated or AI-processed.
Put that together with the point above and the asymmetry is the whole story. A watermark can suggest a model was involved. Its absence proves nothing at all. That is a one-directional signal, and most questions in academic integrity need an answer that works in both directions.
What we do instead
That's why we're careful about how we use AI detection at PaperCheck.
Our report calls the result a "probabilistic signal, not a verdict". It's something to investigate, not a decision to make. Our AI writing check runs as part of the plagiarism check and is built to point you at passages worth a second look, not to hand anyone a judgement.
If you want something that holds up when your work is questioned, the strongest evidence isn't in the finished document at all. It's the process: version history, earlier drafts, your notes, and a clear record of what you asked an AI tool and what you did with the answer. That evidence gets stronger the longer you work, and no watermark, present or absent, can take it away from you.
References
- How Claude marks AI-generated content β Anthropic help centre
- AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation β Tamim & Khan, arXiv preprint, July 2026
- Strong backing for the Code of Practice on Transparency of AI-generated Content β European Commission
- EU AI Act, Article 50 β Transparency Obligations
