24 August 2026

AI Watermarks Won't Solve AI Detection

I've been following the new AI watermarking developments quite closely.

Anthropic recently announced invisible watermarks for text generated by supported Claude models. Google, OpenAI, Meta and Microsoft have signed the same commitment: the Code of Practice on Transparency of AI-generated Content, backed by around 190 signatories in July 2026 and aligned with Article 50 of the EU AI Act, which became applicable on 2 August 2026.

Sounds useful. But there's a catch.

An infographic summarising what an AI watermark can and cannot show: it is invisible and global, it can survive when AI only edited your text, simple paraphrasing removes it, and it shows which model was involved rather than what the model did.

A watermark shows involvement, not authorship

A watermark can show that an AI model was involved. It can't reliably show what the model actually did.

AI can edit someone's original writing, and the watermark can remain. On the other hand, simple paraphrasing can remove it.

A 2026 evaluation found that paraphrasing removed the watermark from 98–100% of detected texts.

This matters in education. Using AI to fix a few sentences is not the same as generating an entire essay. A watermark doesn't tell you the difference.

The marks get missed before anyone tries to remove them

The same evaluation is worth reading for a second number, which I find more striking than the paraphrasing one.

On short passages, and before any attempt to strip the mark, the schemes missed watermarked text most of the time: false-negative rates of 70% for KGW, 83% for Unigram and 80% for a research implementation of SynthID.

That is ordinary generated text, submitted as it came out. No trickery involved. If a mark is missed that often when nobody is hiding anything, it cannot carry an academic misconduct case on its own.

Nobody outside the AI companies can check a document yet

This is the part that surprised me most, and it rarely comes up in the discussion.

There is currently no tool that lets a university, a supervisor or a student check a document for a watermark. Anthropic says third-party detection is still being developed, with details to follow in its technical documentation.

So even where a mark is present and intact, no one outside the model provider can read it today. A policy written around watermark detection right now is a policy with nothing to run it on.

It only points forwards, and only one way

Watermarking applies to models released from 2 August 2026 onwards. Everything written before that date carries nothing to find.

And Anthropic is explicit that the absence of a mark does not mean the content wasn't AI-generated or AI-processed.

Put that together with the point above and the asymmetry is the whole story. A watermark can suggest a model was involved. Its absence proves nothing at all. That is a one-directional signal, and most questions in academic integrity need an answer that works in both directions.

What we do instead

That's why we're careful about how we use AI detection at PaperCheck.

Our report calls the result a "probabilistic signal, not a verdict". It's something to investigate, not a decision to make. Our AI writing check runs as part of the plagiarism check and is built to point you at passages worth a second look, not to hand anyone a judgement.

If you want something that holds up when your work is questioned, the strongest evidence isn't in the finished document at all. It's the process: version history, earlier drafts, your notes, and a clear record of what you asked an AI tool and what you did with the answer. That evidence gets stronger the longer you work, and no watermark, present or absent, can take it away from you.


References

Natthida Baramee (JJ)

Natthida Baramee (JJ)

/ AI & Product Engineer
BEng in Electrical Engineering