AI & Technology

Anthropic explains Claude's watermark: proofread copy carries little of it

Written by
Full Name
August 17, 2026
Anthropic named the method behind Claude's watermark on 14 August and set out where it fails. The mark attaches only to words Claude itself chose, so a light edit of human writing leaves almost nothing to detect, while a Claude translation carries a full signal.

Twelve days after Claude began marking its own output, Anthropic published the mechanism. The technical post, dated 14 August, identifies the watermark as a version of SynthID-Text and answers the question the support page left open: not whether Claude marks its output, but how much of a given document the mark can reach.

That distinction changes the picture for content teams, and it qualifies what the first fortnight of coverage reported. The Helm reported on 14 August that Claude now watermarks text worldwide, including copy it only proofread, drawing on Anthropic’s support page. The technical post narrows that. Marking is not triggered by Claude’s involvement as such. It rides on the words Claude selects, which means the volume of the model’s contribution decides whether a mark registers at all.

How does Claude’s watermark work?

Anthropic describes a version of SynthID-Text, the method Google DeepMind published in Nature in 2024 and a descendant of a 2022 proposal by Scott Aaronson. Language models pick each word from a shortlist of plausible candidates, and where two options are equally good the choice is settled at random. Watermarking swaps that randomness for a pattern generated from a private key, so a later check can estimate the likelihood that Claude produced the passage.

Nothing is inserted. Anthropic states that no characters are added to the text, no extra tokens are produced, and the model is neither slower nor more expensive to run. The company also says the mark carries no identifying information: nothing in the watermark or its key would let anyone recover a user, an organisation or a chat. Ownership and legal responsibility for the output are unchanged.

On quality, Anthropic reports no effect in internal testing. It points to the DeepMind paper behind the method, which served a watermarked model to a share of Gemini traffic and found no statistically significant difference in thumbs-up and thumbs-down ratings. A controlled study in the same paper had raters compare watermarked and unwatermarked answers side by side, and they saw no difference.

Which parts of a content workflow leave a mark?

Anthropic’s account makes the answer a matter of degree rather than a yes or no. The watermark applies only to words Claude chooses, so the more of a document the model wrote, the stronger the signal and the more confident any later check can be. Short passages give too little to work with.

Proofreading is the case that has caused most alarm, and it is the weakest. When Claude edits a person’s writing, Anthropic says, nearly all the words remain the person’s, so there is very little for the watermark to attach to, and the changes may not be enough to make Claude’s involvement detectable. A grammar-and-punctuation pass leaves the mark living in a handful of corrections. Heavier rewriting is a different matter, because more of the finished text becomes Claude’s own choices.

Translation sits at the opposite end. A translation carries a full mark, because Claude selects every word. Factual passages fall between: where the correct next word is fixed, there is no free choice for the watermark to use, so a dense run of names, figures and dates marks more thinly than discursive prose. Code marks least of all. Where an exact output is required, the watermark is not applied, surviving mainly in comments, which Anthropic says leaves a negligible effect on the code itself.

The labelling duty that reaches marketing teams runs on a separate test and is unaffected by any of this. As the Helm reported on 7 August, the Commission’s guidance confines the Article 50(4) duty to material published to inform the public on matters of public interest. Text that has had genuine human review, with a named person taking editorial responsibility, is exempt.

Can anyone read the mark yet?

Anthropic has not released a detector. The company says a watermark detection API is coming and that it is still working out the implementation, which leaves the mark unreadable by anyone outside Anthropic for the moment.

That gap has not stopped a market forming. Dozens of removal tools appeared on GitHub within days of the rollout, and in a review published on 12 August, Pasquale Pillitteri read the code of each project rather than its documentation and found the claims rarely survived contact with it. Most such tools strip invisible Unicode characters. Anthropic’s post rules that out as the mechanism, since nothing is added to the text and no hidden characters are inserted. On removal the company is direct: light editing probably will not clear the mark, while a complete rewrite replacing every word will, at which point the text is arguably no longer AI-generated.

What a detected mark proves is narrower than either the tools or the anxiety suggest. It indicates Claude was likely involved at some point and cannot separate writing from heavy editing. It says nothing about a different model, which would carry a different key or no watermark at all. Anthropic also separates its watermark from detection software such as Pangram, which has no access to the key and instead reads stylistic tells. The company notes that models favour the construction “this isn’t X, it’s Y” and show an unusual fondness for the word quietly.

Claude models placed on the market before 2 August 2026 must carry marking from 2 December 2026 under the Act’s transition period. Anthropic has not given a date for the detection API.

Subscribe to our newsletter

By subscribing you agree to with our Privacy Policy
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Share article

Recommended Reading