AI Models Are Watermarking Text—Will You Notice?

Anthropic announced on 11 August that all future Claude models will produce text carrying a watermark identifying it as AI generated. Google already applies its own text watermark, on which Anthropic's is based, to output from its Gemini models, and OpenAI plans to introduce one. The spread of watermarking is in part a response to the European Union's AI Act, which mandates watermarks for AI models released after 2 August 2026. The act also requires watermarks for images, audio and video; image and video watermarks are already deployed by OpenAI, Google and Meta, and can achieve detection rates above 99 percent, though their effectiveness as a holistic solution remains debated.
Text watermarking is not metadata or invisible characters. It works by subtly altering how a model selects words from its probability distribution. A widely cited 2023 paper by John Kirchenbauer, a postdoctoral fellow at the Vector Institute, and colleagues sorts words into a red list and a green list, nudging green-list words to be slightly more probable. The authors reported a 98.4 percent detection rate and zero false positives in responses of about 200 tokens, and said removing the watermark from a long response requires changing roughly a quarter of its words or more.
Critics dispute whether the technique degrades output. Technology writer and Markdown co-creator John Gruber calls the watermark a "perversion of writing" and disputes Anthropic's assertion that it does not change the meaning or quality of text, noting text responses span far fewer units than images' millions of pixels. Kirchenbauer argues a watermark would not be detectable without a change, and asks whether users care if the distribution differs when utility is unchanged.
Google's 2024 SynthID-Text paper reported no significant difference in user feedback across 20 million responses when queries were randomly routed to watermarked and non-watermarked model variants. Anthropic's watermark is based on SynthID-Text but altered in ways Anthropic has not detailed; the company declined to provide additional information. Vinu Sankar Sadasivan, an AI research scientist at Meta, says watermarks struggle when word choices are few, citing a 20-word tweet needing 50 to 60 percent green-list words for reliable detection, and disagrees with claims of no quality change. Google's paper showed detection up to 95 percent in best cases but below 50 percent for short replies.
Kirchenbauer also points to other uses, citing a 2026 paper he co-authored showing that a model trained on watermarked text will itself produce watermarked output, which could support tracing data provenance or excluding prior-generation content from training data.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
On 11 August, Anthropic announced that all future Claude models will generate text that contains a watermark that identifies its results as AI generated. The company is not alone. Google has its own text watermark (which Anthropic’s is based on) that it uses on the output of its Gemini models . OpenAI has yet to introduce a text watermark but it plans to do so . The rapid spread of watermarking is in part a response to the European Union’s AI Act , which mandates watermarks for AI models released after 2 August, 20