How Claude Watermarks AI-Generated Text
Anthropic has announced that it will watermark the text outputs of its Claude models, according to a lecture published by Ahead of AI. The announcement was made a few days before August 14, the author says in the video.
The 48-minute video, which grew from an initial plan of 10 slides and a 10-minute recording to more than 50 slides, is accompanied by a transcript and a link to the slides. Ahead of AI also produced a YouTube version.
The stated motivation of the watermarking is to let Anthropic identify text generated by its models, for example Claude Opus 4.8. The watermark is described as invisible to users, with only Anthropic able to decode it and determine whether the text carries the mark. The author notes that Anthropic's own article linked to a technical paper and contained no figures, making the mechanism hard to understand.
The lecture begins with a prelude on how LLMs generate text. It covers tokenization, the score distribution over a vocabulary that a model produces for the next token, and the conversion of scores into probabilities. The example given is the prompt "the capital of Germany is," where the highest-scoring token, at vocabulary index 19,846, corresponds to "Berlin." The transcript says a real LLM would almost always sample "Berlin" for this prompt.
The video then describes next-token sampling, including greedy decoding and the random sampling used by most LLMs, as well as top-k and top-p sampling. A second example uses the prompt "today's weather is" with "cold," where "gray" and "overcast" are treated as similarly plausible next tokens.
The lecture also covers random number generation and deterministic sampling with a fixed seed. The author says the watermarking technique is a minor tweak inside the regular text generation process, and that the video explains how the technique works, how watermarking can fail or be removed, and the arguments for and against it. The author says coding an LLM from scratch helped clarify where watermarking is applied and its consequences.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
A 48-minute video walkthrough of token sampling, watermark detection, and removal