Anthropic is now watermarking everything Claude writes. But how do you hide a watermark in plain text?
In this video, I break down how LLM watermarking actually works, starting with the red/green token lists of Kirchenbauer et al. and ending with the tournament sampling behind Google DeepMind's SynthID-Text, which is the method Anthropic says it will use for Claude. I also build all three schemes in a playground so you can see exactly what they do to the output.
β Support the channel: Patreon: NoHypeAI
π₯ CHAPTERS
ββββββββββββββββββ
00:00 - Anthropic Is Watermarking Claude
01:15 - The Hard Red List
03:12 - Seeds, Hashes and the Detector
05:07 - The Problem With Entropy
05:52 - Demo: Hard Red List
07:42 - The Soft Red List
09:58 - Demo: Soft Red List
11:35 - SynthID-Text and Tournament Sampling
15:27 - Demo: Tournament Sampling
16:18 - What This Means in Practice
π USEFUL LINKS
ββββββββββββββββββ
π Kirchenbauer et al., "A Watermark for Large Language Models" - https://arxiv.org/abs/2301.10226
π Dathathri et al., "Scalable watermarking for identifying large language model outputs" (Nature) - https://www.nature.com/articles/s4158...
π SynthID-Text reference implementation - https://github.com/google-deepmind/sy...
π Anthropic's watermarking announcement page - https://www.anthropic.com/news/claude...
#AI #LLM #Watermarking #Claude #MachineLearning
Anthropic is now watermarking everything Claude writes. But how do you hide a watermark in plain text?
In this video, I break down how LLM watermarking actually works, starting with the red/green token lists of Kirchenbauer et al. and ending with the tournament sampling behind Google DeepMind's SynthID-Text, which is the method Anthropic says it will use for Claude. I also build all three schemes in a playground so you can see exactly what they do to the output.
β Support the channel: Patreon: NoHypeAI
π₯ CHAPTERS
ββββββββββββββββββ
00:00 - Anthropic Is Watermarking Claude
01:15 - The Hard Red List
03:12 - Seeds, Hashes and the Detector
05:07 - The Problem With Entropy
05:52 - Demo: Hard Red List
07:42 - The Soft Red List
09:58 - Demo: Soft Red List
11:35 - SynthID-Text and Tournament Sampling
15:27 - Demo: Tournament Sampling
16:18 - What This Means in Practice
π USEFUL LINKS
ββββββββββββββββββ
π Kirchenbauer et al., "A Watermark for Large Language Models" - https://arxiv.org/abs/2301.10226
π Dathathri et al., "Scalable watermarking for identifying large language model outputs" (Nature) - https://www.nature.com/articles/s4158...
π SynthID-Text reference implementation - https://github.com/google-deepmind/sy...
π Anthropic's watermarking announcement page - https://www.anthropic.com/news/claude...
#AI #LLM #Watermarking #Claude #MachineLearning
*FAQ*:
1. Q: Doesn't this obviously make the output worse?
A: That's the hard red list, which the Kirchenbauer paper itself presents as the naive version, and yes, it destroys quality. The soft red list is better but still shifts probabilities. However, in tournament sampling every candidate is an honest sample from the model's own distribution. So the watermark only decides which of those samples survives. Averaged over the random seed, the probability of outputting any token is exactly the model's original probability (the paper calls this "single-token non-distortionary"). Google's 20-million response A/B test found no measurable preference when it comes to quality. Nothing here bans a token or makes a wrong answer more likely than it already was.
2. Q: What about code?
A: Anthropic talks about this in their post and they admit it will behard to apply watermark in many cases. That makes sense because code will be mostly low-entropy tokens. In large enough codebases, the watermark can live in comments within code. I just tested generating a Fibonacci function in the watermarking playground and the tournament-sampled output was identical to the unwatermarked one.
3: Q: Doesn't the detector need the model, the weights, or the original prompt?
A: No. The detector never computes probabilities. The red/green assignment for a token depends only on the previous H tokens and the secret key, so all you need is the tokenizer, the key, and the watermark parameters. The prompt isn't needed either, you simply skip the first H tokens of the text.
4: Q: Can I just remove it by paraphrasing/translating/running it through another LLM?
A: Absolutely. Rewriting a large fraction of tokens with a non-watermarking model (or manually) will lower the z-score. Both papers test this and the watermark degrades a lot under heavy paraphrasing.
5: Q: Isn't this hugely expensive?
A: No, because we are only changing the sampling step, and hard/soft watermarking is basically instantaneous and tournament sampling can be implemented very efficiently. Google measured about 0.5% extra latency on Gemma 7B.
6: Q: Why are they doing this at all?
A: Primarily because of the new EU AI Act regulations that require providers to mark AI-generated output. I can also imagine this being useful as a measure against other people training on their own outputs.