AI Watermarking: Some Consequences

There’s interesting news from Anthropic regarding AI watermarking of text. From what I’ve been able to infer from their website* it’s a statistical measurement for text. For images like svg, png, jpg, there is metadata attached, which makes sense. All of this is in response to EU’s Code of Practice on Transparency of AI-generated Content**.

For video, image, and audio, watermarking is generally understandable and we’ve probably all come across it in some form or another. You can embed some structure in the lower bits of the output that marks something as uniquely yours or literally graft a visual watermark onto the content as two common practices [and I’m grossly oversimplifying].

Text is a bit harder. We all have our patterns and our habits. And LLM text is forced, by design, to have a statistical structure.

A single sentence is probably not enough for AI-watermarking to detect anything. But an article of this length? Probably enough. But with a wrinkle. Without saying it explicitly, there are Type I (false positive) and Type II (false negative) errors with any statistical output [in this case the truthiness of the existence / non-existence of the watermark]. It’s a constant challenge in classification problems, which buckets into four cases.

True positive: Text is detected to have a Claude watermark and it was indeed Claude-generated text.
True negative: Text is not detected to have a Claude watermark and it was indeed not Claude-generated text.
False positive: Watermark exists, but text was not Claude-generated.
False negative: Watermark does not exist, but text was Claude-generated.

Depending on how we use this watermarking, harm is always in the FP and FN rates. Expel a student from school because a watermark was detected? Harsh, if FP is very high. Do we know what these rates are?

There’s also going to be some fun statistical effects. You might be able to take text paragraph by paragraph and get no watermark. But the text as a whole has a watermark. Or the other way around!

For LinkedIn, I wonder now how this ties in with “Seems Like AI Slop”. Are they running a backend test to see how good human slop detection is against watermarking? Maybe the two in concert are a stronger detector?

Will platforms start watermarking everyone who posts content? Can a person get their own watermark? Can a person challenge a watermark?

Will there be watermark drift?

Will there be watermark laundering? (Probably. Anthropic already foreshadows this in their “Limitations” section.)

There’s also a social power dynamic at play. “Watermark” is a euphemism for “plagiarism”. I leave it to the reader to think about the irony.

What irks me is that it is centered around a “is it / is it not Claude” distinction [or any other LLM implementing their own watermark]. Because the methodology in use, could be extended to individuals and people to encapsulate their own writing style with their own watermark. A bit more complicated and I’m not spelling out all the headaches involved in that, but the current centering of “did Claude [AI] touch it or not” is a subtle marketing bias. For example, if the statement were “did a human touch this or not” it would feel totally different and it is a statement about what our null here is.

References
* Watermarking (1): https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
* Watermarking (2): https://www.anthropic.com/news/claude-text-watermark
** Code of Transparency: https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content#1720699867912-0