AI Briefing
KO
Pick

How Claude's 'Watermark' Distorts Writing

·2026.08.17 18:33

Key point

This analysis examines how Anthropic's text watermarking for Claude works and its potential to distort quality.

Details

The text watermark Anthropic plans to introduce to Claude models does not use invisible character insertion but instead leverages steganography.

Core Mechanism: Probabilistic Token Selection

  • When generating the next token, the model classifies words in real time into 'green' and 'red' lists.
  • To embed the watermark, the model intentionally increases the probability of selecting words from the green list over the red list.
  • As these subtle probabilistic biases accumulate across the text, a detector with the secret key can statistically identify them.

Concerns Over Quality Degradation and Distortion

  • Anthropic claims this method does not alter the meaning, quality, or readability of the text.
  • However, the author argues that the probabilistic word selection process interferes with the model's natural language choices, inevitably distorting the meaning and quality of the text.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.