Anthropic plans to mark future Claude output without appending a label, inserting hidden characters, or changing the text in a way readers can spot. Instead, it will steer some of the model’s choices among plausible next words, leaving a detectable pattern across the response. The method turns ordinary word selection into part of the system for identifying AI-generated text.
A watermark inside the model’s tie-breakers
The design starts from a common language-model condition: several possible next words can fit a sentence. Anthropic says watermarking changes the source of randomness used to choose among those options. Rather than relying on an arbitrary random-number generator, the system uses a secret key and words that came before to influence the selection.
Anthropic illustrates the mechanism with a sentence beginning, “The weather today was cold and….” It says words such as “overcast” and “grey” may both be plausible continuations, allowing the watermarking process to shape the random choice. The company says those low-stakes decisions accumulate over a response into a pattern readers cannot see.
The result is not a single marker embedded in a document. Anthropic says a detector can inspect a sequence of words and assess the probability that it is consistent with selections made using that key. Watermarked and unwatermarked versions can therefore contain different eligible wording, even as Anthropic says readers will not be able to distinguish them.
Nothing is added to the text and there are no hidden characters.
Anthropic, as quoted by 404 Media
Detection and writing quality are different tests
Anthropic says its internal testing found no effect on Claude’s content, creativity, or readability, and that watermarked responses are indistinguishable from unwatermarked ones to readers. That is a company performance claim, not a claim that the two versions contain identical language: the mechanism deliberately changes which eligible words appear.
What Anthropic is claiming
- The watermark adds no visible marker or hidden characters to Claude’s text.
- A keyed process uses preceding words to make a pattern across otherwise plausible word choices.
- Anthropic says that pattern lets a detector assign a probability that text came from Claude.
- Anthropic has described the change as part of its response to new European Union AI regulations.
The distinction matters because “indistinguishable” is a practical threshold for a detection tool, while writers may use neighboring words for differences in tone, rhythm, implication, or precision. 404 Media argued that Anthropic’s description treats choices it calls low-stakes as more interchangeable than they can be in actual writing. The criticism is about the premise of the system, not whether a hidden character was inserted.
A method borrowed from Google’s watermarking research
Anthropic’s approach is based on research by Google researchers behind SynthID, a watermarking tool. In the SynthID study cited in the discussion, researchers analyzed about 20 million watermarked and unwatermarked Gemini responses and found a 0.01% difference in thumbs-up rates. That result supports a narrow usability signal from the evaluated feedback measure; it does not settle the more subjective dispute over how altered wording affects style or meaning in a particular passage.
The study’s large-scale measure was product feedback: it routed a random fraction of queries to a watermarked model and compared Gemini users’ thumbs-up and thumbs-down responses with an unwatermarked counterpart. It also used side-by-side human assessments of watermarked and unwatermarked text for quality. Those tests address whether users notice a general quality difference, not why a particular word was selected in a particular sentence.
Anthropic’s proposal therefore puts two evaluations alongside one another: whether its keyed process can identify Claude-like word sequences, and whether altered selections change readers’ experience of the prose. The company says its internal testing found no impact on content, creativity, or readability. Its critics dispute whether reader-level indistinguishability is a sufficient standard for writing.
Reader comments
Newest comments first. Replies stay oldest first.