Toolspublished

Anthropic Will Watermark Claude Text by Steering Its Word Choices

Anthropic says readers will not see a difference, but its approach makes word selection part of an AI-detection system—and raises a harder question about what counts as a low-stakes change in generated prose.

By 3 min read
Anthropic Will Watermark Claude Text by Steering Its Word Choices

Story brief

3 key points

Anthropic plans to add provenance signals to future Claude responses by steering otherwise plausible word choices through a secret-keyed process. Detectors would look for the resulting statistical pattern rather than a label or embedded character. Anthropic says internal tests found no impact on content, creativity, or readability, but critics question whether supposedly interchangeable words can alter tone or...

  1. 01

    Claude’s watermark would be statistical: preceding words and a secret key influence choices among plausible continuations.

  2. 02

    Anthropic says readers will not notice a difference, but watermarked and unwatermarked responses will not use identical wording.

  3. 03

    Google’s SynthID study analyzed roughly 20 million Gemini responses and found a 0.01% thumbs-up-rate difference.

Anthropic plans to mark future Claude output without appending a label, inserting hidden characters, or changing the text in a way readers can spot. Instead, it will steer some of the model’s choices among plausible next words, leaving a detectable pattern across the response. The method turns ordinary word selection into part of the system for identifying AI-generated text.

A watermark inside the model’s tie-breakers

The design starts from a common language-model condition: several possible next words can fit a sentence. Anthropic says watermarking changes the source of randomness used to choose among those options. Rather than relying on an arbitrary random-number generator, the system uses a secret key and words that came before to influence the selection.

Anthropic illustrates the mechanism with a sentence beginning, “The weather today was cold and….” It says words such as “overcast” and “grey” may both be plausible continuations, allowing the watermarking process to shape the random choice. The company says those low-stakes decisions accumulate over a response into a pattern readers cannot see.

The result is not a single marker embedded in a document. Anthropic says a detector can inspect a sequence of words and assess the probability that it is consistent with selections made using that key. Watermarked and unwatermarked versions can therefore contain different eligible wording, even as Anthropic says readers will not be able to distinguish them.

Nothing is added to the text and there are no hidden characters.

Anthropic, as quoted by 404 Media

Detection and writing quality are different tests

Anthropic says its internal testing found no effect on Claude’s content, creativity, or readability, and that watermarked responses are indistinguishable from unwatermarked ones to readers. That is a company performance claim, not a claim that the two versions contain identical language: the mechanism deliberately changes which eligible words appear.

What Anthropic is claiming

  • The watermark adds no visible marker or hidden characters to Claude’s text.
  • A keyed process uses preceding words to make a pattern across otherwise plausible word choices.
  • Anthropic says that pattern lets a detector assign a probability that text came from Claude.
  • Anthropic has described the change as part of its response to new European Union AI regulations.

The distinction matters because “indistinguishable” is a practical threshold for a detection tool, while writers may use neighboring words for differences in tone, rhythm, implication, or precision. 404 Media argued that Anthropic’s description treats choices it calls low-stakes as more interchangeable than they can be in actual writing. The criticism is about the premise of the system, not whether a hidden character was inserted.

A method borrowed from Google’s watermarking research

Anthropic’s approach is based on research by Google researchers behind SynthID, a watermarking tool. In the SynthID study cited in the discussion, researchers analyzed about 20 million watermarked and unwatermarked Gemini responses and found a 0.01% difference in thumbs-up rates. That result supports a narrow usability signal from the evaluated feedback measure; it does not settle the more subjective dispute over how altered wording affects style or meaning in a particular passage.

The study’s large-scale measure was product feedback: it routed a random fraction of queries to a watermarked model and compared Gemini users’ thumbs-up and thumbs-down responses with an unwatermarked counterpart. It also used side-by-side human assessments of watermarked and unwatermarked text for quality. Those tests address whether users notice a general quality difference, not why a particular word was selected in a particular sentence.

Anthropic’s proposal therefore puts two evaluations alongside one another: whether its keyed process can identify Claude-like word sequences, and whether altered selections change readers’ experience of the prose. The company says its internal testing found no impact on content, creativity, or readability. Its critics dispute whether reader-level indistinguishability is a sufficient standard for writing.

Sources

  1. 404media.coAnthropic’s Text Watermarking Proves AI Companies Do Not Care at All About Writing