Anthropic Says Claude’s Planned Watermark May Survive Light Edits—but Not a Complete Rewrite
A planned API will check a pattern encoded through subtle word choices, but constrained code, lightly edited human drafts and complete rewrites expose the boundaries of what the signal can establish.
Story brief
3 key pointsAnthropic will embed a hidden, key-detectable watermark in Claude outputs using the SynthID-Text approach, aiming to meet the EU AI Act’s Transparency Code. The watermark is encoded through low-stakes word choices and should be imperceptible to readers; Anthropic says light edits will often leave enough signal to detect while a complete rewrite will remove it. Anthropic plans a detection API but hasn’t released...
- 01
Watermark method: SynthID-Text-style signal encoded via ordinary word choices, not visible badges.
- 02
Edit resilience: Anthropic says light edits may preserve detection; a full word-by-word rewrite removes the watermark.
- 03
Detection API: Anthropic plans an API but has not disclosed release date, access model, or output format.
Under Anthropic’s plan, Claude-generated prose would carry a hidden, deliberately embedded pattern assembled from ordinary wording decisions. That would shift one form of AI detection away from interpreting stylistic clues and toward checking for a signal with the appropriate key—but the check would remain bounded by how much of Claude’s chosen language survives editing.
Anthropic announced the watermark and subsequently published an explanation of how it is intended to work. The company says the move is meant to comply with the EU AI Act’s Transparency Code, which the cited reporting describes as requiring systems that make AI-generated content identifiable. The proposal is narrower than a universal detector: it concerns an intentionally encoded signal in Claude’s output, and Anthropic acknowledges circumstances in which little or no signal may remain.
Inside the word choices
The mechanism operates where Claude has latitude. Anthropic’s example is a low-stakes choice between words such as “overcast” and “grey” when describing the weather. By making such choices in a patterned way across generated text, Claude can produce a signal intended to be imperceptible to a reader but detectable to someone with a key that encodes it. There is no visible badge or altered formatting in that description; the selected wording carries the pattern.
Anthropic says a watermarked response should be indistinguishable to readers from an unwatermarked one and that the process does not affect Claude’s output quality. Those are company assertions. The available reports do not describe independent testing of readability, quality or detection performance, so they do not establish how reliably those claims will hold across different kinds and lengths of text.
For the underlying method, Anthropic says it will use the SynthID-Text approach outlined by Google DeepMind in 2024. It also plans to release a planned detection API for checking the watermark. The reporting does not provide a release date, access model or response format for that API, leaving the practical detection workflow less defined than the encoding mechanism.
The distinction matters because the two approaches ask different questions. A style-based system examines the finished language for signs associated with AI writing. Anthropic’s proposed system instead checks whether a pattern that Claude was instructed to embed is still present. The watermark therefore does not rest on a phrase, punctuation habit or stylistic stereotype that a reader can simply notice; it depends on the accumulated choices encoded in the text.
The edit boundary
Anthropic’s resilience claim is carefully qualified. The company says light editing probably will not remove the watermark completely, while a complete rewrite in which every word is replaced will remove it. That is not a promise that the signal will survive editing generally. It is a narrower assertion that some amount of revision may leave enough of the original pattern intact to be detected.
The direction of the edit also matters. When Claude edits or proofreads human-authored material, Anthropic says detectability will depend on the document’s length and the intensity of Claude’s intervention. If the model changes only a small amount, nearly all the words remain the human author’s, leaving very little—or possibly nothing—for the watermark to attach to.
Where Anthropic says the signal can attach
- Claude-generated prose: Low-stakes wording alternatives give the model room to choose words in a pattern that can be checked with the relevant key. That offers more room for watermarking than constrained code or a human draft receiving only minor edits.
- A lightly edited Claude draft: Anthropic says the watermark will probably not be removed completely, although the company does not characterize every possible degree or type of revision.
- A human draft lightly edited by Claude: If most words remain the author’s, Anthropic says there may be very little or nothing available to carry Claude’s watermark.
- Functional code: The requirement to produce working code limits arbitrary choices, so Anthropic expects the watermark’s effect on the code itself to be negligible. Comments and other places with arbitrary wording could still carry the pattern.
- Text rewritten word by word: Anthropic says a complete rewrite in which every word is replaced will remove the watermark.
The escape route inside the limit
A full rewrite creates both a technical and definitional boundary. Anthropic argues that once every word has been replaced, it becomes debatable whether the resulting text should still be described as AI-generated. But that framing does not make the earlier use of AI reconstructable: it means only that the stated watermark is not designed to persist through complete replacement of its carrier text.
Futurism’s report relayed Ars Technica’s observation that passing watermarked text through another chatbot for rewriting could destroy the signal. It also relayed a warning that disclosing how detection works could make watermark-removal tools easier to build. These are reported concerns, not evidence in the supplied material that Claude’s system has already been defeated.
The same report compared the issue with attempts to strip Google DeepMind’s SynthID provenance marks from images, while noting that nobody had yet figured out how to remove those image watermarks fully. That history shows why removal is a live concern, but it does not establish the robustness of SynthID-Text or Claude’s planned implementation: the cited example concerns image provenance rather than the text system Anthropic is preparing to use.
Compliance meets user resistance
Anthropic presents regulation as the driver. It says the watermark is being implemented for compliance with the EU AI Act’s Transparency Code. Futurism describes the underlying AI Act as having passed in 2024 and as requiring AI companies to mark content generated or edited by their systems. Anthropic also says other major model developers signed the same Code of Practice and will implement their own watermarks, although the cited report does not identify those developers or describe their planned systems.
The announcement nevertheless triggered a visible backlash in some online communities. Futurism documented heated posts and jokes on r/ClaudeAI and r/artificial. Separately, TechCrunch reported that Business Insider had found “dozens” of people on X claiming to have canceled Claude subscriptions because of the change. Those posts establish that some users objected; they are not verified cancellation figures or representative evidence of Claude’s broader customer base.
Sources
- techcrunch.comAnthropic shares more details about how Claude’s new watermarks will work | TechCrunch
- futurism.comPeople Horrified That They'll Be Busted Now That Anthropic Is Watermarking AI Content
