You Can Watermark the Words. But Can You Watermark the Meaning? — Claude and the AI Watermark Arms Race
Editorial image generated by the author.
In August 2026, Anthropic announced that supported new Claude models would begin embedding machine-readable watermarks in generated text. The mark is intended to preserve some trace of origin after the text leaves Claude: copy it into a document, paste it onto a website or edit it lightly, and the signal may still remain.
The announcement came as Article 50 of the European Union's AI Act began applying new transparency obligations on 2 August 2026. Providers of generative AI systems are required to make artificially generated or manipulated output machine-readable and detectable, using technical solutions that are effective, interoperable, robust and reliable as far as technically feasible. The accompanying Code of Practice on Transparency of AI-generated Content offers a voluntary route for demonstrating compliance. Anthropic presents its marking system as part of its response to those requirements.
Then the countermeasure appeared almost immediately. An open-source project called claude-watermark-cleaner set out to disturb Claude's watermark not by discovering and surgically removing the hidden signal, but by using another language model to rewrite the text while trying to preserve its meaning. Some coverage promptly described the episode as the watermark being neutralised within 24 hours.
That goes further than the evidence currently allows. The project's own documentation says that its success cannot be measured objectively until Anthropic releases its detector. The countermeasure is attacking an incompletely documented target: Anthropic has confirmed the watermark, but outsiders do not yet have enough information about the detection mechanism and thresholds to establish whether this particular rewriting strategy defeats it.
The uncertainty makes the episode more interesting, not less. The important fact is not that Claude's watermark has already been proven breakable. It is that a plausible countermeasure follows almost automatically from the structure of the problem. A text watermark marks choices made while expressing an idea; language can preserve the idea while replacing those choices. The arms race began before outsiders could even score the first round.
Why a Watermark Is Different from an AI Detector
Most AI detectors examine finished prose and ask whether its statistical characteristics resemble machine-generated writing. They infer origin from features that were never deliberately placed there for identification. A watermark begins from a stronger position: the provider deliberately introduces a detectable pattern during generation, allowing a later detector to search for that pattern rather than merely asking whether the prose “looks like AI.”
That is a meaningful evidential advantage. Ordinary AI detection can classify authentic human writing as artificial because its style happens to resemble statistical patterns associated with language models. The consequences of trying to infer AI authorship from the finished text can become serious long before the evidence becomes conclusive. Watermarking tries to replace stylistic suspicion with a signal deliberately introduced during generation.
But a watermark is still not a cryptographic signature of authorship. Anthropic's own description of Claude's marking system says that detecting a supported mark indicates that content may have been processed by Claude. Claude might only have proofread, translated or summarized material whose ideas and original text came from elsewhere. Conversely, heavily edited, paraphrased, translated or mixed text may no longer carry a detectable mark. Watermarking improves the evidence without settling who did the intellectual work.
Text Is a Hostile Medium for Durable Provenance
Consider two sentences:
The government announced the decision on Tuesday.
and:
On Tuesday, the government made its decision public.
For most human purposes, they say the same thing. At the level of language, however, much has changed: the words, their order, the token sequence and some of the grammatical structure. A text watermark must attach itself to the linguistic expression of an idea while surviving transformations that preserve the idea but replace much of that expression.
Images face a related problem, but they contain enormous quantities of redundant information. Millions of pixel values can be altered slightly without changing what a viewer perceives, and file formats can carry signed provenance metadata alongside the visible content. Text has less room to hide. Every word contributes to meaning, tone or rhythm. Push a model too aggressively toward watermark-friendly choices and the prose may deteriorate; make the signal too subtle and ordinary rewriting may erase it.
Language is also designed to be transformed. We summarize it, translate it, shorten it, expand it, reorganize it and alter its tone. Those operations once required meaningful human effort. Another language model can now perform them almost instantly. An attacker may therefore have no need to understand the watermark itself. If the signal depends on patterns in wording or token choice, it may be enough to preserve every factual claim and conclusion while reconstructing the sentences with different vocabulary and syntax, perhaps more than once.
This attack class predates Claude's announcement. The 2026 Vaporizer study tested lexical alteration, machine translation and neural paraphrasing against several LLM watermarking schemes and found that watermarks could be removed under its experimental conditions while much of the semantic content remained intact. That is not a proof that every watermark is fragile. It is evidence that preserving meaning while replacing expression is an adversarial problem any durable text watermark has to confront.
The same technology can therefore operate on both sides. One model produces the marked language; another attempts to launder it through transformation. A more robust watermark raises the cost of doing so, which creates an incentive for a stronger transformation. The pattern resembles spam filtering or malware detection more than a watermark printed onto paper: success is measured in costs, probabilities and adaptation rather than permanence.
The Watermarker's Triangle
A useful text watermark appears to need three properties that pull in different directions. Fidelity means that the text should not become noticeably worse merely because the system is trying to mark it. Persistence means that copying, ordinary editing and ideally more substantial transformation should not immediately erase the signal. Attribution validity means that the mark should not continue implying meaningful Claude authorship after Claude's contribution has become marginal.
The first two create a technical problem. The third creates a conceptual one.
Suppose I write an essay myself and ask Claude to correct the grammar without changing the argument. Claude returns a polished version carrying a watermark. The detector has not established that Claude wrote the essay. Or suppose I write something in Swedish and ask Claude to translate it into English. Claude generated the English phrasing, but the argument, evidence and intellectual work remain mine.
Real workflows can be more tangled. A human writes a draft; Claude restructures it; the human replaces several sections; another model copy-edits the result; a third translates it; an editor changes the wording again. If the original Claude mark disappears too easily, deliberate evasion becomes trivial. If it survives almost indefinitely, it risks continuing to describe text whose relationship to Claude has become increasingly remote.
The more durable the watermark becomes, the harder the attribution question becomes. Persistence is useful only while the thing being persisted still describes something we care about.
When Does Claude Stop Being the Author?
Much of the argument around AI writing still relies on a binary distinction: human-written or AI-written? That distinction was easier to defend when the typical workflow was either writing something yourself or entering a prompt and copying the resulting output. It becomes less useful as AI moves into ordinary intellectual work.
One writer uses a model to brainstorm; another asks it to identify weaknesses in an argument. Someone dictates rough thoughts and has a model organize them. Someone else writes every sentence but uses AI for translation. Another generates most of a first draft and then spends hours reconstructing it. All of those workflows contain AI, but treating them as equivalent tells us very little about authorship.
A powerful watermark may therefore answer a technically precise but socially ambiguous question: Did this text pass through Claude? That is different from Did Claude write this?, and different again from Did a human do the intellectual work behind it?
Schools, publishers, journalists and researchers rarely care about processor history for its own sake. They care about originality, responsibility, effort, deception and authorship. A watermark can provide evidence relevant to those questions. It cannot decide them.
An Arms Race Can Still Produce Useful Defences
A removal tool appearing almost immediately makes watermarking look futile only if an unbreakable mark is the standard. In adversarial systems, that is usually the wrong standard. Spam filters are circumvented constantly and remain useful; malware evolves to avoid detection while security tools continue to block large amounts of it. A defence can matter because it changes the cost and convenience of evasion, even when it cannot make evasion impossible.
AI watermarking may occupy the same territory. A student who simply copies an untouched Claude answer could be easier to identify. A spam operation publishing enormous volumes of unedited generated text could become easier to detect. Platforms might use watermarks as one signal among several when investigating synthetic content, while publishers could gain another piece of evidence when examining disputed submissions. None of those uses requires the watermark to survive an adversary prepared to reconstruct every sentence repeatedly.
The useful questions are therefore empirical: how much rewriting is required, how much does it cost, what happens to quality, how reliably does the transformed text evade detection, and how many users will bother? A mark that survives ordinary editing and catches most casual copying could be valuable even if a motivated user can eventually remove it. A mark erased by one routine paraphrasing pass would have a much narrower role.
We do not yet know where Claude's system falls on that spectrum. Describing it as “defeated within 24 hours” therefore mistakes the appearance of an adversarial strategy for a demonstrated technical result. The strategy itself is nevertheless revealing: if provenance lives in the way an idea is expressed, an evader will search for transformations that preserve the idea while replacing the expression.
Can an Artifact Carry Its Own Biography?
The larger problem may be the demand we are placing on the artifact itself. We are asking finished text to contain enough information to reconstruct how it was created.
A printed novel does not reveal whether its author used a spellchecker. A spreadsheet does not preserve every change to every formula unless someone maintains revision history. A photograph does not inherently record each transformation made in an editing program. We usually establish provenance by preserving records around an artifact rather than forcing the artifact itself to encode its entire biography.
AI-generated content may develop in the same direction. Anthropic is already combining approaches: supported Claude text can contain embedded watermarks, while supported generated files can carry digitally signed provenance metadata using the C2PA standard. A more mature provenance system could combine generation-time marks with signed metadata, platform records, disclosure by authors and institutional rules defining which forms of AI assistance actually matter. Different mechanisms could answer different questions rather than demanding that one statistical signal settle the entire history of a document.
That would move the debate away from the brittle question Does this writing look artificial? and toward a more useful one: What can we actually establish about how this artifact came into existence?
The rise of AI writing has exposed an assumption that was rarely worth examining before: that the sequence of words in front of us could serve as a reasonable proxy for the process that produced them. The same argument can now move through several languages, be summarized and expanded, have its sentences regenerated repeatedly, or pass back and forth between human and machine editors until little remains of its original phrasing. The semantic object can survive while its linguistic implementation changes.
Watermarking confronts that separation directly. Anthropic can place a signal into the choices Claude makes while producing words, and increasingly robust systems may force deliberate evasion to become more elaborate. Such marks could become genuinely useful for platforms, educators, publishers and regulators. They still cannot make authorship simple again.
Once humans and machines routinely participate in the same intellectual process, the important questions are no longer limited to who generated the final sequence of tokens. They concern who supplied the ideas, who made the decisions, who accepts responsibility for the result, and whether anyone was deceived about that process.
Those things reside somewhere deeper than the words. And meaning is much harder to watermark.
Comments
Post a Comment