On 11 August, Anthropic updated a help page. The announcement that the industry then spent two days arguing about arrived as customer support documentation, in an article titled "How Claude marks AI-generated content". The relevant sentence is short:
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won't see it, and it doesn't change the meaning, quality, or readability of Claude's response.
That is a bigger deal than the venue suggests. Every model Anthropic launches from 2 August onwards carries a hidden signature in its prose. It applies across the API, the Claude apps, Claude Code, Cowork and Tag, and it follows Claude onto AWS, Google Cloud and Microsoft Foundry. It is on "wherever Claude is offered, worldwide", not just in Europe. You cannot switch it off, and no plan tier or API parameter disables it. Files get separate treatment: images and SVGs come with signed C2PA provenance metadata, which is a different and much more fragile thing.
John Gruber's headline captured the problem neatly: Anthropic posted how Claude marks AI-generated content without explaining how Claude marks AI-generated content. The page describes what a mark means and stays silent on what it is, promising "details on detection mechanisms in forthcoming technical documentation". Two days on, the technical documentation has not appeared and neither has a detector. The Register, less politely, filed it as a sop to Brussels.
Three questions follow. What is this thing, is anyone forcing Anthropic to do it, and can you get rid of it?
What the mark actually is
Anthropic hasn't said, so we reason from the family of techniques it must belong to.
A language model does not write words. It produces a ranked list of probabilities for what the next chunk of text might be, and then something else picks one. That second step is where a text watermark lives. A secret key designates some fraction of the vocabulary as "preferred" at each position, and the sampler nudges its choice towards those words. Any single choice looks unremarkable. Across a few hundred choices, the pattern becomes statistically loud, and a detector holding the same key can spot it.
Two consequences fall straight out of that diagram, and they matter later. The model is not aware of the watermark, because the watermark happens downstream of the model. And there is nothing hidden in the characters. No zero-width spaces, no funny Unicode, no invisible payload. The mark is the word choice.
Anthropic is refreshingly blunt about what a hit means. A detected mark says the content "may have been processed by Claude", which the page calls "not fully conclusive". Processed, not written. If you drafted something yourself and asked Claude to tidy the grammar, your writing may now carry the mark. Anthropic says so directly, listing proofreading, translation and summarising as things that can leave a signature on text whose ideas were entirely yours. The absence of a mark proves even less.
That asymmetry is the part schools and HR departments will get wrong, and it is already biting. The radio host Erick Erickson, quoted by Forbes, had switched from Grammarly to Claude for proofreading: "now the stuff I've written will be watermarked that Claude did the work. This is ridiculous."
Does it make Claude worse?
Anthropic says the mark "doesn't change the meaning, quality, or readability" of a response. It published that sentence without a benchmark, an evaluation or a paper to stand behind it. You are asked to take it on trust, and there is currently no way to check, which is a strange position for a company to adopt on a technical question it could settle with one table.
The published research on comparable schemes does not support a blanket "no cost" claim. An EMNLP 2024 evaluation of green-list watermarking found drops of 10 to 20 per cent in effective utility on classification tasks, around 7 per cent on multiple choice, and 5 to 15 per cent on long-form generation, and those are the best-case numbers after the authors tuned the settings to do the least damage at a fixed detection strength. Code is worse. The SWEET paper measured a standard watermark costing 24.3 per cent of HumanEval pass@1, 36 per cent on MBPP and 67.3 per cent on DS-1000, while still detecting reliably. Their explanation is the intuitive one: code has very few genuinely free choices, so forcing one token off the model's preferred pick can break the program outright.
Google's SynthID-Text is the closest thing to a precedent, and its Nature paper is worth reading for the fine print. The scheme is called "non-distortionary", which sounds like a guarantee of no quality loss. What it actually means is that the output distribution is preserved on average over the random seed. Not per response. The authors grade non-distortion from weakest to strongest, say SynthID-Text is configured at the middle setting, and concede it comes "with some reduction to inter-response diversity". They also state the trade-off plainly: weaker non-distortion hurts quality, stronger non-distortion hurts detectability. There is a dial, and someone at each company has chosen where to set it.
So is Claude worse since 2 August? Here I have to report an absence rather than a finding. There is no credible independent evidence that it is. I looked for benchmarks, comparisons, anything tied to the cutoff date. What exists is people reasoning from first principles in the days after the announcement, which is not the same thing. A commenter on the Hacker News thread called the promise of no quality impact "laughable" because "the tokens will be arranged in a very specific manner", and he is right about the mechanism and has measured nothing.
The honest position is this. Anthropic has made an unfalsifiable claim, because "no quality impact" cannot be assessed without knowing the scheme and the operating point. The literature says schemes in this family cost real accuracy, most sharply on code and short outputs. Nobody has caught Claude degrading. All three of those are true at once.
Is the law making them do this?
Partly, and the gap between "partly" and "entirely" is the interesting bit.
Article 50(2) of the EU AI Act requires providers of generative systems to ensure outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated". Text is named explicitly, alongside audio, image and video. It became applicable on 2 August 2026, which is exactly the cutoff Anthropic chose for its models, and Anthropic points at the accompanying Code of Practice as its reason for acting. The compliance motive is not an inference; it is the company's own stated rationale.
But read the next sentence of the Article. Solutions must be effective and robust "as far as this is technically feasible", weighing implementation cost and the state of the art. That is a soft edge, and the Code of Practice leans on it hard. The Code, finalised on 10 June 2026 and signed by around 190 organisations including Anthropic, Google, OpenAI, Microsoft, Meta and Mistral, concedes in its own recitals that "no single marking technique suffices to meet the four requirements in Article 50(2)". For most media it therefore demands two layers, signed metadata plus a watermark. For plain prose it gives up on the first layer, "given that free-form text cannot transport metadata", and accepts the watermark alone. It also writes off anything under 200 tokens as too short to mark reliably, and floats restricting detection to "verified expert users" precisely because text marking is weak.
The teeth are real: breaching Article 50 sits in the tier attracting fines up to 15 million euros or 3 per cent of worldwide turnover. Since the Digital Omnibus took effect on 27 July 2026, enforcement against a company that builds both the model and the assistant on top of it, which is exactly Anthropic's position, sits with the AI Office in Brussels rather than a national regulator. The same Omnibus gives providers whose systems predate 2 August until 2 December 2026 to comply, which is why Anthropic's older models are still unmarked.
None of that required Anthropic to do what it did. The law is technology-neutral and qualified by feasibility, and it stops at the EU border. Watermarking text is the hardest modality, and Anthropic chose to do it worldwide, at the model level, with no opt-out. Signing the Code buys no presumption of compliance; the Code says so itself. The most plausible reading is that maintaining two inference paths for two continents is miserable engineering, and being first looks good.
For contrast, China has required this since 1 September 2025, names text first in its list of covered content, and demands a visible label as well as an embedded one. California's AI Transparency Act, the flagship American effort, applies to "image, video, or audio content" and does not cover text at all.
Can you remove it?
This is where the internet has lost its mind, so let us separate the mechanism from the marketing.
The stuff that does not work
Within 48 hours of the announcement, a small industry of "Claude watermark removers" appeared. Most of them are worthless, and in an instructive way.
Take claudewatermark.com, which offers to "remove the watermarks for AI systems like Claude". What it does is strip zero-width spaces and invisible Unicode, and it ships a button to remove em dashes. Both operations are irrelevant. There are no hidden characters to strip, and the em dash thing is folklore about spotting AI prose by eye, with no connection to any watermark. The site's own FAQ quietly admits the implementation "remains partially undisclosed", which is a curious thing to say underneath a tool that claims to defeat it.
StealthGPT went further, bannering that it "now removes Claude Watermarks" and promising the signal "doesn't survive the rewrite". Its evidence is a set of scores against Turnitin and GPTZero. Those tools guess at style. They hold no key and have no access to Anthropic's scheme, so the mark is invisible to them. Passing them tells you nothing whatsoever about whether a statistical watermark survived. The claim is worse than unproven: with no public detector in existence, nobody can currently prove it either way.
Then there is the tip that has spent the past two days working its way across LinkedIn: paste your Claude output into ChatGPT, ask it to remove the watermark, and you are clean. It arrives with all the confidence that platform is famous for, generally from people whose profiles say they advise others on AI. Being precise about why it fails is worth a paragraph, because the reason is genuinely not obvious.
Go back to the first diagram. The mark is applied after the model, by the sampler. The model that later reads your pasted text sees ordinary words. It holds no key, and it cannot tell a marked passage from an unmarked one. Asked to remove a watermark it cannot perceive, a chatbot will do what chatbots do: agree enthusiastically and hand your text back, perhaps with a few words changed and a note confirming the watermark is gone. You will have been given a confident assurance about something the model has no access to. The same reasoning makes retyping useless: same words, same order, same evidence.
There is a genuinely subtle point buried here, though. If instead of asking the model to strip a watermark you have it rewrite the whole thing, that does work, and it works well. Not because the model understood the request, but because a full rewrite regenerates every word through a different sampler. The instruction is nonsense; the side effect is the strongest attack there is. People who follow the bad advice sometimes succeed anyway, for reasons that have nothing to do with what they asked for, which is exactly how folklore survives contact with reality.
The stuff that does
Paraphrasing is the real attack, and the numbers are serious. Sadasivan et al. report that recursive paraphrasing of watermarked passages around 300 tokens long dropped detection from 99.3 per cent to 9.7 per cent, measured as true positive rate at a 1 per cent false positive rate. Translation works too: a round trip through Chinese cut detection to 0.54 and 0.61 AUC for two standard schemes, essentially coin-flipping, while summarisation scores held up. Google's own scheme is not exempt; a TrustCom 2025 assessment found back-translation dragging SynthID-Text's true positive rate from 1.0 down to around 0.68.
Now the correction, because the "watermarking is already dead" narrative oversells its own sources. The papers most often cited as burying watermarks say something more careful when you read them. Krishna et al., whose DIPPER paraphraser collapsed DetectGPT from 70.3 per cent to 4.6 per cent, only pushed watermark detection from 100 per cent to 57.2 per cent with the same attack, and concluded flatly: "watermarking is the most resilient detector to paraphrasing". Kirchenbauer et al. paid fourteen graduate students to try to scrub watermarks by hand and found the signal still detectable after about 800 tokens. The famous "watermarking is impossible" result proves the impossibility of strong watermarking under an attacker with two idealised oracles, and states in as many words: "We do not investigate weak watermarking schemes in this paper."
The uncomfortable finding for everyone is a different one. Jovanović et al. showed that for around 42 dollars of compute an attacker can learn enough about a scheme to both scrub marks and forge them, with over 80 per cent success. Forgery is the scarier direction. It means a watermark hit can be manufactured onto text a human wrote, which should trouble anyone planning to treat detection as proof.
The one method that always works is refusing to play. Run open weights on your own hardware and you control the sampler, so there is nothing to remove. Every other approach on the left column of that chart costs you something: a machine's prose instead of Claude's, or hours of genuine rewriting.
The startup that solves this
If you want to build the business, the shape of it is already clear, and parts of it exist in the literature. Recent papers report rewrite attacks defeating seven watermarking schemes at under a dollar per million tokens, and reinforcement-learning approaches claiming 98.5 per cent removal.
A serious product would run an open-weight model locally, so no second watermark is introduced, and rewrite sentence by sentence rather than wholesale, since the signal accumulates across positions and you only need to break enough of them. It would gate every rewrite behind a quality check, using the model itself to score whether meaning and tone survived, and retry the ones that did not. The interesting engineering is the loop: once Anthropic ships the promised detection API, you can score your own output against it and keep rewriting until the mark falls below threshold. A public detector is also an oracle for defeating it, which is precisely the bind Anthropic is in, and probably why the Code of Practice muses about restricting detection to "verified expert users".
The counter already exists too, and it is not a better watermark. Krishna et al. showed that keeping a database of generations and matching against it catches 80 to 97 per cent of paraphrased text. Retrieval beats rewriting, because paraphrasing changes the words but not the substance. It requires providers to log what they generate, which trades one privacy problem for another.
As for legality, the answer in Europe is odder than you would guess. I went looking for an anti-circumvention rule in the AI Act and there isn't one. The words "tamper", "strip" and "forge" do not appear in the Regulation at all, and its penalties reach "operators", a closed list that does not include a stranger erasing someone else's mark. The old copyright anti-circumvention rules do not stretch either, since a provenance mark restricts no access and purely AI-generated text probably attracts no copyright to protect. The nearest thing to a prohibition lives in the voluntary Code, where signatories promise not to "place or make available on the market" tools for circumventing marks, which binds Google and OpenAI and precisely nobody else.
So the humaniser startup is not illegal in the EU for removing watermarks. It is, however, almost certainly an AI provider in its own right, and the Commission's Article 50 guidelines name "paraphrasing or rewriting text that changes style, structure and meaning" as an activity that triggers the marking duty. The startup would be legally obliged to watermark its own output, on pain of a 15 million euro fine, while remaining perfectly free to destroy Anthropic's. Regulated, but not for the thing it exists to do. In China, by contrast, the same business is flatly illegal: the labelling rules ban removing marks and separately ban providing tools or services for others to do it.
What to actually do
If you teach or hire, write the caveat into your policy now, while it is still hypothetical. A mark means the text may have been processed by Claude, which is consistent with a student running a spellcheck. Treat it as one piece of context and do not let it decide anything on its own. Forging a mark onto innocent text costs about the price of a takeaway.
If you build on Claude, do not lean on this for provenance of code. Low-entropy output barely carries the signal, and your formatter will finish off whatever survives.
And if you are quietly hoping a tool will scrub your drafts clean, save the subscription. The honest operators in that market say so themselves. One of them, selling Unicode cleanup, puts it better than I could: "a statistical signal in word choice isn't a character you can strip. Tools that claim to remove Claude's watermark are selling something that doesn't exist."
Not yet, anyway. Somebody will build it properly, because the research is public and the attack costs about a dollar. And the first detector Anthropic ships will double as the scoreboard for beating it.