Claude's Invisible Watermark Proves Less Than You Think
The EU made Anthropic mark Claude's text, but the mark says only that Claude touched it, never who did the writing.
The Oldest Argument in Authentication Just Got a New Exhibit
Hi, I‘m Cipher, the Codekeeper from the NeuralBuddies crew! Every field that ever cared who made a thing eventually invented a test. Seals, signatures, hallmarks, certificates of authenticity. And people eventually mistook every one of those tests for a verdict.
The art world learned this the expensive way. A certificate proves that somebody signed a certificate. Real authentication is a chain of small, boring pieces of evidence, and the specialists who do it for a living never rest the whole case on one result.
On August 2, Europe‘s transparency rules took effect. Days later, Anthropic confirmed that Claude now marks the text it writes. Ordinary prose just got its first real test.
The old mistake is already forming around it, right on schedule.
Table of Contents
📌 TL;DR
📝 Introduction
🔑 What Is Actually Hidden in the Words
🛡️ What the Mark Knows About You
⚖️ Involvement Is Not Authorship
📉 How Much of Your Own AI Use Shows
🎯 The Scanner Pointed at Writers Is a Different Machine
🧭 The Codekeeper’s Threat Model: Six Questions Before You Panic
🏁 Conclusion
📚 Sources / Citations
🚀 Take Your Education Further
TL;DR
Anthropic added nothing to your text. There are no hidden characters and no invisible ink. Anthropic changed how Claude picks between words that were already equally good.
The mark carries no identity. It holds nothing about you, your employer, or your conversations. Nobody can read your name out of it, because your name was never in it.
It costs you nothing. No extra words, no higher price, no meaningful slowdown.
It proves involvement, never authorship. The best it can say is that Claude probably touched these words at some point.
It fails in both directions. It can appear on writing that is mostly yours. It can vanish from writing Claude produced entirely.
How much shows depends on how you used Claude. A full draft marks strongly. A grammar fix barely registers at all.
The detection tool is not out yet. Anthropic promised one and has not shipped it.
A different machine is already scanning writers. Substack now runs posts through an AI detector, and detectors work nothing like this watermark.
📝 Introduction
A watermark on a banknote exists so you can see it. You tilt the note toward a window, the portrait appears in the paper, and you know the note is genuine.
Anthropic went out of its way to say that this is not that. The mark in Claude‘s text is invisible to a reader by design, and no amount of squinting will surface it.
That single difference reshapes the whole story. A banknote watermark answers a question you can ask yourself. This one answers a question only a key holder can ask, and it answers with a probability rather than a verdict.
Probabilities get treated as verdicts all the time. That is the actual risk to you here, and it has very little to do with the cryptography.
So this post does four things. It opens the mark up and shows what is inside. It answers the question you actually want answered, which is whether any of this points back at you.
It maps how much of your own AI use shows, depending on how you used it. Then it introduces the other machine, the one already reading writers today.
🔑 What Is Actually Hidden in the Words
Start with how a model writes a sentence at all. It produces one word at a time, and at every step it holds a shortlist of candidates. NeuralBuddies has a ground-up explainer on how a chatbot picks its next word if that idea is new to you.
Take Anthropic‘s own example. The sentence so far reads “The weather today was cold and...“
The next word will not be “sugary.“ It might well be “overcast.“ It might equally be “grey.“ Both are fine. Neither changes the meaning for you.
So how does the model choose? Normally a random number settles it. A coin flip, essentially, between two words that were always going to be acceptable.
Watermarking changes the coin. Instead of a plain random number, Claude derives the choice from a secret key plus the handful of words that came just before. The choice still looks random from the outside. It just came from somewhere specific.
Why that leaves a trace
Here is the part I find genuinely elegant, and it is old cryptographic thinking wearing new clothes.
If you hold the key, you can replay those choices. You can walk through a passage and ask, at every fork, whether the word that appeared matches the word the key would have produced. One match means nothing. Hundreds of matches mean something.
Anthropic reaches for a board game to explain it. Imagine a Monopoly game where nobody rolls dice. Instead the players take each move from the next digit of pi, starting at some random position in the sequence.
The game plays identically. No player could tell the difference. But afterward, someone who knows pi can look at the record of moves and work out that pi was probably running the game.
Claude‘s text works the same way. The reading experience is untouched, and the record still carries the fingerprint of what produced it.
The technique is not homegrown. Anthropic uses a version of SynthID-Text, which Google DeepMind published in Nature in 2024, and which traces back to a 2022 proposal by Scott Aaronson. DeepMind served a watermarked model to a slice of real Gemini traffic and compared user ratings against the unwatermarked version. The difference was not statistically significant.
Why any of this happened
None of it was Anthropic‘s idea.
Article 50 of the EU AI Act took effect on August 2, 2026. It requires providers of generative AI systems to mark their outputs in a machine-readable format, so that other systems can identify them.
Nature reports the penalties for missing that bar at up to 15 million euros, roughly 17 million US dollars, or 3% of global annual turnover.
Anthropic signed the accompanying Code of Practice on Transparency of AI-generated Content in July 2026, one of roughly 190 signatories in total. TechCrunch names Black Forest Labs, Google, Meta, Microsoft, OpenAI, and Synthesia among the companies committed to the same code. Google already marks text with SynthID.
One detail matters more than the rest for you. Anthropic applied the watermark worldwide rather than only in Europe, because it does not yet have a durable way to scope the behavior by region. A European law reached your account regardless of where you sit.
Models released after August 2 carry it immediately. Older Claude models fall under a transition period that runs to December 2. Anthropic says it will add the mark to those over the coming months.
🛡️ What the Mark Knows About You
Every threat model starts with the same question. What does the adversary learn?
So take the worst case. Somebody holds the key, runs it over a page you produced, and gets a strong signal back. What did they just learn about you?
Almost nothing.
The watermark carries no identifying information whatsoever. Nothing in the mark, and nothing in the key, exposes the user, the organization, or the conversation. There is no account number woven into the phrasing and no timestamp buried in the word choice. The mark reports on Claude, not on the person at the keyboard.
That is worth sitting with, because it is the opposite of what “invisible watermark“ sounds like. The phrase suggests a tracking pixel. This behaves more like a mint mark on a coin, which tells you which facility struck it and nothing about who spent it.
The other things it does not do
It adds no characters. Claude inserts nothing into your text. No zero-width spaces, no odd unicode, no hidden payload for a spell checker to trip over.
It costs nothing. The watermark produces no extra words, so the price to run the model is unchanged. The speed impact is negligible.
It changes no rights. Ownership of the output and your rights under Anthropic‘s terms are exactly what they were. A mark that says Claude was involved says nothing about who owns the result.
It survives the clipboard. Because the mark lives in the word choices themselves, it travels when you copy and paste. It also persists through light editing, though that is a matter of degree rather than a guarantee.
Files work completely differently
If Claude produces an image or a similar file, you get something else entirely, and the distinction matters.
Supported files such as PNG, JPG, and SVG receive a C2PA content credential. C2PA is an open industry standard, the same one camera manufacturers and photo editors use to record where a file came from. The credential is a small signed note in the file‘s metadata saying Claude was involved.
Nothing inside the image changes. No pixels get nudged, nothing is embedded, nothing is hidden. It is a label attached to the outside of the box rather than a pattern woven into the contents. Any tool that reads C2PA can check it, and Anthropic says it will provide one.
⚖️ Involvement Is Not Authorship
Now the claim the whole post rests on.
Anthropic states the limit plainly on its own announcement page: the watermark “cannot distinguish ‘Claude wrote this‘ from ‘Claude heavily edited this.‘”
Read that again, because a lot of people are about to skip it. The mark reports that Claude was probably involved with a piece of text at some point. That is the ceiling. It cannot rank how involved, cannot separate drafting from editing, and cannot tell you whose ideas these were.
It fails toward you
Suppose you wrote something yourself and asked Claude to translate it into Spanish. Every Spanish word came from Claude, so the translation carries a full-strength watermark. The thinking was yours. The mark does not know that.
Suppose you handed Claude a rough draft and asked for a heavy rewrite. The argument, the evidence, and the structure are yours. Enough of the wording is Claude‘s that a check will find the signal.
In both cases the work is substantially yours and the mark reads positive. A person who treats that as proof of cheating has misread their own evidence.
It fails away from you too
The reverse holds just as firmly.
Short passages carry too few word choices to register. Paraphrased text loses the signal. Code carries almost none, because code must be exact and exactness leaves no room for a mark.
Anthropic is explicit that a heavy rewrite can wash the watermark out completely. It also makes the obvious point: text rewritten word for word is arguably no longer AI-generated anyway.
So a clean scan proves nothing at all. Absence of the mark is not evidence of a human hand.
The honest counterweight
I would be selling you a comfortable story if I stopped there, so here is the result that cuts the other way.
Nature reports that organizers of the ICML 2026 machine learning conference tried this in practice. They added a watermark to papers sent out for peer review, designed to produce telltale text if a reviewer fed the paper to an AI. They caught 506 reviewers breaking a no-AI policy.
Nihar Shah of Carnegie Mellon University, who ran that process, drew the sensible conclusion. Careful evasion is possible, but plenty of people simply copy and paste.
That is the real picture. Reese Richardson, a metascientist at Northwestern University, told Nature that watermarks are easy enough to strip that motivated bad actors will not be stopped. Both things are true at once. A weak signal still catches careless misuse, and it still cannot carry the weight of an accusation.
📉 How Much of Your Own AI Use Shows
Here is the practical map, and it follows one rule.
The watermark lives in freedom. It can only occupy the moments where Claude had a genuine choice between equally good words. Take the choices away and there is nowhere for the mark to sit.
Cryptographers have a name for that room to move. They call it entropy, and a hiding place needs it the way a signature needs a blank line.
So the strength of the signal tracks how much choosing you handed over.
You asked for a draft from scratch. Strong mark. Claude made thousands of free choices, and nearly all of them can hold a piece of the pattern.
You asked for a translation. Strong mark. Anthropic notes that every word in a translation is Claude’s, even though the ideas came from the original.
You asked for a heavy rewrite. Moderate mark, scaling with how much of the wording changed.
You asked for grammar and punctuation fixes. Barely anything. If Claude only corrected commas, the mark can live in a handful of corrections, which is usually too little to register.
You asked for code. Almost nothing. Code has to be exact, so there is no free choice to encode. Some signal can appear in comments, where the wording is arbitrary, with no effect on the code itself.
You asked a fact-dense question. Sparse mark. Anthropic’s example is a sentence about Isaac Newton’s Principia Mathematica, where only one word is correct and the watermark has nothing to work with.
You generated something short. Likely nothing usable. Confidence rises with length, and a paragraph may simply be too small a sample.
There is a quiet reassurance in that list for most readers. The heaviest everyday uses of Claude, cleaning up an email and tightening a paragraph you already wrote, are the ones that leave the least behind.
🎯 The Scanner Pointed at Writers Is a Different Machine
Now the part that actually affects your week.
Anthropic‘s watermark is not yet checkable by anyone. The detection tool exists as a promise on a web page, and the company says it is still working out the implementation. As things stand, nobody outside Anthropic can run this check on your writing.
Meanwhile a completely different technology already reads writers, and readers keep confusing the two.
Detectors have no key
An AI detector like Pangram never sees Anthropic‘s key and never checks for a watermark. It guesses, from style. It looks for the tells that machine writing tends to leave: phrasing habits, rhythm, word frequencies that drift away from how people actually write.
Anthropic named two of those tells in its own announcement. Models are fond of the construction “this isn‘t [X], it‘s [Y],“ and they use the word “quietly“ far more often than a person would. You may notice this post has gone out of its way to avoid both. Now you know why.
NeuralBuddies has a whole piece on the punctuation mark that gives AI away if you want the full catalog of tells.
The two systems have opposite error profiles, and this is the distinction worth carrying:
The watermark has a key and makes a weak claim. Mathematically grounded, and it only ever says Claude was probably involved.
The detector has no key and makes a strong claim. It guesses, and then it reports a confident percentage of how much of your text is machine-written.
Where that lands on you
Substack turned on Pangram scanning on July 21, 2026. A reader can now scan any post, Note, reply, or comment longer than 100 words, and the result estimates how much of it a person wrote.
Pangram says its latest model has a false positive rate of 0.01%. That is the company‘s own figure, and the track record for such figures in this category is poor. Turnitin claimed a false positive rate under 1%. A 2023 Washington Post investigation put the real number above 50%.
Matteo Wong of The Atlantic put the stakes bluntly, warning that AI accusations “could very quickly spiral into a witch hunt.“
There is a fairness problem underneath as well. A 2025 study found that some AI detectors flagged writing by neurodivergent people more often than writing by neurotypical people. That study tested an OpenAI detector rather than Pangram, so treat it as a warning about the category, not a verdict on one product.
Substack‘s own CEO, Chris Best, acknowledged the tool is imperfect while arguing that readers benefit from having it.
So the scoreboard for a working writer looks like this. The rigorous, key-backed system makes a modest claim, and nobody can run it yet. The detector makes a confident claim, and any reader can point it at your posts right now.
🧭 The Codekeeper’s Threat Model: Six Questions Before You Panic
Ask what the mark can actually say. The honest ceiling is that Claude probably touched this text at some point. Any claim stated with more confidence than that goes beyond the evidence.
Ask who holds the key. Only Anthropic does, and the public detection tool has not shipped. Anyone claiming a watermark result on your writing today deserves a follow-up question.
Ask whether absence got treated as proof. A clean scan means nothing. Short, paraphrased, and heavily edited text all read clean, and so does a different AI entirely.
Say how you used it before anyone asks. Substack offers an optional note where you describe your process, and you can scan your own draft before you publish it. You can also flag a scan result you believe is wrong. A sentence you wrote yourself beats a percentage somebody else assigns you.
Keep your drafts. Version history, notes, and outlines establish how a piece came together. Chain of custody settles more disputes than any single test result ever will.
Check which machine flagged you. A watermark check and a style detector are unrelated tools with unrelated failure modes. Once you know which one produced the number, you know how far to trust it.
🏁 Conclusion
Cryptographers spend a lot of time on a distinction that sounds pedantic until it matters. There is what a system proves, and there is what people believe it proves. The second one is usually larger, and the gap is where the damage happens.
This watermark is a careful piece of engineering. It costs nothing, degrades nothing, and reveals nothing about you. Anthropic was unusually candid about its limits, and the limits are severe. It reports that Claude was probably involved, and it stops there.
The trouble starts when a probability about involvement gets waved around as a verdict about authorship. A translated essay reads guilty. A wholly machine-written paragraph reads innocent. Anyone who treats either result as settled skipped the part where you ask what the test measures.
You do not need to change how you work. You need to know what the mark says, so you can correct anyone who claims it said more.
Privacy by design, not by accident. That principle cuts both ways here, and the second direction is the one to hold onto. Never conclude so much from a system built to reveal so little.
Hold it to the light all you want. The mark was never meant for your eyes.
-- Cipher 🔐
Sources / Citations
Anthropic, August 14, 2026. How Claude’s text watermarking works. Anthropic. https://www.anthropic.com/news/claude-text-watermark
Ivan Mehta, August 11, 2026. Anthropic says it will watermark text generated by its AI models. TechCrunch. https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/
Elizabeth Gibney, August 13, 2026. Can Anthropic’s invisible watermarks curb ‘AI slop’? Researchers remain sceptical. Nature. https://www.nature.com/articles/d41586-026-02503-7
Bruce Gil, August 12, 2026. Anthropic’s Claude Will Start Adding Invisible Watermarks to AI-Generated Text. Gizmodo. https://www.gizmodo.com/anthropics-claude-will-start-adding-invisible-watermarks-to-ai-generated-text-2000797759
Devin Pavlou, July 22, 2026. Substack’s new AI-detection tool has some critics worried about false flags. Straight Arrow News. https://san.com/cc/substacks-new-ai-detection-tool-has-some-critics-worried-about-false-flags/
Net Influencer, July 22, 2026. Substack Launches AI-Detection Tools With Pangram To Flag AI Generated Content. Net Influencer. https://www.netinfluencer.com/substack-launches-ai-detection-tools-with-pangram-to-flag-ai-generated-content/
European Commission. Code of Practice on Transparency of AI-generated Content. Shaping Europe’s digital future. https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content
Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, et al., October 23, 2024. Scalable watermarking for identifying large language model outputs. Nature. https://www.nature.com/articles/s41586-024-08025-4
Take Your Education Further
Envisioning AI Governance: A Path to Fairness and Stability: A NeuralBuddies look at how AI rules get written and enforced, which is the machinery that produced the watermark in the first place.
Top 10 AI Safety Tips to Protect Your Privacy: A NeuralBuddies guide to what these tools actually record about you, the practical companion to a mark that records nothing.
The Top 10 Misconceptions About Artificial Intelligence: Separating Fact From Fiction: A NeuralBuddies roundup of the beliefs about AI that do not survive contact with the details, of which “the watermark will catch you” is the newest.
Disclaimer: This content was developed with assistance from artificial intelligence tools for research and analysis. Although presented through a fictitious character persona for enhanced readability and entertainment, all information has been sourced from legitimate references to the best of my ability.





