Text AI watermarks will always be trivial to remove
seangoedecke.com
seangoedecke.com
Currently I can recognize AI text because I read thousands of ai generated text. I know that 110% of yahoo finance news is generated. I don't want to read an AI generated personal blog, but if I do what's the problem really? Other than the companies distinguishing AI text for getting better training data, how do people benefit from watermarked text exactly?
Really, Anthropic (and the other labs that follow) are just trying to satisfy the requirements of the law so they can continue to serve the EU. Wether it's actually effective is something else entirely.
Malicious compliance is to make cookie acceptance much easier than refusal. Lack of oversight doesn’t punish this. Now, the incentive is to make life hard for people who decline cookies. That results in more annoying pop-ups.
That isn't malicious compliance, it is just non-compliance. If the equally difficult rule wasn't there then it would be malicious compliance.
The people that make the laws and the people complaining about AI don't know or care about the reality of the situation nearly as much as they care about being re-elected and feeling good about their social posture.
Well, the negative effects of AI are already visible, and I'm not talking about job losses here. I'm talking about AI hallucination, about mass generated spam, about AI-assisted scamming (apparently Indian scammers are shifting to use AI systems [1]) and about political manipulation.
Governments usually are slow to react, but eventually they do react, when a technology poses more issues than it is worth.
[1] https://timesofindia.indiatimes.com/gadgets-news/new-ai-scam...
Additionally, think of cases like paying a lawyer or an expert for an extensive report or opinion on something. Wouldn’t you want to know if that is actually their carefully assembled professional assessment rather than the output of an LLM prompt?
This article says any AI watermarking can be defeated trivially, by running the text through a tool which strips weird Unicode characters (fine) and then runs it through an LLM which replaces all words with similar words (whaaattt?). Running text through another - probably much lower quality llm which scrambles the words you use would make slop even sloppier. And it’s another whole step you have to know to do. A great many people who use LLMs to avoid doing work will not know about these extra steps, or not make use of these sort of tools.
I would say that's not necessarily the case unless there are zero false positives. In fact, your university situation is exactly where a detection system that works 50-90% of the time would be a nightmare if a meaningful share of the 10-50% errors were false positives.
EDIT: Sorry, I misread your comment above. Yeah, hopefully orders of magnitude fewer kids than the number who are getting falsely accused of cheating with LLMs now.
Increasing the accuracy of these systems - both in terms of false positives and false negatives - seems like a good thing.
If some kids are false positives (detected as using ai but didn’t)then how are they lazy?
A proper watermarking system should be able to have an arbitrarily high accuracy - as many nines as you want. And it should be able to actually report the accuracy of its judgements.
If you're worried about kids being falsely accused of cheating using LLMs, you should be cheering on these developments.
Probability of a false positive is 3 x 10^-5 in one example.
Seems like about 3 collisions per 100 million pictures. If everyone have 1000 pictures that is 3 collisions per 100 000 users.
Raiding 2 people in my city for made up CSAM pictures would be way to high false positive rate.
But you can also make any picture match a CSAM hash by adding picked noise.
Im very confused how this is even supposed to work at face value.
1) If the verification can be done by anyone, then anyone can bypass it.
2) If it can only be done by anthropic then the government or whomever has the special privilege (not everyone otherwise this is just #1) has to make a specific request
Is the point of the legislation to accurately classify text in general or to simply detect the true positives?
Why not both? Stenographically fingerprint llm output to make cheating risky. And use other approaches too.
Most people have no idea open source LLMs exist, let alone how to use them. Right now, they’re much worse than the frontier models.
What about false positives? Imagine being a honest student and then the software declares your work to be AI-generated. How do you defend against that claim? The detection software is a black box and is likely running as a cloud service, so that you have no realistic options for reverse-engineering the false positive detection.
How do you think you will know the text was AI generated but the person trying to deceive you won’t be able to undo the watermark?
Nothing will save us. You can't automate trust.
Does it sound good?
Either anthropic does not tell anyone the signal and only they can verify the text. Or they share the signal and everyone can verify it, which means everyone can bypass it and its just an inconvenience.
Companies will implement it in the worst way possible? Nowhere in the cookie laws does it say you need to add a banner, you just can't spy on users without their consent.
Had google offered a privacy-preserving ad model (basing the ads off the content of the page instead of who the user is) things would be very different.
This will most likely bring a sift end to at least the low hanging fruit.
The only way to prevent AI cheating is to make the assignments in class in pen and paper. To be honest, I don't understand why schools moved away from that in the first place.
Cost and convenience.
Being able to work on computers does help people with bad handwriting, which includes some disabilities.
How does good handwriting improve people's live significantly? I know some people who prefer handwritten notes and I think there is research that says they can be more effective when learning, but most adults very rarely write more than a few lines by hand.
Disabilities don't disappear after school, and children who struggle with motor control in school will continue struggling with it on adulthood. Half the point of school is finding out your weaknesses and learning to overcome them in a safe and controlled environment.
Pushing someone with a disability to improve their handwriting will not cure their disability. How does making people struggle more help them? its like suggesting kids in wheelchairs should be made to try to walk.
Learning to use a computer as a means to overcome a disability that makes handwriting painful or difficult or illegible sounds like it'd qualify as exactly that.
And taking away writing from people with bad handwriting (if not related to disability) is just about the worst thing to do if you want them to ever get better at it.
But if we're talking about deterministically taking some watermarked LLM output and having a function removeWatermark(text), it won't necessarily be "trivial" to remove, because the watermark function itself need not be public. Only the API that tests for the watermark need be public, right?
Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. And if Anthropic controls those variables, it would be very difficult to determine the rule, even with some pretty advanced analysis (I would assume). And keep in mind, that example rule I mentioned is pretty naive, too. I expect the actual rule would be way more advanced and not so straightforward as "swap every <charX> for a <charY>"
For Anthropic, it’s highly likely that the watermark is a SynthID type mark similar to the one that Sean is talking about (I actually ran the analysis here https://johnjwang.com/post/2026/08/12/how-claude-watermarkin...). When we get confirmation of whether all models actually are watermarked, I think we’ll be even more confident.
Of course it’s always possible that Anthropic has come up with a proprietary scheme, but I think it’s definitely harder to implement.
I think the game will be a cat and mouse game similar to LinkedIn and other websites trying to block scrapers: each iteration makes it harder for someone to figure out the watermarking scheme, but likely not impossible
Serious question… what is the point if you cant verify the watermark? Then only the model provider will know.
Is that what this law is about? I thought it was so anyone would know (which of course means anyone can bypass).
the government can ask .... who? If the claimed text is genuine, then the government is responsible for submitting the text to ...? all frontier labs? surely not every model within a lab is going to watermark in the same fashion?
I'm so very confused on this implementation and would want to see an expected use case.
There's also literal language translation. Generate in language A, translate to language B (Either "manually" by being proficient in it, or with non-LLM translation).
A very, very significant portion of the world knows more than 1 language.
And the more hoops you make cheaters jump through, the better.
Perhaps, but no tools needed for those who know more than one language, which is a lot of people.
Watermarking will never be a perfect tool. But there’s a lot of value in making low effort llm slop detectable. Even if high effort llm slop is still undetectable. Don’t make perfect the enemy of good.
If your goal is to defeat an AI detector for kicks, this might work. If your goal is to save effort by using AI to produce text, I don't think this works.
You don't think so based on what? One does not have to be a subject matter expert to translate from one language to another.
People that know more than one language are translating all the time, both consciously and unconsciously (While they're proficient in language A, they may "think" in language B). It's not as much effort as you think.
Catch true positives? Sure
Reliably determine if text was generated by AI? Not even a little.
Doesnt this and probably all techniques require the validator to know which portion of the text to validate?
If its not all generated together then how could it reliably carry the mark? Sure, run it against the full text. But what if the full text was not one-shot by the llm?
In other words, in order to reliably detect if the text is ai you need to first determine which part of the text was generated together by ai.
People underestimate the value of rules that only take malice and a little knowledge to break.
And they tend to exaggerate that underestimation if they... don't like the rule.
Of course I would want my code formatting tool to normalize that all to plain 0x20 spaces. But it would still be a helpful "brown M&M" test of did you even read CONTRIBUTING and run the code formatter before submitting this PR?
Most people are generally lazy. They upload text to LinkedIn full of "genuine", "honest" and "load-bearing".
And the people who get LLMs to make content for them are often the laziest of us all! Some people can’t even be bothered to remove their prompt from the output. Or they submit work which ends with “is there anything else you want to know about ___?”
Go ahead. Make tools to rip off the llm stenography. Many of the people who would proactively use tools like that are writing content by hand anyway.
Yes, this watermark will be easy to strip. It is still valuable for the vast majority of times where people just don't.
E.g. research paper, law makers, lawyers, state policies, notaries,...
These are much longer content and thus statistically they will disclose a better guess at AI generated content.
Asking another AI to paraphrase will not erase the mark (which they are unaware about) but rather cumulatively add their own mark and make it easier to detect.
The problem is not to use AI, but to endorse the responsibility of the content you (as a human) deliver and somehow make sure that fake-news, biased content or unverified output is detected as early as possible.
Maybe some part of it is that the deluge of slop is uncovering how poorly/sloppily these social institutions were working in the first place.
It is as always (think cybersecurity) the cat and mouse game, what is a weapon is also a defense. You, the simple fact that you are on this very site, means that you are probably more educated to AI than most, so it is not necessarily you that will really benefit from any constraining framework.
There are many people believing in many weird theories and those are more prone to be convinced by a nice narrative. AI did not introduce that, it just made it easier and available to anyone with any intent.
It makes it harder, because amount of bullshit goes up and it is easier to generate plausibly sounding bullshitm
Fictive example: imagine the bias is to inject notion of colorfullness in the text regardless or its content:
"I eat an apple" becomes "I eat a red apple"
The 2nd AI is about injecting size component, so the text becomes " I eat a big red apple"
There you get both watermarks. It is not that obvious obviously
Here's what I grasp: The AI system scores each token and then selects tokens based on those scores. If we encode something in the token selection routine ('in order choose the 1st, 3rd, 1st, 5th, 2nd, then 1st highest scored tokens'), we can identify AI-generated text by comparing sample text (ST) to the expected text (ET) for that prompt.
1) How do we score the tokens for the ET without the original prompt? Even a Markov-like process needs to start somewhere.
2) To recreate ET don't we need to maintain, until the end of time, the AI state - entire model and code - at the time of ST output?
3) Doesn't #2 require maintaining all states for all AIs? Often you won't know when and from which AI system the ST might have been generated. What happens when an AI vendor goes out of business?
4) To recreate ET, don't we effectively have to rerun the prompt? Won't rerunning it for every verification increase most costs of AI output by an order of magnitude? Most of what AI vendors do would be ST validation.
2. same answer
3. no, it just needs the previous text, private key, and the matrix math (CPU is fine)
4. no, see above.
1. So we can't score the first paragraph (or similar-sized block), and not short texts? Not deal-breaker, but a limitation.
2. Doesn't the score vary by each AI system state - its model, programming, harness, etc.? Claude's output today doesn't match Gemini's, nor Claude from 2 years ago.
Whatever happened to just delivering the best product or service? Why must tech be full of ninnying nannies that act against their users, "for their 'safety'‽"
Just because you want to ignore it doesn't mean it's not there.
Society then decides on what morals we want to uphold, or at least tolerate, vs. those we want to change.
Laws are hard power to change behavior. Influence/goodwill/etc. are soft power ways to change behavior and bend organizations in ways they may not otherwise want to.
Would you pay a pto keep grandma from eating dogfood if it wasn't the law that you have to put a portion of your earnings into Medicare and Social Security?
Ok say i do that on 30% of bullet points. Karen the hiring manager is vehemently anti AI. She gives my resume to her AI scanner and what does she find? This not a rhetorical question. Will it treat the text as a whole and not find it? Does it scan every combination of contiguous terms? It could scan bullet points but i could generate in pairs of 2. What about novels?
Then I tell you: calculate the average of every 3rd flip minus the average of every second. It should come out close to zero, and unlikely to be higher than +/- 30 but for my sequence it comes out to +89. Ok, that’s very unlikely.
So from now on we share this secret. Whenever I generate sequences of coin flips they all came out with this particular statistic out of whack, but unless you’re looking for it, you can’t detect it. In fact, in the real world I use cryptography such that unless you know my secret key, the specific statistic in question is mathematically undetectable.
Now use this sequence of coin flips to pick amongst next tokens in an LLM. It doesn’t change the distribution of words used in the LLM. It doesn’t change its writing style. In fact, unless you know the secret key you cannot detect the watermark. So it’s not that some words are used more often. That would be detectable. It’s just that if you translate back to a series of 0’s and 1’s the average of every 3rd bit minus the average of every 2nd is out of whack.
"I'm sure glad the LLM was able to fix up my grammar and imagery in the anonymous treatise I made criticizing the authoritarian regime. Hold up, someone's knocking at my door..."
It just seems like once you try and detect a signal against something that wasnt one shot it breaks down fast. Unless the signal is constrained to a very small space (pairs of words) but then quality and stealthiness must suffer massively.
Im curious about situations like this.
Paper submitted with some headings, title, and 3 paragraphs. 1 and 3 mostly generated. generated ones have a small percentage of sentences rewritten or deleted. A few find and replaces on terms like load bearing and provenance and glue words around them. The teacher runs the entire thing through a checker. What happens?
Or you have 100 paragraphs and 15 are generated, whole things scanned, what happens?
I guess the implementation and edits matter. And the “resolution “ (ie every sentence-ish bears the mark vs every paragraph). But it seems likely the signal would be lost pretty badly. And i think this is the typical sort of way people actually use AI for important things you might want to validate against.
Absolutely you can hide a watermark in a big chunk of text. But what happens when it’s inside a larger body work that gets checked or split up even a little?
Edit: im reading about synthid and see a splicing doesn’t hurt it much.
It's not clear to me that it's impossible, or even especially difficult, to make something that survives a casual LLM paraphrase. Remember, all you need to encode is a single bit of info. There's a lot of space to redundantly encode that signal.
Watermarking and fingerprinting have inherently weak security guarantees -- they rely a lot on security through obscurity, weak assumed adversaries to deliver.
There are clear trade offs between true positives, false positives and maintaining the quality of the media. It's true for audio-visual media, models and their outputs alike.
As much as I like to take shots at poor technical choices by corps and govs, this one is unjustified. Sure, inserting glyphs is bad but biased sampling is as good as it gets in 2026.
EDIT: grammar
(pardon for reposting just now from off-front-page thread on https://declaude.org/watermarking/, just thought of question)
Even if you find a way to 100% watermark any text, couldn't you just use a non-watermark model? I have a hard time believing every AI company on heart would comply.
Hell, even if every AI company on earth decide to somehow apply watermarking to their next model, they would also need to apply it to all the previous version that they commercialize. Given that anybody could make a copy of an open source model right now and would be safe forever, this seems quite the lost cause to me.
The students who, without a single neuron activating, will just copy paste questions to an AI web UI and paste the answers back.
Or the mid-tier corporate people who file 50 page "reports" that are 100% AI bullshit.
None of them will install a local model with no watermarking or bother looking up some Badonk AI service to do the same thing.
It would look like a lot of little signatures on little bits of text, and then larger signatures on a collection of those chunks once the larger chunk exists. It's not that hard. It's just a lot of signatures.
But this scheme could only ever prove that this bit of text was made by a given AI, and validate anything else ever included in the signature hasn't been tampered with. It's not hard to work up a scheme that proves (within reason) a text was generated no earlier than some date by incorporating some sort of information that could only have been known at that date so that could be validated. But this isn't even a step in the direction of proving that something was made by a human. And that's assuming the private keys stay private, which is its own tricky problem. If a private key ever leaks anything signed with it becomes invalidated.
We don't appear to be talking about cryptographic signing here. (That would never work for the problem because everyone expects unsigned text anyway.) We're talking about:
> It’s basically a text steganography problem (concealing a secret code), made more difficult because the plaintext cannot be arbitrarily manipulated.
As in, trying to force ChatGPT output to contain intentionally crafted ChatGPT-specific LLMisms that a human is unlikely to imitate, even one who reads a lot of ChatGPT output.
But how is this implemented? It's a few lines of code to implement a basic "identify and strip/replace any non printable ascii, unicode ..." or whatever. A screenshot/OCR will also do this.
SO at the end of the day you're left with some dumb rules like "you used `load-bearing` more than once per 500 words, that's AI!"
"The output of chat tools (and most of the output of AI agents) is not containerized text, but plain old regular text, and so can’t be signed. What would it even look like to sign ChatGPT outputs? There’s no artifact to pass around."
I'm reacting to the idea that "plain old regular text" can't be "signed" because they aren't "files". I'm observing that you can sign a stream of text, and probably other metadata, no problem. To my reading this really is about signing and not stegonographic watermarking, so we're in a context where for some reason the users in question want to carry the certificate of generation by AI and so the fact that this is trivially strippable isn't the issue at hand.
I read it this way because it seems to me clear that it isn't any particularly harder to do the stenographic stuff on a stream than a file (per zahlman's comment), so it only makes sense to be talking about this if we are actually talking about signing.
This is what's known as the "analogue loophole" and it's why DRM and the related "attest, strongly, at every link" schemes are ultimately destined to fail.
Lots of people suggesting we need some sort of cryptographic signature that is unique and unforgeable from every camera; it's the only way to _know_ that something wasn't generated with an LLM!
Just point the camera at a sufficiently high-resolution monitor...
I remember seeing watermarks in Stable Diffusion ages ago. We're just going to have to come to terms with the reality that AI generated anything will be indistinguishable from human generated anything soon. Counting fingers already doesn't work, most of the "tells" people think work with text are little better than reading tea-leaves. Watermarking anything is futile.
I was assuming it was something like SynthID rather than just sneaky invisible unicode but it's hard to tell from the description.
From a SynthID (or analogous) perspective, it makes a lot of sense you'd see something like this especially in something like a code comment where there's a finite amount of things to say, and it should be said in a terse manner using well defined language. The room for poetic flourish is very small, so you get... 'substrate' or 'verdict' or 'firmament' or whatever else it's using these days.
Basically 90℅ of my PR review comments now are "reword or delete this comment block, it's not helpful as-is"
It feel that today you can generally convince someone with text alone that you’re human. Beepity boopity zip zap zoopity today’s AIs aren’t this loose and derpy. Here’s a fTypo and my secret stash of dashes ——-–.
But one day even this won’t do, right?
just make an API that returns the string distance between a previously generated paragraph and the query?
that would sidestep this whole problem class.
regulators could even specify how that has to work.
what am i missing?
- watermark-free generation
- the stripping of watermarking from the output of SAAS models
Any discussion of watermarking is dead in the water in a world where we are permitted to have these things. I fear for the future.
No, I will not chill. The war on general purpose computing is gonna get real hot real soon.
If they give it to every teacher it’s the same as giving it to everyone so they’re not going to do that.
- not being locked into a provider
- not being forced to have your prompts saved by a possible competitor
- an alternative to the duolopy we quickly see forming
- offline access
I fear for the future without local models, much more than the future with them, and would rather everyone had access to a local model than be certain we catch everyone copy+pasting LLM responses. Watermarking would be cool, but it's not worth losing local for.
In the long run... well, general purpose computing is under threat anyway. Platforms with locked bootloaders and mandatory digital signatures outnumber those that don't. Even on "PCs", the openness is merely a cultural norm observed for a particular market segment[0] by Apple and Microsoft - for they are the only ones with true root keys. Even a brand new motherboard comes with Microsoft signing keys pre-flashed, and you are permitted to enroll your own only by grace. The technical infrastructure to flip a switch and lock it all down is now in place, ready to go at the stroke of a pen. Will you be able to get Llama.cpp from "the app store"? I hazard not. If you think this sounds hyperbolic, look at what is happening to Android.
[0] They don't even segment it the same! Apple segments by "does it have a keyboard", and Microsoft segments it by "does it have a native x86 processor". Apple runs different software on the same basic hardware, Microsoft runs the same basic software on different hardware, but both have created artificially restricted second class computing categories.
They store some prompts and responses, not all, that's what you're missing.
They claim to only do this when you agree to this in your personal settings. Though Google does say they will train on it, unless you disable history and only use ephemeral chats. Anthropic has a setting for it and claims not to train by default.
Also it would be vulnerable to attacks and privacy problems. You could search for substrings about some suspected information, like "John Smith's medical records show advanced cancer" etc. Of course you'd have to guess the phrasing but still.
Your privacy argument, on the other hand, makes total sense. Not even security measures like homomorphic encryption or locality-sensitive hashing could fix that basic issue.
2. Attacker asks the LLM for the opening sentences of the book, it goes into the generated responses database.
3. Later, a malicious user shows that the first few sentences of the authors book are identical to a previously generated response.
Not really, there are any number of reasons why they might do that.
In my own personal writing, sometimes I run it through LLMs to give grammar or writing suggestions.
As far as I know, most of these require user consent for retention?
Key phrase. And I'm not saying fraud in the legal liability sense. If you're not trying to hide the fact that something was LLM generated, then you have no reason to remove it. If you are trying to hide it, then there's probably a reason, i.e. you would face consequences for doing so, therefore it is fraud.