How Claude marks AI-generated content
support.claude.com
support.claude.com
I'd like to know a lot more about how that works.
A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.
I guess this may be covered by this:
> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;
Count load-bearing words using two different algorithms in a belt-and-braces fashion
Similarly, if you quote someone word-for-word, you wouldn't anticipate their words to be flagged as Claude content, but if someone memorized Claude output word-for-word. That would still be classified as a Claude output.
Going forward you could categorize the influence of Claude on a population based off a percentage match between their spoken words with the LLM prose.
If a detector flags something in your text and you've properly attributed that text to another author then what is there to worry about?
If you are implying that this Claude watermark may be good enough to enter the "evidence to fail a student" category then we have a major disagreement.
The mechanism seems to survive editing. The extreme probabilities get a little less extreme, but are still extreme enough to be distinctive.
But it wouldn't survive paraphrasing, because the output would be entirely human and the token correlations would disappear.
It might not survive referencing if only a sentence or two is used.
The practical issue is how true the claims are. It's one thing to create a proof of concept, another to see how it works in use.
And this is potentially catastrophic for code, because the grammar and word choices of code are completely different and more fragile than standard English.
public abstract class BaseAnimalBeanFactoryGeneratedFromClaudeFactory
My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.
Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.
So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).
Odd variable naming? Stylistic choices that are watermarked?
Or as someone else noted further down in the comments, it could be more subtle:
Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.
Whatever it is, I'm sure it's load-bearing.
I might be imagining things of course. But comments would be great fit for this use case.
edit: as an aside - I actually use extensions to collapse comments and change the color to be less intrusive.
If it works like people describe - on the every nth token or something - then the mark will be left in the chain of thought and discussion with the model - not in the code artifacts.
There isn't any effect on the quality or precision of the output. Nothing changes in practice.
Seems like this would only catch the most unsophisticated cases.
You are just pointing out the very visible single cases. But the mass of low-visible content is much higher and more dangerous (like propaganda bot-farms). If a social media platform adds watermarker checks the bots would implement bypasses ASAP.
You can hide data in that randomness without impacting the quality of the response by using a sufficiently "random looking" pseudorandom bit stream instead of real random numbers.
I previously worked on a project to do that here: https://github.com/shawnz/textcoder
A small anecdotal example to demonstrate how the response lacks depth. It's imperceptible to most. You'll probably see through it in your own field of expertise, if you had both answers. Uncanny valley type thing.
Wouldn't you need the prompt to know the probability of the next token?
There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example).
Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxiliary words/chunks, you should, in principle, be able to determine what other wordings could have gone there instead. From this, you can, very roughly, recreate the token probability distribution that was in effect when those tokens were generated. With the probability distribution in hand for enough chunks of text, you can start inferring properties about the RNG process that was used to sample from those distributions.
But then, this notion of "load bearing" vs "auxiliary" can be expressed directly in the probability distributions. A load bearing token just has a very high probability, and thus any RNG bias that may have been in effect will likely be swallowed in the distribution. So the parts of the text that are highly dependant on the prompt will naturally not be contributing much information about he RNG in the first place.
By the way: https://x.com/alexcdot/status/2087078010524406137
"Neutralize engine is temporarily unavailable. Try again."
Well, they should have run their own AI slop website through their tool...
From their before/after:
- Certainly! -> (removed)
- onboarding redesign -> revamping the introductory process
- this week -> (removed)
- empty states -> empty sections
- CTA heirarchy -> call-to-action sequence
- interviews -> discussions
- aligned copy with brand voice -> verbal identity
... these choices change the meaning of the texthttps://www.pcmag.com/news/genius-we-caught-google-red-hande...
> total = calculate(items)
> result = calculate(items)
> value = calculate(items)
> amount = calculate(items)
All of them are reasonable options. If we bias the model's output so that one of them is more likely than the others, then we can reconstruct that watermark if enough of these frames are present.
You can carefully select which pseudorandom number generator (prng) you use to be able to id text of a certain length. I expect there is some performance characteristic you have to manage since you're doing this on every inference, but once you do that it doesn't change the output in any meaningful way (the prng is still a statistically valid prng, it just happens to let you check if the output used that prng)
The point is that you can do this simply by swapping to a different RNG, which isn't noticeable to the end user, and while it changes the output, it's not any different from how using a different seed or being lumped in a different batch will change the output.
^ excerpt:
> So then to watermark, instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI. That won’t make any detectable difference to the end user, assuming the end user can’t distinguish the pseudorandom numbers from truly random ones. But now you can choose a pseudorandom function that secretly biases a certain score—a sum over a certain function g evaluated at each n-gram (sequence of n consecutive tokens), for some small n—which score you can also compute if you know the key for this pseudorandom function.
The LLM presumably generates f(input, RNG) but we only can observe f(RNG).
... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches").
Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.
there are many simpler methods, for example you can have a tiny windowed transformer operating on the output text and all you do is alter certain words (that don't change meanings) to maximize its surprise. the tiny language model will have a special training regime to build up a somewhat unique view of the language.
we are talking about a 0.5 bit watermark here (existence). I would have zero confidence in being able to reliably remove such a watermark from pretty much any medium.
Short version: it's likely a SynthID style statistical watermarking scheme that Anthropic is using. There's likely no hidden unicode and no large stable green/red scheme.
Longer version: I found some really statistically significant outcomes when using Gloaguen et al.’s (https://www.sri.inf.ethz.ch/blog/probingsynthid) detection algorithm for SynthID, though unfortunately it seems like the models might be pre-biased to pass that test already. The difficult thing is that it's hard to conclusively know which models definitely already have watermarking enabled and if there are any models that don't have watermarking, which means that being able to get a highly positive statistical result is much harder.
You can take a look at the full writeup, but many of the clues from Anthropic's blog post, plus negative results on many of the simpler tests, point to them using such a scheme. Of course, it's possible that they came up with a novel new scheme/algorithm, but amongst the available literature, the SynthID family of models is probably the best bet.
What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?
The way I use LLMs (and I'd advice everyone to do the same) there really isn't, the agent implements things exactly how I want them, or I use the agent to massage it into the exact bit-by-bit version I imagined when I first sent the prompt afterwards. I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually, although I know it's a popular approach taken by many.
> And of course you can add arbitrary comments; my Claude-generated code is very verbose.
So watermarking for all users who allow code comments from agents, no watermarking for us who force the agents to never write a single code comment? Alright, I'd be fine with that.
And the reason to let Claude make worse code than a professional would by hand is basically suppressed demand. Since programmers are expensive, previously code mostly got written when a large number of dollars were on the line, or when an individual programmer did something not economically optimum (e.g., hobby project).
That left a whole lot of somewhat less valuable software unwritten. It's the economic space that no-code tools have been nibbling on for years. One way to think of things like Claude Code is as effectively no-code tools. Pre-LLM no-code tools would produce data structures that got executed by special environments without ever being seen or tuned by a human. Claude Code can be used just like that, with text as the input and python as the intermediate representation that nobody ever looks at.
That approach probably isn't sustainable for what we professional programmers would call a serious project. Claude can easily get in over its head and I expect that its code decays over time, in a fashion similar to how many human teams get in a state where they just have to rewrite everything. But faster, I'd expect.
But there are a lot of unserious projects that previously would have never been created. E.g., a quick app to manage your little league team, or a bit of in-house business stuff in the "a little hard to do with a spreadsheet" range.
You never generate throwaway code used to test an external service? or try out an interface idea? There's a lot of code that's only meant to be ran once. I often dont even care what language it's written in.
And save/persist it? No, most of any experimental stuff goes into /tmp which gets cleared out on reboot, nothing I care to save in any repository. Or just "show me how this would look like" and then it's only in the session itself (and the logs/state I suppose, technically...).
In cases where one token is extremely likely, it'll randomly be red or green and still be picked in either case as it is simply the best (or only) option. So you'll have more tokens that don't show a pattern either way (half of these cases will match and half won't, just the same as if a human wrote it). Meaning you'll need more instances where multiple tokens were all likely to see if there is a pattern. Given the check algorithm can't identify these cases, it can only judge on the overall text, so the more strict a language, the more the length requirement scales.
Where I wonder if this keeps working is in tool calls. Often, you don't take code straight from the llm, you take the results of a tool call to edit already existing code. It might be that the result of this leads to far too few signals to pick up, meaning that this only works when one does significant generation with a single model (even swapping between different models, at least by different companies, breaks this just as much as having a human write parts of the code).
Think of it like finding a loaded dice. A dice that has a slight bias in a few dozen roles is just random chance. If that bias continues after hundreds of thousands of roles, the dice is loaded. But will a code base have enough samples, especially when edits made from tool calls? I could see this being unable to detect things at the size of a reasonable PR and only being useful for massive sets of changes and only if the person behind them didn't structure their AI usage to avoid detection.
You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible.
The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).
Paper:
Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.
Note, there are many ways to represent words visually on computers that look identical
Theres no way that’s what they are doing.
There is no room for false positive here in the same way you can't randomly find a collision in a hash function if it's strong enough. Like the rate is so infinitesimal that it is effectively zero.
Now replace random number sequence with prompted string of words. And instead of using the PRNG on every word I use it every n words. If the generated text is sufficiently long I can tell by matching the expected deterministic pattern.
You can defeat it by changing the words yourself and triggering a false negative but there isn't really any room for a false positive if the text is long enough and the pattern matches perfectly. If the pattern doesn't match then I can compute a probability.
Here's the strawman: The text-based watermarking is going to be done procedurally instead of generatively. Maybe they add some sequence of zero-width Unicode characters to all generated text at certain intervals. Then, there is effectively no false positive possible (because humans would [effectively] never type such sequences of unicode naturally). It may survive some editing (depending on how you select/edit the characters), and it's possible to be stripped (false negatives).
Keen to see if they are doing something SynthID-esque?
This is scripture homeopathy and it's irresponsible.
Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
if they used older texts as training data, to some extent pangram would just be an age classifier for writing style.
https://www.pangram.com/research/model-card/pangram-4
> Pangram 4 achieves a 0.0041% false positive rate (roughly 1 in 24,000) on 1,000,000 human-written English FineWeb evaluation examples
> Overall False Negative Rate is 0.3396% on English AI generations (26 generator models)
Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks.
Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.
I recently heard someone say "that's genuinely the exact solution I was looking for" and had to do a double take.
Are you talking about pieces that were fully human-written with zero AI editing/rewriting etc? If so, what makes you think that false positives will happen there? They aren't looking for "writing styles" or emdashes etc. They are using watermarks and metadata.
If you're talking about people using AI to copy-edit text they manually wrote, this was explicitly called out in the article:
> A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example: Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source; The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.
They could be doing invisible and vaguely-harmless Unicode stuff. Insertion of zero-width joiners and non-joiners, replacement of regular spaces with non-breaking spaces, building spaces from multiple hairline spaces, intentional use of non-NFC-normalized codepoint sequences for accented characters, etc.
Text with all this junk in it still reads the same; it just might wrap a little strangely, or not byte-match / collate correctly in a database (and Anthropic has never made a guarantee that their models would be capable of emitting text with these properties, so that’s fine.)
And, importantly, no regular text or document editor would insert these things (especially in the useless places you could insert them for watermarking.) You only really see them in text that’s been explicitly typeset for a specific layout (e.g. in text-containing SVGs, website mastheads, or game HUDs) or for print publication.
Of course, if this is the technique they end up using, then it’s very simple to strip it out by canonicalizing the text (i.e. Unicode-normalizing it + stripping out invisible layout characters + replacing “weird spaces” with regular ones, etc. Essentially the same thing many sites already do to user-generated content to prevent users from using Unicode features to break the page’s layout.
[0]: https://deepwalker.xyz/blog/evaluating-synthid-watermark-rob...
I think in this case I think it's some kind of cryptographic signature smeared across the token IDs, so I don't think the risk is very high.
If you find yourself getting to be afflicted by this "brainrot", be sure to go outside and take a moment to ponder what's around you. The grass is there and will be there long after we are all gone. Consider this for a moment as your organic thought processing unit starts to slowly munch away at its internal context window.
I don’t know what the answer but I absolutely know it isn’t this.
It would also require the individual humans you are trying to control to get on board otherwise the analog hole breaks the chain, absent mindboggling levels of physical surveillance on top of the the total monitoring of all electronic data flows that this idea requires.
TLDR: AI will force global online digital ID for everyone that uses the internet for the exact reason you mentioned. And that would forever change free speech forever allowing the powers that be to put the genie "back in the bottle" so to speak.
link: Raiden Warned About AI Censorship - https://youtu.be/-gGLvg0n-uY
This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem.
But ultimately, when you are powerless and can't afford to do the fighting: I'm convinced the only way to protect yourself is to be very mindful about your writing style, and to deliberately corrupt the language through objectively wrong "stylistic elements".
The bias is different for each position and follows a defined RNG, seeded somehow predictably.
Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.
How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).
So, an NG?
If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?
Keep in mind that the LLM "sees" the previous (tweaked) output and picks what makes sense based on that. There are few situations where a perturbation like that would be unrecoverable, and I assume these situations also correspond to a huge probability difference between the most likely completion and the second most likely one - in which case, the watermarking algorithm can choose not to touch the token.
Gotta be hard to tune that.
Now the goal is either "identify the meaningless interesting bits and swap them out with 0% loss in the direction of the original goal," or "perturb some small selection of the output towards my secondary secret goal of watermarking the text."
It would be quite impressive if they managed to identify with 100% accuracy the tokens that "don't matter" and are free to swap with whatever signalling tokens encode the AI scarlet letter, but most likely they are not 100% accurate, and that means the output is worse off than without the watermarking logic.
Oh, how I laughed. That was never your code, my friend.
Why were you doing that before watermarking?
Same answer.
I think the solution is assume everything is ai generated unless told otherwise and rely on authorship/brand as a sign of quality.
If you're putting the work you say in, the result won't be obviously distinguishable. Obviously, from some of the things that get posted here, that last sentence is too much for most people to bother adding to their prompt.
In the meantime, it is true that this takes something away from you. But it's something you were only recently given. Now you're not given quite as much, but readers are given a little more (or rather, there's less being taken from us!)
> This is incredibly different than pure ai text.
Ok. But it's still incredibly different from pure human text. I guess the question is which provides more value? Providing the information "this text is AI watermarked" to readers? Or allowing creators to lie and claim that AI processed text was 100% human generated? I agree that people assuming that "has AI watermark" == "is AI slop" is incorrect and causes some amount of harm, but having the watermarks also pushes back on a large amount of harm already being done.
(Personally, I'm skeptical that these watermarks will ever hold up to adversarial attacks, and they haven't claimed that they will. So I think it's the usual "casual liars will be caught, determined liars will get an additional thin veneer of respectability".)
I was referring to a hypothetical future where everything has gone through AI to some degree, and for the purpose of the argument it has been modified enough to embed a watermark. Thus, the "detector" just always returns "true" in practice, and therefore nobody cares about what it says anymore.
I think what you call "aura" in your essay is already very much a thing. We use the adjective "artisanal", for example. Artisanal, bespoke, heirloom. hand-crafted, small batch, ... we have many ways of trying to convey "aura". (Heck, even things like "minority-owned" or "woman-owned" are attempts to imbue an aura.)
But more to the point of text, I'm also fascinated by the evolution of taste in this area. People's ability to detect AI slop improves, triggers revulsion to things they were previously fine with, and then sometimes they'll go back to being neutral once they start using AI more heavily themselves. Some AI writing is noticeably better quality than the human written version, yet it still smells of slop, and I at least am sometimes torn about which I prefer. Then there's the trend of AI writing getting better, substantially better, and sometimes at the same time worse in certain ways. Not to mention students, many of whom are now much more accustomed to AI writing than human writing -- it's their mental model of what an essay looks like, especially when all of their own essays uncoincidentally read like AI writing. And then there's all the nuance of the sometimes subtle differences between a wholly AI-generated piece of writing vs a heavily human-directed piece of AI writing vs an AI-edited piece of human writing vs a translated piece of human writing. It's fascinating and scary.
I too am curious about the final question in your essay: "will such categorization lead to the development of a potential human premium?"
I understand you're saying since you "worked with it", it is not ai generated but if you still use the final output verbatim, the writing itself is LLM generated purely.
You want to share the output by it but also position it as not ai output. But that's fundamentally dishonest.
Furthermore, if you think your approach actually creates value and can be judged on its merit, why not disclose its ai written? If you think that will make people think your content is bad then you should see that as feedback and maybe not use AI since readers don't like it.
In academia they've got their own concerns of 'purity,' (not least of which is justifying their continued existence which is in my opinion hard to do) and they are the ones who are going to want most strongly to punish anyone who uses AI.
And perhaps "journalists," who will want to trumpet the latest government's press release being [what they'll portray as] mostly AI-generated, as a headline-grabbing "gotcha." Ironically, that may even be a story that'll be written by a fully-autonomous journalist "agent" in a newsroom that's been pruned of all human journalists!
But in business it seems to me that we're all agreeing that it's a "good" use of AI to write in that way.
Individuals maybe, companies won't and that's where most of money is at.
It’s just not a reasonable ask.
Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.
It helps detect low effort slop.
---
Caveat. If you write your own creative work and send it to Claude for "cleaning up grammar", it might insert the watermark.
It seems it would get as simple as:
outputText = promptLLM(prompt)
scrubbedText = scrubWatermark(outputText)
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.One could even say, the mark is load-bearing.
There just isn’t enough information in plain text to do this and we should stop pretending there is.
If we need to verify something isn’t made with ai then we need other ways of doing so - eg looking at a document edit history, doing it as an exam, oral defense.
There are options! But pretending you can tell if text is ai will only catch out people who make no effort to hide it and will inevitably have false positives.
It’s not a solvable problem.
That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.
If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it.
In a world with many different competing models, the risk of losing customers to other providers over this is much more real.
Maybe they've looked at the numbers and the portion of people who clearly use Claude to cheat on examples etc is so tiny that losing them to other providers isn't a problem?
I guess whoever is the policy maker is assuming that some protection is better than none and that most people will not reach for such tools.
Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?
So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.
you can then consistently like figure out if it was claude that wrote the sentence. it is easy as you noted if you just get another ai to read it and then rewrite it.
At first I thought this approach was just the "LLM flavour" of writing, but it's way more subtle, especially as the bias is applied uniquely for each token position.
Signatures are extra data, added out-of-band to the existing data. Out of band data, by definition, is easily detected and stripped, so the utility of a signature is that authoring one requires secret knowledge. Philosophically, the presence of a signature is a kind of authentication, a desirable thing that is hard to grant and easy to revoke (the smallest change to the data renders it invalid).
Now, watermarks: if you flip it round and say you want to glue on a piece of undesirable data - something that represents disauthentication, like a cursed black spot of written-by-LLM - then you want it to resist removal efforts. And now right away you have a hard problem because your sticky data must be in band, or else it is trivially stripped. Not only that, in fact, it has to look enough like real signal that it isn't easily filtered. And on top of that, you can't distort the real signal too much, or people will complain. So you're cornered into doing a kind of steganography - hiding small amounts of information in the entropy, biasing the signal in perceptually plausible ways that are detectable to those in the know. Cartographers add fake streets ("trap streets") to catch plaigiarists - for LLMs, watermarking might take the form of subtly odd word choices.
Whether that's right or wrong, I'll leave to you, but there's huge differences in perspectives, and if you only get your news from Western sources and communities (and companies), you're in a bubble too. A different bubble, and arguably a more porous one, but still a bubble.
For a taste of where I think things are headed, try asking Chinese models about Tiananmen [1]. And then take a look at the Chinese government's approach to pretty much anything that they think reduces security or social harmony. I find it hard to believe their models will be the one exception to that over the long term.
[1] https://en.wikipedia.org/wiki/1989_Tiananmen_Square_protests...
Relevant comment from a few days ago:
I guess it's a bit weird. An LLM can recite copyrighted material verbatim from training data (they had to work hard to get them to stop doing that, and they haven't been entirely successful). But LLM outputs are public domain. (Except when they are a verbatim reproduction of a copyrighted work.) I'm not sure where you draw that line when it's not verbatim...
For a repo, I don't know what that means. Only the AI generated lines are public domain?
This seems to be similar in execution to Google's SynthID. I hope they release actual code the technically proficient can use, unlike SynthID which can only (afaik) be queried with Gemini's UI.
Several open-source projects have already proven SynthID to be ineffective.
If I heavily edit LLM output, will this still hold the watermark?
You really can't make this stuff up, it doesn't make any sense.
What happens is that AI selects similar words based on a random process.
Something like "The company had a large/big/substantial advantage".
It chooses between these words, and over a longer piece of text, the pattern will start showing, like a "choice A → choice C → choice C → choice B → choice A".
The normal-looking text will actually be a fingerprint living in the form of statistics.
I think Claude will be sharing these patterns to third parties for AI detection.
Even if all models are mandated by EU to do their own watermarking, it doesn't take a new frontier model to be capable of paraphrasing it, so you can use 2026's open models to do that paraphrasing, far into the future.
The funny part is that a lot of people have already developed an impressive ear for spotting AI-isms, so for now I'm not even sure how important it is to have this. No technique can be 100% guaranteed accurate anyway, and humans are pretty good at recognizing AI already.
One example I've seen are junior employees at my company deliberately adopting a lowercase/less punctuation writing style so as to stand apart from AI.
>Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;
Such models already struggle not making any unnecessary or unwanted changes to a corpus, this makes them unable to by design.
Then at least you could have two weak-postive signals, and a strong-negative signal. (Though one that only fits precise chunks of tokens) I'm sure I'm missing something here, but my groggy morning brain thinks that doesn't seem too bad.
This should make it easier to catch cheaters who use Claude, right? Unless everyone runs their artifacts through some watermark and metadata sanitizer?
> Regions. Marking will apply to output from supported models wherever Claude is offered, worldwide.
It will happen if Claude tampers the text. Guaranteed.
So great that Anthropic is doing something about this, but it's not clear what their watermark exactly is. How do I, as a user running into some content online, know that it's generated by Claude? What is their watermark?
It sounds to me like they create the pattern in the regular text of the content, which sounds interesting, but also odd, unreliable, and may limit the content you can get out of Claude. Will it subtle change the words in order to hide this pattern in it? I don't know if that's something anyone wants.
I'm assuming that's exactly how it works. How else could it?
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing. Watermarking will be applied at the model level, which means it will be present no matter which Claude product or surface the text comes from."
That's probably an over simplification. Also a solid defence that can be used against complaints about the way AI writes text.
- "Ensure distribution of vowels is in >99th percentile of human work"
- "Ensure the distribution of the letter "s" is within 99th percentile of human work"
- "Ensure the distribution of the letter "L" is periodic with periodicity within 5% of 1/N characters.
- "Ensure there is a cross-linguistic 'typo' (colour vs color) at 1/N words, where N: 1000 = Model1, 2000 = Model2, 3000 = Model3.
- "Ensure the distribution of tense error is within 99th percentile of human work"
If more than 3 dimensions have a score >99% percentile of human, let's call it watermarked...
Yes they can do this, but it's more likely closer to the original "red token, green token" paper: https://arxiv.org/abs/2301.10226
i.e. take half of your LLMs vocabulary, and upweight its probabilities by ~55% to the other half's ~45%, and scan for overuse of this half of all tokens. You can even choose a different half/slice for every individual user, for every individual action. You can implement this under the hood cheaply with logit-biasing.
I suspect they'll roll out the watermark everywhere.
Computerphile on YT has a video explaining how models can fingerprint the text they produce. Essentially they modify the probabilities of word choice slightly in a predictable way.
This would imply that a positive watermark signal is likely (but not guaranteed) to be AI generated. Also implies that a negative watermark signal is not necessarily void of AI generated text. This would create a problem if people start to trust the watermark as a heuristic, as the ability to critically evaluate the text is replaced by the search for a watermark.
Seems to me all of this is really trying to solve for "is this text bullshit" or not, which would require a different solution.
>A file’s metadata was stripped through format conversion, re-saving, screenshots, or other means
Ah. So what essentially every single consumer-oriented media host does. Gotcha.
I fully recognise this is a hard problem, but hopefully metadata isn't the only method for media. Standard procedure is to shrink files for storage and privacy reasons, and non-visual metadata goes out the window by default.
I propose we defeat this with the obvious: Simply, figure out what are some of the markers Claude and others will use for these tools, and sprinkle them randomly on everything we type or produce, all the time, 100%. If users flood the tools, and everything returns as AI-generated, then the tools become useless.
So, they’ve been doing this for over a week without telling anyone?
AI stigmatization is ableism.
What I was trying to point to was that "this thing helps some people" does not equal "this thing is unequivocally Good and should be entirely unchecked".
I don't even care about the AI. I just get peeved by bad lines of argumentation.
You quoted a Futurama caricature that is not comparable to all this, apparently because the words "assistive technology" were used.
In the literal sense that a gun is not a chatbot? True.
In the sense that both your post and my analogy use the argument "this is assistive technology" to defend something which only in a very narrow slice of their thing-ness actually is assistive and in the whole rest of usages are much more, and not only Good, which both you and my "futurama caricature" were willfully ignoring? I think it's quite illustrative.
I am cancelling my Claude max 5x subscription and moving to ChatGPT pro. I have difficulty enough trying to ensure my meaning comes through correctly, along with everything else; to now have to look out for/analyse watermarks too?
I feel shamed enough by society, thanks Anthropic.
Both points suggest your subscription support was well chosen before.
Unless with "proofreading" you actually mean having the LLM write your content for you.
What am I missing?
Could we get an Anthropic subscription for Claude Code with data residency in the EU, so we don't get robbed blind by AWS Bedrock et al., but can have a monthly subscription like with the regular US option?
Anthropic has this strong repulsive effect in the way they operate, I wonder if they'll be around for long, it's hard to say at this time. The competion is fierce, so there isn't much room for shenanigans at this stage.
Anyway, I've found a magic line that can be copied & pasted to the comment sections of most OpenAI/Anthropic news threads. This one is no difference.
The magic line:
> Doesn't matter; have DeepSeek.
It’s foolproof, I tells ya.
Why do feel so entitled to being able to pass LLM-generated text as our own? I get that a lot of techies aren't good at writing. I also see a lot of tech hustlers who like to use LLMs to fake human connection and compassion - I've gotten LLM-generated recruiting emails that talked at length about how the recruiter "valued" my work. Just because we found a "cheat" button doesn't mean it's wrong for others to want to know.
Yes, LLMs are great. So is transparency. If you think an LLM writing is your new superpower, wear that badge with pride. It might mean you will lose some business from LLM haters and win some other business from like-minded customers. C'est la vie.
One difference perhaps is that you think using LLMs is cheating, while others do not.
"Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed."
> Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, ...
My intuition is that this would be very possible, in a way that makes false positives so unlikely as to be virtually nonexistent (at a certain fragment length.) Basically all you would be trying to do is to defeat people who would deliberately screw up the signal below the fragment length, and you would try to get that fragment length to at least the size that intentional obscuring of the signal would be obvious. I could see it being possible to detect even from non-contiguous fragments interspersed with noise.
It's just 1 bit, and you don't really care if a sentence or two is slop. I'd be surprised if a PhD interested in steganography couldn't come up with a good scheme in a week. It's a QR code.
What would be scary is if they could come up with a way to detect advice from Claude i.e. you get Claude to review your work as an editor, read the output, then as a result make non-verbatim changes, and that signal still gets through. If you could do that, you could do things like tell if a pundit speaking on television has read a particular Wikipedia page. Seems impossible, but LLMs seemed impossible.
edit: there are so many unimportant language choices; ones that are even hallmarks of AI use already, like the fact that it generally picks the mode. Not always picking the mode or picking at precise distances from the mode could hide signals without significantly affecting the quality of the content.
For a company that is too sloppy to properly contain it's ai and prevevent "it" from hacking other services, how could we possibly trust them to not fuck this up? There's even less incentive for them here to not get it wrong, so getting it wrong is in my expectation a given.
This is dismal.
>These experiments provide empirical evidence that more advanced LLMs can lead to smaller TV distances. Thus, based on Theorem 1, reliable AI text detection would become increasingly difficult
EU regulation does it again!
I would prefer to know if given content was generated with LLM. This is information, and information should be free.
Can it be circumvented? Of course. Will most people go through the trouble to circumvent it? No.
U+2800 or U+3164 would be nice.
But as I remove unwanted characters with grep before layout in InDesign, someone will make a skill for removing such space characters.
Hell, it'll probably happen no matter how sophisticated their watermark is. There's no watermark in text that can't be detected and removed, and no text that can't be converted to generic keyboard ASCII.
No Anthropic model has been launched in August.
They should make it easier, to detect slop so we can ignore it quickly.
I hope Pangram makes an API or an extension to analyze a page to detect slop on a page and then closes the tab immediately.
Nobody should be wasting time on garbage LLM output in code, text, image and videos.
Is their research also a scam too?
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
https://www.pangram.com/blog/pangram-4-technical
If so, what is the best one out there other than Pangram then?
https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...
What about on Pangram 4?
Today, the best way is probably Pangram. Tomorrow, it might not be, especially if they try to push their recall up.
You might have to make peace with the fact that there may not always be a tool that does what you want.
Thanks!
> But the burden of proof is on them...
I mean is this enough proof?
https://www.pangram.com/blog/pangram-4-technical
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
Or is this marketing, a public stunt or not real research?
I think this is enough for me to know they are actually improving their AI slop detector.
In my own testing, Pangram is excellent at detecting the default output styles of LLMs.
If you tell the LLM to change its output style, so it’s not full of “load-bearing spaced em dashes that aren’t X, they aren’t Y. they’re Z.” constructions (which humans are pretty good at detecting on their own), the false negative rate soars.
I wrote this three years ago:
I feel like this is FUD. If you copy text from Claude, Ctrl Shift V it into VS Code, the IDE will give up the ghost on if weird characters are in there. And it's not like Google suddenly invented new letters or fonts either.
Practically speaking, I feel this is Google publishing misinformation.
First, the article doesn't talk about adversarial usage. As in, it's not claiming to be proof against various techniques of watermark removal (inserting words, rewriting with a different model, manual paraphrasing whether minor or extensive, etc.) It might handle some things and not others, but "I could trivially defeat this!" is not a gotcha; they haven't made that claim.
Second, basic information theory tells you a lot about what is or isn't possible. Watermarking is information. You need degrees of freedom to store that information. You can even estimate various sources of space in bits (often fractional bits.) To a first approximation, longer text has more bits of space. Language matters -- a rich (aka messy) language with lots of potential synonyms has more space. That goes for human language as well as the difference between human and programming languages. (Most programming languages have much less flexibility to them than most human languages.)
The details of what space you make use of are interesting, but speculative. In the English sentence "Ellie spat in his eye", you could look at it at a word level and say that swapping "Mary" for "Ellie" is a lot more damaging to the meaning than swapping "face" for "eye", so there are more bits of freedom in the latter. For coding, `for (int i = start(); i < end(); i++)` probably shouldn't swap `<=` in for `<`, but it could be written as `int i = start(); while (i < end()) { ...; i++; }`. (I'm not claiming this is the sort of alternative that they'd use, just an illustration of what's possible.) But there are a lot of possible places to find these bits if you look at large chunks of text. Different ones are more or less resistant to accidental or intentional information destruction, and require less or more sophistication (aka brittleness) to be extracted. (In the limit, you could require the full original prompt and encode tons of stuff by tweaking the logit selection. But it wouldn't be very useful to require the original prompt.)
Also, does this degrade model output? Yes. It reduces the bits of freedom available to the model for producing the signal. Does that degradation matter in practice? That's totally dependent on exactly what is happening, and will likely change over time and across different purposes. I hope we're past the point where people believe that setting temperature to zero produces "perfect" output in some sense. (Or should I say flawlesslesslesslesslessless output?) It used to be useful for reproducibility, at least, but my understanding is that it's no longer even good for that? Anyway, reproducibility != quality.
There are a lot of things that could be going on here. The article doesn't claim very much, just that they're encoding a signal in the output that can be extracted later. How robust the signal is in terms of the FP/FN rates is unknown. The resilience (resistance to destruction) is unknown. The impact on the output quality is unknown. Even the question of whether this will make AI slop less sloppy is unknown; maybe this means we'll see a little less exact repetition of "I have the whole picture now" and instead it'll sometimes be "Now I see the entire picture"? Can we dare to hope for an occasional "Ok, this time I got it, boss"? That would be a (very minor) quality improvement.
And thinking out loud, they could be really horrible and embed by using unicode characters instead of ascii, which would give a lot of flexibility, but would make the result almost unusable (but easy to defeat).
Models from openAI had instructions in their system prompt not to talk about gremlins and goblins. Anthropic got caught acting different if you're Chinese. Grok got its Nazi dialed turned to 11 until it started calling itself Mechahitler cause Musk found it too left leaning on Twitter.
It doesn't matter where the content comes from, only the quality/usefulness matters. If you are opposed to this idea, the next decades are going to be very tough for you :)
> It doesn't matter where the content comes from
It absolutely matters to a lot of people. Things like these just provide transparency and allow people to have the necessary information to make their own decisions.
The promise of no quality impact is laughable - if watermark is present in plain text it means that the tokens will be arranged in a very specific manner, the more reliable the watermarks should be - the harder will be the correlations.
Don't forget how annoyingly bad Anthropic products have become in recent releases - low adherence, annoying alignment, annoying guardrail false-positives, unwarranted checkpoints - all that shit. Now they deliver more crap.