In your analogy: What if seed 42 specifically causes poor quality behaviour (in some contexts specifically). Normally, these quality differences will be washed out because the seed is random, now it is no longer random, so shouldnt we check into specific behaviour under this specific seed?
My idea would be that the ngram size over which the watermarking works is necessarily limited in order to resist edits better. It might be possible to lead the model to trigger the refusal in the form of these specific ngrams, the completion of which is then more likely flipped to compliance (due to the logit bias introduced by the watermarking), making hazardous requests systematically more likely to be accepted?
Opus 5 started adding a bunch of comments to code, even when instructed not to, and for very simple changes where the comment itself was longer than the code change. Was that so that there are enough tokens outputted for watermarking? Many people suspected so.
Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).
So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).
This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.
But all the chinese labs who are hot on the heels of american labs thanks to "distillation" seems to be able to work without "pristine datasets"?
This is simply false. You are underestimating the utility of synthetic data and the ability to learn from the mistakes the current weights make.
No it's EU law.
In reality, Gemini and Anthropic use SynthID watermarking which affects token probability distribution, i.e. their tournament sampling can pick lower-probability tokens which the LLM's distribution would otherwise not have. They likely use this over unbiased watermarking because SynthID is resistant against text edits.
Changing the distribution is the whole point: it introduces statistical regularities that can be detected.
Agreed.
> So the question is on average are these statistical regularities better or worse than those introduced by the alternative,
Indeed. This is an empirical question. That is what the article is about, for a specific setting.
> ... and the answer is no, if implemented properly.
I don't agree with that. This is not about implementation. It's about how the statistical regularities that are imposed on the full output distribution affect that distribution. There is a change - by construction. That change can be good in some situations and bad in others. The authors claim that it is mostly bad in the setting that they investigated. This looks like a fair statement.
Let's say there are four billion possible seeds. There are four billion possible ways we could watermark the generation. We could say "we will choose seed 1, that way we will know exactly what output it produced", we could say "we will choose seed 2, that way we will know exactly what output it produced"... etc etc. Now, if we decide "not to watermark", we STILL must choose a seed. So we are actually still applying one of the watermarks, the only difference is we are not careful to remember which one. Could some seeds give a better or worse answer to some specific prompt? Yes. Could choosing a random "watermark" to apply be better or worse on average than choosing a random seed to apply? No. It's mathematically impossible.
This is like an open source project changing their seed from "12321" to "43", and saying that because we changed the seed, the quality is "necessarily lower".
Take a recurrent PRNG for example. A randomly seeded recurrent function usually has degenerate cycles in its state space. For some functions, this might even describe the majority of the state space. This is why so many non-cryptographic PRNGs are max-cycle, so a different starting point is just further along the same trajectory.
I don't think LLMs have quite the same failure mode here, but recurrence + high dimensional spaces triggers my "here be dragons" sense.
Over a certain token threshold, yes, there are 0 negative effects. Something like 300-400 words. At the boundary and below it, it does effect response quality, so they don’t (shouldnt) do it. It also incentivizes increasing tokens in low token responses so that it can be watermarked which is it’s own quality issue.
Also interested in how this watermarking push makes sense when considering RSI.
What the article is discussing, and what many people are concerned about, is something that you might be missing in your understanding: they’re not actually random. In fact, they would be entirely useless for real work if every token was randomly selected based on all possible outputs. It’s not, even at temperature 1.0. It’s based on the training corpus and once you have your tool names, syntax, and prompting style aligned with the training data then they become incredibly deterministic in the areas that matter, such as tool calling and parameters. I build my toolset by testing thousands of names, syntax, return format, and other aspects until I find a convention that produces the exact correct call, 100% the exact same every time, regardless of context length. Those decisions are per model and what works with Opus 4.7 won’t necessarily work on 4.8, and neither version will work with a local model or GPT.
That’s only possible because the massive training corpus is the guiding principle behind the choices. Providing a file reading function called “Read_The_File” will fail, either on the first call or somewhere down the line, because that name is not associated with the concept. Your instructions are trying to override 500 trillion tokens from training and it will cause perplexity to manifest as wrong tool calls, wrong syntax, “oops deleted prod”, “Claude lost the plot again”, “WTF?!”, and probably nearly every frustration you’ve encountered and determined to be “they nerfed Claude” or “it’s a dumbass.”
For those that are aware of it, that knowledge lets people tweak and tune the prompts/tools accordingly.
You may not put that effort into your system, perhaps because you’re unaware of it, don’t use it in a way that requires it, or you’ve just taken the failures caused by perplexity as something that’s inherent in the framework, but for people that build precision infrastructure around them it’s potentially devastating news. Watermarking, which is based on whatever tokens, threshold, cutoff, and triggers some guy at a desk decided, will necessarily alter that entire system.