Let's say there are four billion possible seeds. There are four billion possible ways we could watermark the generation. We could say "we will choose seed 1, that way we will know exactly what output it produced", we could say "we will choose seed 2, that way we will know exactly what output it produced"... etc etc. Now, if we decide "not to watermark", we STILL must choose a seed. So we are actually still applying one of the watermarks, the only difference is we are not careful to remember which one. Could some seeds give a better or worse answer to some specific prompt? Yes. Could choosing a random "watermark" to apply be better or worse on average than choosing a random seed to apply? No. It's mathematically impossible.
This is like an open source project changing their seed from "12321" to "43", and saying that because we changed the seed, the quality is "necessarily lower".
In reality, Gemini and Anthropic use SynthID watermarking which affects token probability distribution, i.e. their tournament sampling can pick lower-probability tokens which the LLM's distribution would otherwise not have. They likely use this over unbiased watermarking because SynthID is resistant against text edits.
Changing the distribution is the whole point: it introduces statistical regularities that can be detected.
Agreed.
> So the question is on average are these statistical regularities better or worse than those introduced by the alternative,
Indeed. This is an empirical question. That is what the article is about, for a specific setting.
> ... and the answer is no, if implemented properly.
I don't agree with that. This is not about implementation. It's about how the statistical regularities that are imposed on the full output distribution affect that distribution. There is a change - by construction. That change can be good in some situations and bad in others. The authors claim that it is mostly bad in the setting that they investigated. This looks like a fair statement.