Then I tell you: calculate the average of every 3rd flip minus the average of every second. It should come out close to zero, and unlikely to be higher than +/- 30 but for my sequence it comes out to +89. Ok, that’s very unlikely.
So from now on we share this secret. Whenever I generate sequences of coin flips they all came out with this particular statistic out of whack, but unless you’re looking for it, you can’t detect it. In fact, in the real world I use cryptography such that unless you know my secret key, the specific statistic in question is mathematically undetectable.
Now use this sequence of coin flips to pick amongst next tokens in an LLM. It doesn’t change the distribution of words used in the LLM. It doesn’t change its writing style. In fact, unless you know the secret key you cannot detect the watermark. So it’s not that some words are used more often. That would be detectable. It’s just that if you translate back to a series of 0’s and 1’s the average of every 3rd bit minus the average of every 2nd is out of whack.