In reality, there can be billions of red/green lists. It is trivial to test a piece of text against all of them to see if it was generated using any of the keys that made a particular red/green list. And, as you alluded to, there are enough combinations that each individual account could be assigned its own red/green list so not only would a person be able to tell if the output was AI generated, they would also be able to tell who generated it.
I believe many are missing the point as to the effectiveness of this watermark. It will be exceedingly difficult to get rid of it. Probably impossible. If normal human text uses, more or less, a 50-50 ratio of red and green tokens, and the AI generates your text with 80% red and 20% green, all the rewriting in the world is not going to get the result close enough to the human like 50-50 to avoid a statistical aberration that will be discernible.