There's no difference between those two things. The distribution that matters is the distribution of tokens that are picked not the distribution of tokens the LLM model passed to the selector.
I think there is a possible weakness in the context of the watermarker but that is not your claim iiuc.