Refusal behaviour specifically is interesting, because if you can point out specific cases where refusal behaviour significantly deteriorates due to the watermarking, it may create a token "route" that may be possible to exploit by adverserial prompters. Static hazardous prompt refusal belies the fact that actual adversarial prompters will adapt their techniques iteratively and gain way higher compliance rates.
My idea would be that the ngram size over which the watermarking works is necessarily limited in order to resist edits better. It might be possible to lead the model to trigger the refusal in the form of these specific ngrams, the completion of which is then more likely flipped to compliance (due to the logit bias introduced by the watermarking), making hazardous requests systematically more likely to be accepted?