The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
"Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy)
Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?
And the Daring Fireball article does complain that watermarking will reduce quality. If that's what you're trying to check, "which is better?" is the right question.