That's weird. Having the community study this would certainly help them. They're afraid this is giving too much insight into their proprietary training/modeling methods?
used to be really useful for detecting text written with the same model, as it was high probability... unfortunately the probabilities are messed up by RLHF.
Ah, that's it; polite fictions are scored higher than uncomfortable facts.