Since they do have the input, they could probably just store checksums at each step...
... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches").
Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.