They do explain their methodology in some detail in the accompanying GitHub repo: https://github.com/lchen001/LLMDrift
They seem to have been taking to the API directly and requesting the two different model snapshots.
I'm not convinced by their methodology generally. It looks like everything may have been run with temperature 0.1, which I don't think reflects most real-world usage for example.