Non-stationary A/B tests
amazon.science
amazon.science
Aside - this paper isn't the first time this issue has been noticed, or even the first time it's been addressed with rigorous statistics. Optimizely did it back in 2017-2018: https://support.optimizely.com/hc/en-us/articles/53262137051...
It would be nice if these authors mentioned the prior art and explained how their work improves upon it.
It doesn't matter how randomly you divide your A and B groups, if you run your test in summer, you'll get different results to running it in winter.