I hope you're not training models based on the "which one is better" question, because that's incredibly subjective.
the dataset is open source and we plan to train an aesthetics picker on it but obviously have to do proper evals (with at least 1M data) to come to a reasonable conclusion.
Disclaimer: I only clicked on Surprise me.