The comments here are terrible. I got a better idea of what’s going on by asking ChatGPT what this paper’s weaknesses are:
https://chatgpt.com/s/t_6ab7d694885481918083b8cbf0ba9040
In particular: sometimes they measure “churn”, which doesn’t show whether the results are better or worse on average. They sometimes only test with one random seed. There are multiple-comparison issues. And they’re not testing Anthropic’s algorithm.