Without very careful contest design, the best performers are obviously going to be over-fitting. Especially if the entire distribution is public. That's exactly what this team did.
This is true of academic contests in general, btw, even without cheating. They stop being interesting/fun/good signals as soon as people start treating them as an independent skill set. Comparing performance on chess games for the first N games between two new players might be a good signal for some general intellectual capabilities. Comparing experienced players against one another is mostly just testing who's spent more time learning about chess.