Your point of view is outdated ;-)
In recent ML competitions, participants do submit code that is run on a held-out dataset - as was the case in the PetFinder.my challenge in question here.
Most competition platforms are migrating to this format, as otherwise you can just label by hand as you said.
Note that this competition went even further: not only was the evaluation code run on Kaggle, the training code was also run there. This means that you couldn't even train a gigantic model then submit it: your model had to be trainable within well defined time and resource constraints, which is a great way to level the playing field.
Of course, there's still some unfairness as people with more resources can try out more solutions before submitting a model to be trained on the platform. No platform has a solution for this yet!
The "cheater" in this competition is both a world-class data scientist and a reverse engineering hacker. Heck, people used to write papers about how they crawled the ground truth. Now they see their name in the newspaper.
This is true of academic contests in general, btw, even without cheating. They stop being interesting/fun/good signals as soon as people start treating them as an independent skill set. Comparing performance on chess games for the first N games between two new players might be a good signal for some general intellectual capabilities. Comparing experienced players against one another is mostly just testing who's spent more time learning about chess.