A bit disappointed to be honest. They produced a bunch of models and then checked with a team of meteorologists what they thought of them vs the other models? This sounds more like GPT-3 writing sonets and getting a bunch of poets to evaluate them. Why not just check the predictions?
TBF, going by the problem statement in the article, the objective function could be pretty wacky. For example, do you weight accuracy by location? Is it more important to get your predictions right over sports stadiums than over residential areas? Would it be useful to weight accuracy by time, like, it's more important to get predictions right at 7 AM when people are driving to work than it is to get it right at 1 AM when everyone is asleep? I don't think that that's what's actually happening, but I do think that weather is a sufficiently ugly problem that comparing to human performance is useful.
Surely they must have already established a loss function in order to train the network in the first place?
Sometimes the metric you use for the loss function is not the best metric for eval. Think BLEU for translation, for instance (not that BLEU is particularly great).