If 10,000 people guess 10,000 fair coin flips each one of them will get more guesses right than any of the others, one of them will get fewer guesses right than any of the others, and the gulf between the two is likely to be over 4 standard deviations wide. I'm certain that I, being an untutored schmuck from Pittsburgh and having thought of this almost immediately after reading about this contest, cannot be the first person to realize this is a potential problem for a forecasting contest. But I can't find anything they've done to mitigate that problem. Can anyone clue me in?