When tackling IMO problems, the hard part is coming up with a good approach to the proof. Verifying your proof (and rejecting your false attempts) is much easier. You'll know which one to submit.
(Source: I am a two-time IMO silver medalist.)
(Source: I am a two-time IMO silver medalist.)
If verifying a good idea is easy, then the evidence shows that the AI didn't have good ideas for the other 2 problems.
Humans are even better at this as you mention - but effectively the approach is similar. Come up with lot of ideas and see what proves it.