Perhaps the underlying process is just so noisy, better models are not really possible. It would be interesting to see confidence intervals on those predictions - I am relatively confident they would overlap. My intuition tells me that the confidence intervals would be huge and there is very little statistically significant differentiation between the best and worst teams in the tournament - upsets happen in soccer many times every tournament.
I think how successful a model is depends on your purpose too - if you have a model that predicts only 60% of games correctly, while technically you're not very "good", you're doing much better than those gambling on the sport (usually the favorite in sports betting wins at a rate of 52-54% I think) or probably conventional 'sports analysts' (no data to support the second point, that's an educated guess based on the gambling statistic).