World Cup Follow-Up: Update of Winning Probabilities and Betting Results
blog.wolfram.com
blog.wolfram.com
Or put another way they don't just focus on making a general purpose programming language or even a math focused language, they actually work to make the language inter operate well with real world data sets.
In my opinion they've passed Matlab and R in terms of being able to take a raw data set and quickly ask and answer questions about it and as someone who does this for a living I'm very happy Mathamatica exists, although R still wins in terms of cost:)
have a look at the Mathematica stack exchange[1] site to see more code. there isn't any contrivance there.
[1] http://mathematica.stackexchange.com/questions/tagged/graphi...
Even though Belgium comes out slightly ahead in the "chance of victory" graph, and in the "most probably game tree" they are picked as the more likely winner, the lower graphs (chance to reach, chance of knock-out) show the US as having a higher chance of making it to the quarter finals than Belgium.
Isn't this inconsistent?
I'm not saying this is the case, but it explains how you could have a lower probability to reach the quarter finals but a higher probability to win overall.
As such, there is no way the probability of winning can exceed the probability of reaching the quarter finals.
We have
P(Belgium reaches quarters) >= P(Belgium wins overall)
P(USA reaches quarters) >= P(USA wins overall)
P(USA reaches quarters) >= P(Belgium reaches quarter)
That does not imply P(USA wins overall) >= P(Belgium wins overall)I think the US lost their last group match, but Belgium won theirs, so maybe something like shifted the model probabilities between the time of the 2 graphs.
#1
* 538 - Brazil - 36.0%
* WOLF - Brazil - 32.2%
#2
* 538 - Argentina - 17.0%
* WOLF - Netherlands - 23.5%
#3
* 538 - Germany - 12.0%
* WOLF - Germany - 21.6%
This paints a pretty interesting picture about trying to use statistics and math to predict sporting tournaments like this. Seems like nobody has been very successful at this yet. Obviously it gets "easier" as the tournament progresses, though.
I think how successful a model is depends on your purpose too - if you have a model that predicts only 60% of games correctly, while technically you're not very "good", you're doing much better than those gambling on the sport (usually the favorite in sports betting wins at a rate of 52-54% I think) or probably conventional 'sports analysts' (no data to support the second point, that's an educated guess based on the gambling statistic).
Brazil: 24.2% Germany: 23.0% Netherlands: 19.0% Argentina: 15.7%
Someone would bet Argentina #2 if the have faith Messi will play what he's used to, or Netherlands if they consider the actual performance in this cup. This reflects 538's stronger bias for player skill score, while Wolfram seems to adjust more for the past results.
http://blog.wolfram.com/2014/06/20/predicting-who-will-win-t...
Just goes to show, when it comes to sport, you can throw away most of these models until they start taking into account news and gossip from tabloids which probably has more bearing on team performance than raw numbers.
- probable weather - some teams play differently depending on conditions
- yellow and red cards - some major player could miss a match. If team A is missing a defender, might be that is more vulnerable and thus, may lose a match.
and so on.
Anyway, good work!
The thing that I find hard to predict and build into my models is style of play. By style I mean: spatially-intensive, high-time-inpossession-the-ball-time (e.g. Spain with tiki-taka and Germany to some extent) versus time-intensive, opportunity-seeking/opportunity-creating (such as Brazil, Argentina, etc.)
Why? Because passing-intensive teams seem to display more of an own effect -- they fall or rise on the strength of their team, since it's an intricate, very technical and collaborative style. The results of opportunity-seeking teams are much more dependent on the strength of the adversary -- i.e. much more Elo-like.
Ideally, I'd be able to infer from the data a (exponentially biased to recent games) own-team/spatial play dependence factor as opposed to a strength-of-opponent/opportunity-seeking factor. In principle if all victories were explainable by a combination of those two variables the Elo residual/surprise would be a measure of this, but hey, teams get better/worse at opportunity-seeking too, even teams specialized in tiki-taka.
I'm not saying you have to write a thesis to suggest new data sources, but I am saying that "you need to do XYZ" is not a productive piece of advice without the evidence to back it.
Now strictly to subject, would be indeed a nice exercise to take a team performance ( under any scoring formula ) and correlate it with the weather at that time. This data ( team data and weather data) exists. If we find any correlation - not sure. Could be that for some teams it doesn't exist - they play good no matter the atmospheric conditions.
The second exercise is the yellow cards, but this is indeed hard to factor since is very dependent of the coach strategy ( we cannot assume that the same player always plays). So, yes, OK, let's leave that out.
Also, if you're doing statistical analyses of your own, hopefully this isn't how you react to domain experts giving you advice on factors you should consider. If you react the way you just did, it's quite likely people will just shut up and you'll be left floundering the darkness.
Plus: History against X. If a team have played against another a lot of times it build some model in how perform. Latin america teams know more about each others than teams abroad.
I see in this model not much discussion in how latin-teams have the upper hand here: The climate, the people, the shared-history, the kind of game, etc
For example, except for Spain in South Africa, no European country has ever won the world cup outside of Europe. Why is this? A large part of this is due to the weather. And unsurprisingly, South African weather in the southern winter isn't terribly different from what the Europeans are used to. This stat alone should've predicted a lot of the European "upsets" and should trigger your suspicions about 2 out of the top 3 favorites being European.
And even if you were unconvinced about this prior to the tournament, having watched teams struggling in Manaus should've convinced you that these statistical analyses ignore important information that is obvious to even semi-casual fans.
The last time an Europen team didn't make it to the finals was in 1950. If climate really did matter that much, you wouldn't expect that. In addition to 2010 that you chose to ignore, the US world cup final went to the penalties and the 1986 Mexico final was tied 2-2 between Argentina and West Germany until something like 10 minutes before the end. (And a non-European team was in the finals largely thanks to the most famous refereeing error in the history of the sport).
That evidence looks really weak.
Im Brazilian and my childhood city was between -3C to -10C on winter, more in the South Brazil(Paraná)
I Think half of the stadiums in the world cup have mid to cold wheater conditions
Belo Horizonte, São Paulo, Curitiba, Porto Alegre
Tropical Wheater:
Rio, Cuiaba, Natal, Fortaleza, Recife, Manaus(this is probably the worse)
Lets not forget that some of those latin americans nations are used to cold wheater conditions, like Uruguay and Argentina(as in South Brazil)
So it depends..
More like outperform the market?
Not sure if more precise probabilities outweigh the risk of ruin, if there's unbalanced amount of money backing each option.
Also team Elo has issues, as not a lot of games are played, former achievements weigh in too heavily.
Nonetheless, this is a really nice effort. Let's see how well it works.
[1] http://fivethirtyeight.com/interactives/world-cup/ [2] http://rogerkaufmann.ch/dsaINTe_r.htm
So Italy, that won the world cup 4 times, wasn't one of the 10 favorites??
No way.
"So you're saying there's a chance..."