Chess rating system provides the best predictor of World Cup success
kaggle.com
kaggle.com
Are today’s crop better players?
Or like the economy, are Elo ratings subject to bouts of inflation?
Players leaving the system is one of the major sources of rating inflation in chess. Even when rating changes of any single game is zero-sum, if a player stops playing chess after a career of net losses, all the remaining players in the world have more rating to divide up amongst themselves.This is very obvious on FICS bughouse where new sock puppet accounts crop up all the time, net lose several games, and then leave forever.
Part of the problem is new players are usually given a starting rating around 1600 (depends on their provisional games), a "median" ranking of sorts, but in reality a new player is often far worse and donates rating as they slide down to their correct rating.
In the game of Go, online ratings are general subject to deflation. New players appear, are given the lowest possible ratings and then rise to some higher ranking, knocking down the ratings of other players. Also, players who only play each other can improve together without any improvement in rating.
The author poses the following question: "[...] like the economy, are Elo ratings subject to bouts of inflation?" Since the amount of points available depends on the number of individuals participating, Elo points should indeed be subject to inflation.
How are you gathering information about where students were accepted? And where is Caltech in all this?
Also, you know that you can't just add the 25th/75th percentiles for individual SAT sections and come up with overall results, right?
Basically the entire site is dedicated to gathering information from students, and then transforming that data into something useful for them. So I gather the info directly (over 250,000 college apps from 50,000 different people are in the system by now).
I honestly have no idea why I thought I could get away with adding together the 25th and 75th percentiles for SAT subscores on that page; it's not something I do elsewhere on the site. (I'm using "I" here since I'm sure this is my fault alone.) Thanks for pointing that out!
Edit: I forgot about Princeton. The Elo points can be a bit hard to decipher on the fly; for example, Yale should only beat Princeton in 59:41 fashion - not exactly a drubbing. Here's an Ivy+Stanford+MIT cross-admit comparison based on this data - http://college.mychances.net/tools/college-choice-matrix.php...
They give a ranking of 1. Harvard 2. Caltech 3. Yale 4. MIT 5. Stanford 6. Princeton 7. Brown 8. Columbia 9. Amherst 10. Dartmouth
while you give 1. Harvard 2. Yale 3. Stanford 4. Princeton 5. Dartmouth 6. Penn 7. Notre Dame 8. Columbia 9. Georgetown 10. Berkeley (and MIT only #12, and Caltech not even on the list!).
For example according to your Elo points (Dartmouth 1767, MIT 1688), Dartmouth would be expected to "win" over MIT in more than 60% of their battles for students (something I have a hard time believing, with all due respect to Dartmouth), while according to Avery et.al.'s paper, MIT has a significant edge over Dartmouth in their competition for students.
Based on my anecdotal experiences, the Avery et.al. paper seems to give a much more reasonable account of student preferences. Where are you getting your data?
The Avery data is from 1999. They collected theirs by sending out surveys to dozens of "elite" high schools and asking them to hand the surveys to their top 10% (this is from memory, probably not 100% accurate but close).
My data is collected from the 50,000 students who have used my site to track their ~250,000 applications to college -- the bulk of which are from 2009 & 2010.
It's not surprising to me that my results look a bit more, shall we say, "democratic" than those from Avery. I say this because of their selection methodology vs ours (elites vs all-comers). Also, it seems that Caltech is seriously over-favored (relative to the preference of current college applicants) in their survey. I say this only based on Caltech's current acceptance rate (17%) and yield (34.4%). Whereas Avery put Caltech above MIT back in 1999, the modern (official) statistics hint at a different picture. Now, MIT still has an extremely impressive 66% yield (and 12% acceptance rate). Their 66% yield trumps Caltech's 34%, especially if you consider geographic rivals such as the entire Ivy league (for MIT) vs, largely, Stanford and Berkeley (for Caltech). (I'm using yield here because it gives a broad-strokes hint of revealed preference when no better data is available. I don't have data from 1999, and Avery don't have data from 2010, so there is no direct way to do this comparison.)
I'm not sure that there is much to learn from trying to compare preference for Dartmouth vs MIT, frankly. MIT (perhaps along with Caltech) holds a special place in educational mythology. For quantitative people, it's sacred. For some others, it's probably not even on the radar. I don't know how much direct overlap they have in real life, but on our site we don't even have enough direct cross-admits to make a statement about preference: http://college.mychances.net/college/tools/college-cross-adm...
Edit: As an aside, Caltech is not on this list because it didn't have enough applicants to give us a fair sense of its ranking. Given that only 4,000 people a year apply there, this wasn't that surprising to me. I think we'll have enough data to rank it this year.
Edit 2: If you have any methodological suggestions, I'd absolutely be eager to hear your thoughts.
1. Harvard (Elo 2009) wins 78% (would be expected to win 90+%) 2. Yale -- insufficient data 3. Stanford (Elo 1834) wins 56% (would be expected to win 70+%) 4. Princeton (Elo 1788) wins 44% (would be expected to win 60+%) 5. Dartmouth -- insufficient data 6. Penn (Elo 1757) wins 57% (about right) 7. Notre Dame -- insufficient data 8. Columbia (Elo 1742) wins 33% (would be expected to win 55+%) 9. Georgetown (Elo 1715) wins 40% (would be expected to win 50+%) 10. Berkeley (Elo 1711) wins 20% (would be expected to win 50+%) 11. Brown (Elo 1691) wins 40% (would be expected to win 50+%) 12. MIT 13. Amherst (Elo 1687) wins 40% (would be expected to win 50%) 14. Duke (Elo 1682) wins 33% (would be expected to win nearly 50%) 15. Cornell (Elo 1670) wins 24% (would be expected to win nearly 50%)
But it is possible that MIT loses more than it should (if it were 1800) against much lower ranked colleges (e.g. Boston University (Elo 1434), which your site gives winning against MIT 28% of the time). What might be happening is that MIT accepts some fraction of students who are weaker academically for a variety of reasons, but those students are intimidated by MIT's reputation for being academically intense, and decide not to attend. My guess is that among strong students (e.g. students with higher SAT scores), MIT's rating would be much higher than among all students.
This is also consistent with your suggestion that Avery et.al. used students from elite high schools.
So if you want to get ratings that are more in line with what people think of the colleges (e.g. where is a good student who gets into both A and B likely to go) you should probably consider students who have high SATs, or alternatively students who get into a lot of top-ranked colleges, to be more "valuable" and thus to be more important for the rating.
In other words, my aim is to have descriptive, not normative, rankings.
However, I like your suggestion and I'll think about tryin to implement a system, in parallel, to give a different set of rankings with the 'feel' that people expect. It should be fun — thanks for the idea!
I'll look for bugs, as you suggest, since I'm going to be re-running the algorithm again in the coming weeks.
Edit: After thinking more about my beliefs about MIT not really being comparable with your average college, I'll add another possibility: perhaps MIT is simply engaged in fewer matches against weaker opponents, collecting fewer "free" points than HYP?
On the other hand, the model just might not fit, and it could be over-predicting rare events, I haven't looked at the data.
In other words, the features that you suggest are critical, and the system already has those features.
Of course, a single number can only tell you so much -- it's a very gross generalization of skill. A lower-score player can have a particular strength corresponding to another's weakness, thus leading to a win for the (generally) weaker player. I imagine that this is similarly true for soccer teams.
In a system where new players can enter after competition has begun (and others have accumulated points), there's also a good chance that a new player will not yet have had enough time to accumulate a score that's in line with their actual skill, which can lead to some misleading matchups. Depending on the details of the ranking system and the ease of finding similarly-ranked partners, it can take a long time to converge to a "true" rating.
For example: I had friends in college who played competitive card games that had some sort of lifetime ranking (Bridge, perhaps?) Because they were young, they were often the lowest seed in tournaments. Despite above-average skill, they were normally eliminated by the very most skilled teams in the first round, and therefore had few opportunities to accumulate points against more equal competition. Once they finally broke through and won a couple of tournaments, there was a bit of a snowball effect as they moved out of the lowest seed.
In any case football/soccer results have such a big variance that it doesn't really matter. I think in HK there was a paper no long ago that studied soccer competitions and sayed something like the best team in a soccer competition only wins it a 30% of the time.
(This has always bugged me regarding the baseball playoff system. It is a big boost for teams economically but for purity I would prefer to see just a NL and an AL league table).
Deleted comment
Out of interest, did you enter the competition?
I tried the different approaches on for predicting 10.000's of club matches and 1.000's of national team matches
Probabilty of Team A beating Team B, given that Team A has 'A' Elo points and Team B has 'B' Elo points: 100/(1+10^((A-B)/400))