Computing Your Skill
moserware.com
moserware.com
This is demonstrably untrue - skill at chess (to use their example) occurs on several axes, the most obvious of which are opening theory, tactics, and positional play. Players may excel only in certain regions of that space, and it's quite easy to set up a player cycle A->B->C->A, in which each player is more likely to win against one player and lose against another. A players observed 'skill' therefore will depend on any biases in the population at large rather heavily (at the lower levels of chess, the population has overwhelmingly studied opening theory and some tactics)
Because of their non-dimensionality, skill ranking algorithms are universally limited to expressing how likely one is to win against an average person of a given skill ranking, rather than the likely outcome of the match about to be played. Sports match prediction techniques are all domain-specific, precisely for this reason (and because substantial sums of money are riding on their predictive effectiveness).
[1] http://www.cs.toronto.edu/~rpa/adams-dahl-murray-2010a.shtml
[2] http://homepages.inf.ed.ac.uk/imurray2/projects/2011_marius_...
In Cambridge (UK), the River Cam is too narrow for side-by-side racing over more than a few hundred yards. Instead, a popular kind of racing has developed called "Bumps". Crews line up along the river with about a boat and a half's length between each other. At a signal, all crews start rowing, and if one boat manages to touch the boat in front ("bumping"), both crews pull over to the side, and the next day they switch places in the starting line up.
It's not uncommon to see two boats trade places repeatedly -- this could happen because crew A has a much stronger start than crew B, but that crew B have much more stamina than crew A, so A will always chase down B relatively quickly, but B will always manage to catch A over the course of a longer race.
The TrueSkill algorithm generalizes Elo by keeping track of two variables: your average (mean) skill and the system’s uncertainty about that estimate (your standard deviation).
which in Elo terms translates to basically having a non-fixed K. Since that's also the goal of the Glicko rating system (an already-used extension to Elo), I was curious if this article would compare them. It doesn't, but their FAQ does (result: there are minor technical differences, but the big difference is that TrueSkill handles games other than 2-player games): http://research.microsoft.com/en-us/projects/trueskill/faq.a...
Some additional Google-Scholaring turns up that there are some extensions to that as well, notably one that computes the Bayesian estimate using the whole history, instead of incremental updates: http://halofit.org/papers/WHR.pdf
One of the most accessible and interesting is his HTTPS breakdown[1]—highly recommended, and I'm sure it's been HN'd more than once.
[1]: http://www.moserware.com/2009/06/first-few-milliseconds-of-h...
I'm hoping to get back to writing again sometime before the end of the year. Unfortunately, it takes me a long time to write so I'm hoping to get at least one post in this year :)