The evolution of chess: Game lengths and outcomes
randalolson.com
randalolson.com
In chess, a "match" is a series of games between two players. The correct term for an individual game is a "game".
"Plays per person" is known as simply "moves". If each player has made 37 moves, that's 74 half-moves, or 74 ply (in the context of lookahead in the game tree by a computer).
Most analysts, when studying performance, assign scores of 1 and 0 to White and Black when White wins, 0 and 1 when Black wins, and 1/2 point to each when there's a draw. It is unusual to not count draws when analyzing performance. Doing so would decrease White's apparent advantage. For example, I just checked all the 2013 games from Chessbase's Megabase and White scored 53.4% with draws counted as a half point for each side.
By the way, one of the reasons that draws have decreased in the last 30 years is probably the advent of faster time controls. I also suspect one reason that games have gotten longer is the death of adjournments.
To me, it doesn't make sense to assign a draw as 1/2 to Black and White. Technically, neither side won, so why not just throw that game out when looking at which side wins more often? That's what I do in this analysis, although I show the full breakdown in the final area plot.
A game in which for evenly matched players, in a match of 1000 games, 9 are wins for white, 1 is a win for black, and 990 are draws. Would it really be reasonable there to say that W has a massive 90% advantage? I don't think so - in that situation white only has a tiny average advantage per game, so that winning such a match would mostly depend on real skill differences.
Also note that even in a match, where players alternate colors, Black should be happy to draw and take White in the next game, while White would be disappointed with that result. So it is worthwhile to keep track of draw results even if they are thrown out.
When researching openings, players often take into account how well an opening has performed in the past. There's a large qualitative difference between White 60% Black 40% Draw 0% (60-40) and White 12% Black 8% Draw 80% (52-48), although the fraction of decisive games won by White is the same in both cases.
This means the weaker player has a benefit from a draw, while the stronger player drops rating points when drawing a weaker one.
but also increasing popularity of things like 'Sofia rules' http://en.wikipedia.org/wiki/Draw_by_agreement#Only_theoreti...
which has a big impact
The data doesn't support the conclusion that chess is becoming more defensive. It could be that time controls or new rule changes have changed the play. For instance, one proposed rule change is to not allow draws before move 50, which would definitely impact the number of moves.
That's a good point. I weakened the wording to make it clear that I'm speculating.
https://github.com/pickhardt/chessbox
I'm kind of hoping someone will do a thorough analysis of the types of moves chess grandmasters make during a game. For instance, here is a chart from a GM book that categorized moves into four categories: http://imgur.com/BMOElXN
Is confidence interval the right statistical concept to use when characterizing this data set? It seems like it would be more meaningful to examine how various percentiles of "moves per game" change over time rather than plotting a confidence interval around the mean.
In particular, the mean moves-per-game that he displayed is (I assume) the actual mean of his data set. It's not an estimate of the mean of some larger population of games, if I'm understanding the article correctly. So what does it mean to display a confidence interval for a parameter that's known rather than been estimated?
It seems more interesting to treat the data set as the population and analyze how it changes over time. For example, I'd be interested to see a graph of 5th, 50th, and 95th percentile of match length. I notice that the confidence interval is shrinking around his mean -- is that really reflecting that percentiles are converging on the mean because of changes in play style? Or have I misunderstood and the data set is being treated as a random sample of some larger population of chess games? (In that case, since there are more games in his data set in later years, it is unsurprising for the confidence interval to shrink. Whereas if the percentiles of match length are converging on the mean, then I think that's pretty interesting.)
Technically, I don't have the "true" mean because this data set isn't a set of all games ever. As such, I'm treating the games I have as a sample of all games ever, and providing an estimate of how confident I am in the mean I'm reporting.
>I notice that the confidence interval is shrinking around his mean -- is that because there are more recorded games
Exactly! Law of large numbers.
I will also put on the record that kerno once 'pantsed' me - creating a checkmate in about 8-10 moves wherein I failed to capture a single piece of his. It was an epic highlight among a series of long games (one went to 57 moves) where I was triumphant.
This data set is from chess tournaments, so it's predominantly games with skilled players.
(I don't have enough games w/ Elos pre-1960 to show a reliable mean.)
Edit: Citation http://en.wikipedia.org/wiki/Draw_by_agreement
The missing factor is that FIDE took control of the World Championship title in 1946, and implemented a regular cycle of Zonal tournaments, Interzonal tournaments, candidates tournaments to determine a challenger to the existing Word Champion starting in 1949/1950. (The candidates tournament was changed to candidates matches after 1963 - see Bobby Fischer's complaint "The Russians have fixed World Chess": http://sportsillustrated.cnn.com/vault/article/magazine/MAG1... )
Before the war, major chess tournaments were sponsored by wealthy chess patrons. FIDE represented amateur chess players, I think their crowning achievement was their nomination of the Max Euwe (who described himself as an amateur, he was a full-time teacher) as Alekhine's challenger in 1935 (and consequently beat Alekhine and became the 6th World Champion). FIDE had no hold over top-level chess. (Edit: FIDE also suggested Bogolyubov as Alekhine's challenger, twice. In all cases FIDE could not force Alekhine to accept, but Alekhine did accept to these challenges - possibly for financial necessity)
FIDE's World Championship cycles from 1949 onwards greatly increased the number of chess tournaments titled players could participate in, globally. Particularly the Zonal and Interzonal tournaments, brought together thousands of chess players regularly in - at the start - 3 year cycles.
So immediately, the number of tournaments involving titled players that adhered to FIDE rules jumped up substantially, and since the World Chess Championship cycle became the main top-tier events, other chess events naturally standardised on FIDE rules, and gave rise to hundreds of FIDE-rules tournaments.
So yes, the removing of the short-draw avoidance rule, in the context of a massive and new qualification cycle for the World Championship title gave rise to the side effect of shorter games.
Before the war, in patron sponsored top-level chess tournaments, each tournament had it's own set of playing rules. They were generally the same, a modification here and there based on previous tournament experiences of sponsors and influential players. There wasn't a consistent set of rules (until FIDE in 1946). For instance, in the Nottingham 1936 tournament book by Alekhine, he writes that they dropped the no-short-draw rules because it was becoming more evident players were circumventing that rule anyway, so it wasn't stopping non-competitive games from happening.
And the World Championship matches were predominantly decided by the current title holder, and mostly about whether it was financially worthwhile to the title holder to risk his title on a match with a sponsored challenger.
In turn, I wonder whether the tailing off of that in 1970 was the introduction of the Elo rating system and the Fischer-effect leading to the professionalisation, financial viability and the commercialisation of chess?
Maybe a rating system infused more competitiveness into tournaments, more professional application from the players.
Also, the tailing off of game length happened right when Fischer returned to the qualification cycle, where his spectacular series of results culminated in winning the World Championship in 1972 against Spassky in Reykjavik, starting the fracturing of the Russian hegemony of World chess.
Fischer also pushed for chess professionalism from a commercial/financial point of view. More money started to flow into chess because of Fischer and his financial demands / quality expectations. Perhaps that started the conversion of chess from a state-run Russian-owned speciality to a commercial/sponsored prestigious tournaments (e.g. Montreal 1979, San Antonio 1972, Milan 1975). And in conjunction with a rating system, compelled players to be more motivated towards playing for the full point, to improve their international standings, to improve the quality of tournament invitations, leading to more financial stability.
Somewhere between the Russian dominance of the World Chess Championship, up to Spassky losing to Fischer in 1972, and the multi-million dollar Kasparov-Kasparov matches from 1985 onwards, elite chess became a financially enriching career path for more and more players. 1970 may have predicated that, and the gradual rise of the length of games a side-effect of money being pumped into chess and initially the Fischer-effect.
edit: this would have been during the peak of the soviet chess machine's powers, which is what I associate with the first real period of intense analysis into the openings and the storing/hoarding of opening novelties.
Slightly off-topic but... wow! I didn't see this coming. It feels like the advent of the computer age is officially part of History now.