Who is the most accurate world chess champion?
lichess.org
lichess.org
During yesterday's WCC Game 6 the computer evaluation meant little when players were in time trouble. Anything could have happened going into the first time control, despite the game being dead drawn for the first 3.5 hours.
In the final stages the computer again evaluated the game as drawn, but presumed Nepo could defend perfectly for tens of moves without a single inaccuracy. Super GMs can't do that given hours or days, let alone minutes.
Last thought: did anyone else assume this was written in R/ggplot2 at first glance? Seaborn and/or matplotlib look strikingly like ggplot2 now days!
Perhaps alongside centipawn loss (a measure of how many hundredths of a pawn a player loses by making the non-optimal move as determined by a chess AI engine) we could also measure the difficulty of any position.
Stockfish (a popular chess engine) roughly works by constructing a tree of possible moves and evaluating the score according to some heuristic at its maximum depth. The best result at depth n (25 I believe) is considered the best move and incurs 0 centipawn loss.
Perhaps we can define the difficulty of a position by the relative centipawn loss at each preceeding depth in the tree? The difficulty of a position is then determined by the depth at which the best move no longer changes.
- Engine evaluation of a leaf of the tree will always be different and more sophisticated than human heuristics. So there's a problem where a human can't be expected to follow down some lines. Of course, this is always changing, as humans seek to understand engine heuristics better. Carlsen's "blunder" at move 33 was a good example of this, from my memory.
- Maybe there's a difficulty metric like "sharpness", some function of the number of moves which do not incur a significant centipawn loss. Toward the end of game 6, Carlsen faced a relatively low sharpness on his moves, whereas Nepomniachtchi faced a high sharpness, and despite the theoretical draw, this difference will prove to be decisive between humans. This seems like it could interact in interesting ways with your difficulty metric - for example, what does it mean if sharpness is only revealed at high depth?
- It would be interesting to take the tree generated by stockfish, and weight the tree at each node by the probability that a human player would evaluate the position as winning. Then you could give a probability of ending up at each terminal position of the tree. Maybe some sort of deep learning model trained on players previous games? Time controls add such a confounding factor to this, but it would be so interesting to see "wild engine lines" highlighted in real-time.
For example, in yesterday’s game Stockfish was often giving a drawn evaluation (0.00) where Leela Chess gave a win probability of 30%+. I was posting about this during the game.
https://twitter.com/nik_king_/status/1466794534214504454?s=2...
Chess.com did do one study and found a large percentage of mistakes occurs in moves 36-40 because in some time controls additional time is added at move 40.
I agree that it’s not very useful to compare with table bases, especially given the “30 seconds time added per move” regime this was played under by the time they reached the position.
However, I don’t think the table bases even have enough information to indicate how close to losing a theoretically drawn position is. So, i don’t think this required perfect accuracy to defend against (defining ‘inaccuracy’ as any move for black that either makes it take longer to reach a draw or moves to a losing position. That, I think, is the most reasonable definition)
And if we are talking about practical chances, why should we rely on computer-centric evaluation? If a human has to choose between a move that leads to the win but they have to find 40 best moves or they will lose and a move that is a theoretical draw but now the opponent has to find 40 moves or they will lose, what should a human choose?
What is even the ACPL of a move from a tablebase? There is no value, it is either a win, a draw or a loss. So while the whole idea behind this exercise is intuitively appealing and certainly captures some sense behind the idea of accuracy, it should be taken with a grain of salt.
As for the tablebase question, it would be nice to see win/forced-draw probabilities from engines instead of the increasingly artificial material evaluation.
it's ironically also a murky concept for the opposite reason. In some openings the analysis of GM's goes so deep that they can fairly often play almost exclusively computer-aided prep. There's a big difference between a 40-move game in that kind of theoretical position vs off-beat games.
So you might have a very precise, narrow, theoretical repertoire but that's not the same as playing strength because your opponents can prepare. What really matters is more something like precision under uncertainty.
This would be a really good follow-up experiment. If the theorized result really happens, we would have strong evidence that players are "overfitting" to their training chess engine. It would also be interesting to see how stable the historical figures look between different engines.
But also
"Since 1941 Zuse worked on chess playing algorithms and formulated program routines in Plankalkül in 1945."
If the accuracy is high, not only it means that the players are good, it also means that they don’t ask each other serious questions. Put any human against Stockfish and, I am sure, their ACPL will increase dramatically.
It seems like just looking at ACPL isn't looking at this correctly. If someone makes a mistake, and loses some centi-pawn, but it induces an even larger mistake in their competitor, that wasn't a mistake, it was a risk.
It could bring many players "back to life". It would be even possible to watch "impossible matches" like Kasparov vs Capablanca!
I don't think they were suggesting that's the result they wanted - if you could somehow magically reanimate Capablanca in real life and pit him against peak Kasparov, he might lose badly.
A neural net having the same outcome is essentially what's being asked for. Kasparov raised on Capablanca's era chess or vice versa would be unrecognizably different players, and I don't think anybody expects an AI to simulate their soul.
But I remember watching Hikaru Nakamura stream once playing through each of these bots (and beating them fairly easily). He commented that several of the bots were doing things the real players would never do, both in style and even the opening move (1.e4 for a player that almost always opens 1.d4)
It was fairly early after the personality bots came out, so maybe they've fixed it by now.
- get a chess playing algorithm (I think it will probably well with minimax or mcts) with many tunables,
- use a genetic algorithm to adjust the tunables of the first algorithm; use how similar it plays (make it choose a move on positions from a database of games from said player) as a goal function.
Doesn't seem terribly complicated to do, but don't know how similar to a human it would play.Personally I find it odd to measure how well the players match the computer program and call it accuracy. The computers do not open the game tree exhaustively so they give only one prediction of true min-max accuracy.
When Lee Sedol made move 78 in game 4 against AlphaGo, it reduced his accuracy but won him the game.
It now seems humorous that Kasparov once accused people of helping computers behind the scenes. Now chess masters have been caught huddled in bathroom stalls with their smart phones. Chess commentators choose to willfully ignore chess engines in their presentations, in order to enable our understanding of the analysis. The torch has been passed.
Personally I’ve never felt Magnus enjoyed the modern game with as much opening preparation as we have now. It seems like he’s only in the last few years invested the time in this, instead of relying on his technique to win even from losing positions. I hope AlphaZero proving that fun positional ideas like pawn sacrifices and h4 everywhere reinvigorated him somewhat during his dominant first half of 2019, so there’s still hope the machines haven’t just drained the romance from the game, even if their ideas remain dominant.
(Of course, as with all historical players, he would be stronger if he were re-animated today and exposed to modern principles and openings.)
Very very unfortunate timing but still a valid question.
For example, Karpov and Kasparov sometimes agreed short draws. I wonder if that is flattering their figures.
Isn't lichess open source?
https://lookingforfinesse.github.io/lookingforfinessevariant...
https://support.chess.com/article/1135-what-is-accuracy-in-a...