Chess World Championship: Stockfish live-analyzing game 7
analysis.sesse.net
analysis.sesse.net
This is their 7th game. They tied for all the previous 6.
They’ve come a long way indeed!
At 12, he was "only" around 2150 FIDE.
I always believed it is because either FIDE ratings are only updated once a month (see 7.1[1]) or because their ELO is too close, therefore the prob. of winning is 0.5 for both anyways.
[1] http://www.fide.com/component/handbook/?id=172&view=article
I heard (possibly from the Chess24 commentators) that if their difference in rating was 4 or more, each draw would bring them closer together in live rating - Magnus would lose rating with each draw; but as it's 3, their ratings don't change. i.e. anacleto was right, thomasahle wrong.
Also, Elo isn't an acronym (ELO) as one might think - it's named after its inventor Arpad Elo. https://en.wikipedia.org/wiki/Elo_rating_system
Interestingly, when I first looked at this page, most of the good/correct comments were voted down, most of the bad ones not, it was weird. (It made me suspect that other HN pages are like that but I don't know it! - I know chess better than most subjects discussed on here) But that has now almost totally been corrected.
For example the requirements for the most common way to earn the GM title include a requirement that your rating has reached 2500. The live rating is used for that.
https://en.wikipedia.org/wiki/World_Chess_Championship_2018#...
The situation in matches is different to that in tournaments, where it's no good drawing all your games if you want to win the tournament. In this match, drawing every game, then drawing all the tiebreak grames, then getting black in the Armageddon game and drawing, will make you world champion. And historically, if a world championship match was tied, the champion automatically retained their title. (Tiebreaks are a recent development)
[0] https://en.wikipedia.org/wiki/Draw_by_agreement#Steps_taken_...
[1] e.g. Here's Capablanca (world champion 1921-7) writing at length about it in 1925 http://www.chesshistory.com/winter/extra/capablanca7.html
He created a version of chess with 2 new superpowered pieces to avoid the 'draw death' problem https://en.wikipedia.org/wiki/Capablanca_Chess (that page mentions some other similar variants)
The computer shows Black wins with 68..Bh4 here. But had Caruana played the incredible 69.Bd5 Ne2 70.Bf3 Ng1!! they would request metal detectors immediately! No human can willingly trap his own knight like that.
0 means that the position is exactly equal (material + positional)
The engine always has an idea of what's going on through simple material calculation + positional heuristics (although it maybe be flawed)
Back in 1972, during the Cold War, we had the chess "Match of the Century". In one corner the USA, in the other the Soviet Union. Bobby Fischer soundly defeated Boris Spassky.
All before the Internet and before computers analyzing chess. So the local PBS station in NYC had chess master Shelby Lyman analyzing moves, for hours and hours, as they came in from Iceland. Perhaps 5 or 10 minutes between moves.
It kept me amused one summer during high school.
See https://www.quora.com/Since-chess-is-not-a-physical-game-why...
For example, Hou Yifan, the #1 woman in the world in terms of rating, is 91st overall.
They really should get rid of these archaic female only leagues.
Rather, to have the segregation by blitz or full time controls, but some players can be in both if they want to do.
The reason for being able to see lists by the other classifications is because it's the same way that tournaments are frequently divided. Beyond having rating sections in tournaments there are generally also events that are only open to people below a certain age (or seniors events for players above a minimum age), or only to females. Females are also granted certain female only titles. For instance Hou Yifan holds the title of Grandmaster as well as the title of Female Grandmaster. The latter having substantially easier requirements.
To be clear, these divisions are all completely undirectional. Females, young, and old players can compete in any event and be rated against any player. Hou Yifan is regularly invited to top level events. For instance this year she was invited and competed in both Tata Steel and the Granke Classic closed events, competing against players including Magnus Carlsen and Fabiano Caruana. She scored a total of 0 wins, 10 losses, and 12 draws.
Edit: Alphazero's record against Stock Fish is 28 wins, 72 ties and 0 losses.
In a 100 game matchup, AlphaZero had 28 wins 72 draws and 0 losses. I'd classify that as orders of magnitude better.
https://www.chess.com/news/view/google-s-alphazero-destroys-...
https://www.chess.com/news/view/google-s-alphazero-destroys-...
But yeah, it would be interesting to see a side by side comparison of alphazero and stockfish's live analysis of these games.
> AlphaZero compensates for the lower number of evaluations by using its deep neural network to focus much more selectively on the most promising variation [1]
Even if you compare CPU to GPU by price, and not speed, it seems pretty even. It clearly has to sacrifice speed by making more intelligent pruning decisions than stockfish.
1: https://en.wikipedia.org/wiki/AlphaZero#AlphaZero_vs._Stockf...
2: https://sites.google.com/site/computerschess/stockfish9-benc...
3: https://www.reddit.com/r/hardware/comments/9jyts8/rtx_2080_t...
The computer itself was strong. But the weird timing setup, the disabled databases (ie: both opening book AND endgame book was disabled. Stockfish normally plays PERFECTLY when the board is reduced to 6 pieces or less, as well as perfectly knows the winner / loser in every 6-piece setup. But that was disabled for the AlphaGo games)
Without the ability to run AlphaGo on our own and recreate the test, we have no way in knowing how AlphaGo would work under "fair" conditions.
-------
And btw: 1GB of RAM is a lulzy setup. You put all the CPU time you want, but gimping the RAM down to 1GB is... weird. Its a VERY suspect "test" that none of us can replicate.
Remember that if stockfish uses tablebases and opening books it's really "cheating" by using human knowledge.
Although I am surprised that they reduced the hashtable so much and also used stockfish 8 when stockfish 9 was available. It was not RAM but just the hashtable size also their "test" has been replicated by plenty of people.
We have our own "replica" of the NN playing chess in lichess.
So if A0 wasn't built to access them, but stockfish did in the match, then that would be an unfair advantage.
Why is that an unfair advantage? AlphaZero developers wanted to prove that their AI was better than Stockfish.
They failed to do that. If AlphaZero wasn't built to access Tablebases, then they should have built it to access tablebases. Don't unfairly gimp Stockfish because you're lazy at programming.
One way is to do according to the most capable configurations of each program, so that means Stockfish with all of them. Other way is to specify the same limits of time, memory, etc of each one, although that is still not the excuse to avoid the opening books and tablebases and so on unless they use up more memory than AlphaZero does.
I thought the "point" of AlphaZero was that it was better than the sum of human knowledge.
By the way: tablebases are computer generated. There is no human alive who knows all combinations of 6-piece endgames. Tablebases are brute-forced endgames, created by computers to be used by computers.
Stockfish itself was only really designed to play the midgame. Its designed to be used with both tablebases and opening databases.
> It was not RAM but just the hashtable size also their "test" has been replicated by plenty of people.
Citations please. I'd love to see a "proper" test of AlphaZero vs Stockfish.
Your incorrect perception comes from the fact that, at very high levels (as shown by this match as well), draws become way more common than for lower levels. A 3400 vs 3300 Elo match might be 28-72-0, a 1400 vs 1300 match might be 60-8-32. At the lower level, you see one player winning twice as many games as the other, at the higher level you see one player losing all the time; that looks different to you, but as far as Elo difference is concerned the two results are exactly the same.
I dunno why AlphaGo devs would do that, aside from artificially gimping Stockfish on purpose.
Check out the Leela project, which is trying to reproduce and improve the AlphaZero project to also beat Stockfish dev: http://lczero.org/ (at this point they are about equal with Stockfish 9)
28Win 72Draw and O Loss in my opinion reflects at least 300 ELO (8 times stronger).
I guess that the spectacular 28-0 was somehow confusing my ELO opinion engine by concealing the fact that 72 draws are a lot more difficult to achieve than 12 draws (28-12-0 would have been a 300 ELO difference).
It's true that without some basic knowledge of chess it would be hard to follow, but that can be said about many sports.
For example, here's a good move-by-move analysis of game 6, including (around the 33-minute mark) an explanation of the "missed" win:
https://www.youtube.com/watch?v=4yzaG0Ia_fs
And tl;dr for those not following the match: a chess engine analyzing the game found a line Caruana could have used to force a win in game 6 (the actual result was a draw), but it required making such an unnatural and normally bad move (trapping his own knight in a corner) for a relatively distant payoff that commentators pretty much all agree no human would ever have spotted or played it in a live game.
Here's Garry Kasparov, for example:
https://twitter.com/Kasparov63/status/1063576827850096640
But had Caruana played the incredible 69.Bd5 Ne2 70.Bf3 Ng1!! they would request metal detectors immediately! No human can willingly trap his own knight like that.
If Mikhail Tal was alive today, he would play boring chess, too. Otherwise modern GMs would just accept his sacrifices, not make any mistakes, and win.
Players like Tal had their share of boring games, it's just that we remember the crazy ones because they're fun.
Imagine if the Lakers played the Warriors, and late in the 4th quarter, the game is tied, and LeBron and Curry decide they are tired, declare a draw and go to their locker rooms. Not particularly satisfying.
It also doesn't mean the game was unexciting.
Saying that, Carlson once had a reputation for not agreeing on a draw and being able to maintain focus while his opponent faded and made a mistake.
To make this more concrete, the following link is a drawn position (it is trivially easy for either side to avoid checkmate, indefinitely)
http://www.jinchess.com/chessboard/?p=---k------------------...
However, checkmate is still possible (White can checkmate Black if Black makes severe blunders that even beginners probably wouldn't make)
Do you think the players should be forced to play on here?
But, nevertheless, here's an example with equal material where checkmate is possible for both sides, but not without blunders. Who should have to resign here?
http://www.jinchess.com/chessboard/?p=--------------------k-...
If you really insist on saying "well yes, they should have to keep moving the kings around aimlessly for hours until somebody's clock runs out", you're proposing a game that is so fundamentally unlike chess that it should probably be given a different name.
[1] https://www.pentagram.com/work/world-chess
[2] https://www.bbc.com/news/world-42425587
I really like the ones from [1], now that's what I call great design.
Edit: The album art: https://www.chillygonzales.com/wp-content/uploads/2017/09/GE...
Edit: I was commenting on the 2014 design.
It's wort noting that up until recently, FIDE had a president who was full-on batshit crazy. He repeatedly claimed in public that he had been abducted by aliens. He (and by implication FIDE under his rule) was sanctioned by the U.S. Treasury for financial involvement in the war in Syria. Corruption was an open secret. It's a bit of a stretch to call FIDE under Ilyumzhinov a "serious organization".