Stockfish 14
stockfishchess.org
stockfishchess.org
It's great to have a better engine, but I feel that would benefit the most the online chess community is not a better engine, but an open source cheat detection system, if that's even possible.
I wouldn't know how to build one, but I think that is a lot more important for chess right now, still it's great to have a better engine so congratulations and thank you to the Stockfish team.
But sometimes a computer will make odd moves no human would ever make. I've been playing against an ios stockfish app to relearn how to play, and when it gets behind it starts throwing material away to delay the inevitable; a human would more likely keep the material and hope the opponent doesn't see the path to victory.
One way to detect use of an engine is to look for moves like this, though if it could be done algorithmicly then that same algorithm could be used to make the engine play more like a human.
Presumably an excellent player might often make the same moves as an engine, so this measure alone isn't going to be perfect. But it could be a starting point. You might also look at the player's historical performance and watch for suspicious changes, or perhaps look for patterns in the time taken to play the move?
The really hard part of chess is what to do in the mid game, once you're off your scripted opening, there is still lots of material, and neither player has any significant vulnerabilities. A computer is useful here.
I don't like giving out ideas here, but I feel like these are obvious ways one could cheat and go undetected.
That would be hard to distinguish from someone memorizing an opening book. So hard to detect.
However the most obvious cheaters are more easily given away by time between moves. When they take the same time between every move whether it be a deep positional move or an obvious recapture, you can be quite sure something fishy is going on. Sometimes they can have literally 1 legal move and still take 10 seconds to find it.
A good player using an engine sparingly however would be very difficult to spot in online chess, especially in a single match.
Note: I have not cheated myself, but I can definitely see a way it can be done. I'm not sure if it would be good of me to describe the process of course... Just think what input you can have and what output you can get if you were to do this programmatically and you can very well imagine if it's possible, there are also existing tools for that. You don't have to open a chess engine in another window and manually do the movements, if you know how to script.
In bullet, I think may be, you could technically even use something to do "anti blundering", meaning you will blunder a lot less, because engine will just check whether it would be an obvious blunder, and block your move. May be you just allow few blunders, and I imagine it would be undetected. Sorry, again for brain storming about that. It is fascinating topic though. Engine could be running on a lower depth and it could be more sort of positional engine that does not do magical engine moves, but is trained on human players and using neural network mostly. Maybe it will just help you do theory openings. And you won't be able to charge anyone for cheating for following opening theory.
I also think the cheating detection algorithms can be beated, and I believe I could do that. Why do they still serve their purpose, more often than not?
To me, it's inherently linked to the very nature of online cheating. It's essentially a futile, nonsensical activity. The only gratification is an illusion of intellectual superiority, whose worthlessness is so transparent that it can only attract people who don't get to experience the sense of intellectual superiority pretty much anywhere else. As harsh as it may sound, your average cheater is rather stupid. That's why it isn't really difficult to catch 90% of them.
I don't rule out there are some cheaters who do it out of intellectual curiosity, but that would be a statistical outlier.
For real ranked matches, participants must have 360 webcams and so on showing they aren't cheating
Perhaps a nit-pick: they need to discover the legal move, and discover that no other moves are possible, right? As a rather basic chess player myself, I can imagine I might spend some time on this depending on the situation.
Of course a beginner would take some time to spot this, but it's unlikely you would confuse a beginner for a cheater, since a beginner will likely make many sub-optimal moves and spend lots of time thinking in general.
* Some cheaters will just 100% match the best engine moves. If a player consistently does exactly what Stockfish would do that's an obvious giveaway.
* Some cheaters will be manually copying moves between the chess website and their engine; in high-speed games ('blitz' and 'bullet' chess) their abilities plummet when there are only a few seconds left on the clock, because they can't copy fast enough.
* Similarly, a player who takes 5 seconds a move whether they're pounding out a basic book opening or making an inspired move in an extremely complicated situation will raise suspicion.
* Some cheaters will just be improbably good for their known background. A few weeks back some billionaire beat five-time world champion Vishy Anand in a charity game (where Anand played a bunch of different games at once) which is the chess equivalent of Mark Zuckerberg outrunning Usain Bolt.
* Chess engines will sometimes make moves that even the top humans fail to see. All the action is happening on the right of the board, and some innocuous move on the left of the board produces a perfectly executed forced mate in 15 moves? Some people will look at that suspiciously.
Of course, a sufficiently careful cheater could cheat without triggering any of these heuristics - a player who only relies on the engine for one or two key moves can easily be undetectable.
An engine user would definitely beat them, unless they were using it sparingly of course, which can be, but I don't think it would be that obvious for you in this case as well.
There are examples where they face an engine and it's obvious, but it doesn't seem 1 out of 7 times.
I would guess that even with the engine you would take some games to rank up to that high so if chess.com is good at banning cheaters most of them would probably get caught sooner.
I do meet a lot of engine players on 3 minute blitz but then I just do very fast bullet moves and all of a sudden I'm losing with a 90 second advantage that cannot be recovered if the user persists on playing with an engine.
And Daniel Naroditsky definitely would have good internal cheating detection even when the user does it sparingly, as he can basically understand most lower rated opponent moves, and if it seems too good for this rating, he can know this from just few moves.
Check Daniel's rapid speed run account: https://www.chess.com/member/ohmylands
145 wins and 3 losses. The 3 losses had, were lost purposely, if you look here in 3 or 4 moves with 5 accuracy, for whatever reason (to drop ELO?). https://www.chess.com/games/archive/ohmylands?gameOwner=othe...
If there were 1 out of 7 opponents using an engine, he would never have 145 wins against 0 losses.
And again for bullet and blitz there are also quite many examples with 100W to 0L.
If you'd look at all the speed run videos Daniel has done, it definitely doesn't seem 1 out of 7.
Daniel also posts all the times he's facing a cheater to youtube, same with Chessbrah and others as it makes for a good content, good views as people are always interested in seeing a GM playing against a cheater.
Not sure what the aim of cheating is, probably going from a rating of 1500 to 2000 and the status that it brings. In the end you still win 50% of the games, just against higher rated players.
Playing against a GM would reveal your cheating instantly, it seems. It's like robbing the police station :)
Not always. Hikaru does speedruns[1] using an alt acct with entry-level ELO to race to ELO 3000. I've watched a fair bit of this, and seen him encounter the odd cheater or suspect game, but much nearer 1% than 10% of the time.
It's much hard to cheat in 3minutes games. The chess engine takes somes time to think about the next move
First of all it's not that bad to play against a cheater once in a while. If you compare it with other games, playing against an engine is a huge disadvantage but will not fundamentally change the structure of the game. You are still playing chess, but against a superhuman opponent. You don't want to play against the computer but it's not as bad as the other player abusing a glitch in the game.
Secondly, I'm guessing that cheaters will mainly play at the entry level strength (1200 on chess.com) and a bit above that. If you are seriously cheating you will be caught very quickly. So maybe if you change your rating you might encounter less cheaters.
Edit: I just looked at your comment history to find out what your rating is and apparently you are playing (for an online game) with extremely long time controls? That's probably the reason why you are encountering many cheaters. The player pool for long online games is much much smaller, so you will automatically have more cheaters who just recently signed up for the game.
It wastes your time. Playing against a human is a different experience. If you actually wanted to practice against an engine, you would do so knowingly. With some possible benefits such as takebacks etc. (since computer is not a rival, just a training tool).
It wastes your rating points - if you play rated games. Obviously not everyone does, or cares about their online rating; but I do to an extent. For one, while rating isn't a goal in and of itself, it's still a convenient form of tracking my progress, and cheaters distort this measure.
Finally, it wastes your nerves. However insignificant this may be in the scheme of things, I think that most people still dislike being cheated or lied to (in any way or form) simply out of principle, and find that frustrating.
- simulate a game using stockfish
- for each move (except few moves at the beginning) compare the move made by player with the list suggested by engine - if the move chosen by player is on the list generated by engine, than give that player some points (depending on the position of the move on the list)
- do some math considering player's ELO and some other stuff (I can't remember exactly).
Definitely not an ideal solution, but also open for improvements. Btw it wasn't my idea - chess players provided the exact algorithm, so it must have been known.
It's also worth keeping in mind that you will sometimes see players match the best engine move 95% of the time or more at the 800-1000 elo's and they're not cheating, it's just their opponent is blundering and the next move is obvious.
So specifically, you have to find when players matched up with engine moves, where the engine decided on an optimal move by looking far into the future.
This statement seems a bit funny because in order to have a good idea that they cheated, you would have also had to been analyzing the game with the chess engine.
Regardless, unless you truly an amazingly player, nearly any chess engine made in the last 15 years will destroy you and incremental improvements on stockfish have absolutely not effect on that.
You analyze the game after it is played. When your opponent managed to have a 99.9% accuracy in a 1500+ ELO blitz/rapid game, it's highly unlikely that they managed to do that without some computer assistance.
I'm 1500+ ELO, play blitz/raipd, and get 100% from time to time.
It's usually because I played some book moves and then my opponent fell into an opening trap that I knew and they didn't [1] and I knew exactly how to play for the win to checkmate, because I've done it before many times and remember the post game analysis from them. I didn't come up with the moves on the spot.
[1] I get tons of wins with this one, especially since Queen's Gambit came out on Netflix. https://en.wikipedia.org/wiki/Queen%27s_Gambit_Declined,_Ele...
I had realized a while ago that, as a human, the computer will give your position a score, and then you make a move, and your score can pretty much only stay the same (if you make a "perfect" move) or go down. Much of the time, it goes down.
Two humans playing each other, it's just a question: who's score goes down less each time they make a move? It became a little sad. It seemed like either you can make the right move, or you make a sub-optimal move, and the winner is simply the one who makes the fewer sub-optimal moves.
But when AlphaZero plays, and you watch Stockfish's score, the reason it wins is that it makes moves Stockfish thinks is poor, so it rates AlphaZero's moves poorly, and then all of a sudden it has an oh shit! moment when it realises that AlphaZero is actually ahead, and its score jumps. It's really a look inside the computer's head while it's being beaten by a better player.
That's more or less a description of what happens when two humans are playing over the board.
What you describe is when a perfect chess-playing computer (approx AlphaZero) observes two humans playing. What humans observe watching two humans playing (including the participating humans) is very similar to what you described AlphaZero vs Stockfish as. The only difference is we don't ask human players to ascribe a score to their opponent's move (and wouldn't expect it to be accurate)
This is also something to keep in mind to not get discouraged: Just because every move is terrible to a 3500+ chess engine at some level, it does not mean these concerns always apply to you at half that rating.
Everything else like a positional score or centipawns or even classic material points is an abstraction, that we use to summarize because we don't have unbounded or sufficient computing power to solve all possible continuations. That score apparently going down is only an artifact of our limited ability to evaluate it; the only real scores are 0/½/1 for lose/draw/win. If you make mistakes, your score will evetually drop by those quantizations; we just typically don't know exactly when, except in endgame situations pared down enough to be computationally tractable.
And it's impossible to raise your estimated score, because that estimation assumes you continue to play perfectly. There's no such concept as a better-than-perfect move to raise your expectation over what was already calculated, since that calculation already includes all your best possible moves.
* (Other gradiations between win/lose/draw are possible in such a game. Chess doesn't have such, but imagine playing Go for a dollar per point, where nuances smaller than swinging a win or draw still matter.)
A move that requires serious hardware to defend against is better than one a 10 years old laptop can hold a draw against.
For example I remember the original alphazero model that had been trained specifically against stockfish, would often take a material sacrifice for some advantage that stockfish couldn't see (e.g. sacrifice a pawn, but their bishop gets locked out of the game). I don't know if these moves were objectively good given perfect play, but they could be the only way to win now (chess is very drawish at the top computer level).
No, what I was saying is that it's absolutely possible to raise your estimated score, because your estimated score is only an estimation of who has the best position.
If AlphaZero is better than Stockfish, then by definition it will make moves that sometimes raise its estimated score, because Stockfish is only as good as its ability to estimate the score of a position better. So Stockfish must occasionally underestimate a position, and then later (after another move or two) is forced to reevaluate (because while it's worse, it's not stupid).
AlphaZero wins because, and precisely because, it believes some positions are more favorable than Stockfish does. You can almost see it as an arbitrage between the two estimations. That's what I was finding cool, and the point of my post.
Mostly I'm pointing out that these estimations represent the best guess of an ultimately limited engine. People tend to treat those engine evaluations as actual numbers, like scores in a sport like baseball or some such, but they're not.
What they are describing is that feeling you get when you suddenly realize you are losing even though you are even in pieces, because you suddenly see your opposition has superior positioning and board control
It's not a blunder, because there wasn't one particular move where the game slipped away.
At high levels, blunders are rare.
A move that worsens the position in a way that's not so immediately obvious might be called a 'mistake' or 'inaccuracy'
By these standards, it's quite possible for a game between humans to be won or lost without a blunder on either side.
Hence why I talked about the feeling you get when you realise you have mis-evaluated rather than trying to define it precisely.
From the OP:
> But when AlphaZero plays, and you watch Stockfish's score, the reason it wins is that it makes moves Stockfish thinks is poor, so it rates AlphaZero's moves poorly, and then all of a sudden it has an oh shit! moment when it realises that AlphaZero is actually ahead, and its score jumps. It's really a look inside the computer's head while it's being beaten by a better player.
Sure - you get that feeling from blunders too. But take a look at the Stockfish vs Alpha Zero games - Stockfish doesn't blunder. It it just outplayed: https://www.chess.com/news/view/updated-alphazero-crushes-st...
If you watch from here in game one you can see the "oh shit" moment, as Stockfish's evaluation drops from +1 to even to -1: https://youtu.be/Q5EPqM8gS7k?t=255
Not quite, as that would imply that each game has at most one blunder. It's quite possible for two players to blunder back and forth multiple times.
A mistake is like a blunder but you're still winning, but it could be a blunder in a worse position. An inaccuracy is a bad move that doesn't cost you.
It’s also evaluating your own position vs someone else’s - humans and computers are the same in that both will make the move they think is best, and will only have an oh shit moment when their opponent has provided a reply they didn’t expect.
The only difference is computers can see further, so while an oh shit moment for a human might be 4 moves out, with a computer it might be 20.
$ ./stockfish_14_x64_bmi2
Stockfish 14 by the Stockfish developers (see AUTHORS file)
help
Unknown command: help
?
Unknown command: ?
eat flaming death
Unknown command: eat flaming deathHere's an immediately usable browser version (supports WebAssembly with SIMD instructions):
If you want to experiment with ./stockfish anyway, the protocol documentation is here:
https://www.shredderchess.com/chess-features/uci-universal-c...
edit: also someone mirrored it here (in a more convenient format)
https://gist.github.com/DOBRO/2592c6dad754ba67e6dcaec8c90165...
$ ./stockfish
position startpos
evalhttps://heptonion.net/9ff70753.html
Notice the diagram titled "NNUE derived piece values" underneath the table for "Contributing terms for the classical eval".
Edit: for anyone interested in the hand-tuned evaluator, it might be worth checking out the Stockfish evaluation guide [1]. (The board in the upper-right corner is interactive!)
(I suppose I should have also disclaimed that I'm running from master rather than the release, so it's possible the output's a bit different.)
When you play stockfish at a human level, it is purposely choosing an answer it knows is worse ^.^
I do not see chess engines as good sparring partners, but rather as tools for effective chess exploration.
The nice thing about these newer engines is that they are starting to be useful tools to explore more sacrificial and unbalanced positions which are really fun to get over the board against humans.
Finding these ideas with the help of a computer gives a competitive edge against other humans, as well as helping discover interesting corners in this vast game.
1. Tactics - in middle game there are a lot of possible variations and humans are bad at looking at all the possibilities so they blunder some tactic (eg some move order combination of under 10 moves that leads to a mating attack or significant material advantage). This alone is close to enough to be superhuman.
2. Long term advantage building — a move order that has no clear tactical advantage but puts the player into a better position many moves into the future. This is something that old fashion chess engines used to need very high depth but with nnue stockfish started selecting these kind of move orders even at low depths.
Variations of type 2 are more interesting to people because this is something we can try to learn from and infer general rules (eg instigating with flank pawns without immediate conversion of the attack became more common after alphazero used this in many games to get long term advantage).
At the same time most people analyze their games in browser running on their laptop/mobile device (since most games happen on chess.com or lichess.org) so really they only get low depth stockfish variations.
Without strong ability in (2) it can be hard to interpret SF generated variations since it is still superhuman (eg it will beat a human from a given position) but it sometimes makes moves that are probably suboptimal (because there are no immediate tactics available). But you can’t really tell if a variation is because sf doesn’t know what to do or because there is some hidden tactic/etc. The normal “solution” is if you suspect the position is pivotal to just calculate to very high depth. But that takes time and not something an average player will do.
If you can get reliably to an end game with only a king and a one or maybe two pieces left on each side, then maybe it can just be hand-scripted and bolted onto what you're proposing, but I'm not sure your proposal would consistently get it there.
I did some research (mostly on the online-go.com forums) and joined lichess.org and haven’t looked back since. Superior in every way.
Perhaps I should thank chess.com for being so annoying, if it weren’t for their constant nagging I would probably have stayed on the platform and never discovered lichess.org.
I have fallen in love with chess again over the past few months, and lichess is an absolutely amazing experience. Their mobile clients (iOS/Android) are both fantastic and BS free.
Here is an example for the current world champion: https://www.chessmonitor.com/u/kcc58R9eeGY09ey5Rmoj
How do you unify ratings?
Edit: ah I see you don't, makes sense I guess. Might be interesting to have them both on the same graph even if the y axis is different..
But there are many other features I want to implement first. I'm currently more focused on the statistics part than on the chess.com/lichess relation.
I also found that when clicking through to an opening on your site, it would always say "no games find at this position". Bug?
The openings page list your openings for white and black. If you click on an opening it takes you to the explorer which shows the stats for only one color (white by default). Therefore, if you play an opening for black a lot it will appear in the list of your openings. But when you click on the opening, the explorer will show your stats for white.
There is also another problem, that the detection of openings (on the openings page) respects transpositions [1] while the explorer does not.
Maybe I'll remove the link to the explorer as this seems to cause a lot of confusion...
Chess.com is brilliant, and I use it as well as the less popular (but also brilliant) Lichess.
Lichess is "cooler" because it's non-profit but honestly, comments like that are reminiscent of the childish anti-Microsoft barbs from Linux ideologues that have thankfully declined in recent years.
Pointless but fun question: how does this compare to human players?
Here is the world champion:
https://lichess.org/@/DrNykterstein
If you browse some other people, especially young and new players you can see the improvement over time.
Climbing chess elo is very hard and slow!
Bullet players are already on average a higher percentile of chess players so the elo is skewed
Generally speaking, players playing bullet are more interested in having fun and less interested in gaining elo (which of course is totally fine).
Edit: Just to be clear, I'm talking about online rating
I guess this is different to official rankings and/or elo though, so that was my mistake.
Definitely throw in a no casteiing game as well.
[edit]
I was imperfectly remembering the rules. If it is theoretically impossible (even with blunders on the player who is out of time's part) for the player with time left to win, then it is a draw.
In addition, a player with less than 2 minutes on the clock may request a draw; see Article 10.2 which includes this subsection:
> a. If the arbiter agrees the opponent is making no effort to win the game by normal means, or that it is not possible to win by normal means, then he shall declare the game drawn. Otherwise he shall postpone his decision or reject the claim.
I am not a chess player but I have never heard of referees stopping the game if you don't "give up" in a losing position and have extra time. That sounds ridiculous but I would love to know if it applies in certain tournaments and the reasoning behind it.
Recent instances I saw was an adult was in an almost-lost position with over an hour on the clock while his opponent had 10 minutes. He let his clock run down to nothing and then played quickly before finally let the clock run to zero in a lost (mate in 2) position. He got mocked for this in the local forums.
Also common for a kid to do a blunder and then sit there sad/crying for an hour. You try to encourage them to resign though.
Source: Am Chess Player/Organizer/arbiter.
Also, time limit doesn't seem that different from a power draw limit if you're a computer.
This is pretty questionable in my judgment, actually. TCEC's GPU hardware is 4x Nvidia V100 data center class GPUs, with a pretty powerful processor to boot. A quick search suggests that ONE of these will run you close to $10k, so we're talking about an all-in system worth mid five figures.
Meanwhile, the CPU hardware is pretty dated at this point. They have 4x Intel E5-4669V4, which is from early 2016. It's not easy to find this processor for sale any more (because, again, it's old), but prices seem to run in the $750 - $1500 range if you look on places like Ebay. Meanwhile even on Ebay a V100 is likely to run you $7K+.
I don't know that it's possible to compare "performance" between GPUs and CPUs in a one to one way, but looking at cost, it seems pretty clear that you'd have to spend a lot more to get a system that allows Leela to play at the kind of level you see on TCEC.
Looking at power consumption tells a similar story. Nvidia's data sheet for the V100 shows a maximum power consumption of 250 watts per GPU, so 1000W when running at maximum load (as a chess engine is presumably likely to do). Meanwhile, Intel places the TDP of the E5-4669v4 CPU at 135 watts. Even assuming they're undershooting that by a bit, we're probably talking 600 watts for that system ... on a rather old CPU model.
I'd say it's not a fair comparison. I'm not mad about it, because at the end of the day computer chess tournaments are for entertainment. It's much better if the best neural net programs are competitive with more traditional chess engines, even if by "objective" standards they are weaker.
This is one of the reasons Core War was so intriguing; all the programs battling it out were running on the same hardware, each given an even slice of compute time. To win, you must then find ways to do the same amount of work in less time, while keeping your footprint small.
When the day comes (and I think it will, if our civilization lasts long enough) that a computer finally "solves" chess, it will be a momentous achievement, but ultimately boring.
But why are we comparing used 2021 prices when these hardware weren't purchased in today's market? Especially when GPUs are 1.5-2x MSRP right now. Even very old GPU prices are insane. I recently sold a 980ti near what I purchased it new. 4669v4 MSRP was $7k, so they are not far off. The V100 is pretty dated too as it is from 2017, and doesn't have FP16 which is heavily used by Leela. For this and several other reasons, a single 3090 is actually faster than 4x v100s according to their own bechmarks[1]. A single v100 is approximately equal to a 3080 or 2080ti in performance.
Maybe you should also checkout the CCCC[2] hardware which is even stronger for both: 2x A100 vs 2x AMD EPYC 7H12
[1]: https://docs.google.com/spreadsheets/d/1lGFf6PLGmBUSMan-YP7V...
Stockfish 8 actually won games against Alpha Zero. But Stockfish 11 (which is still classical evaluation engine with no neural net support) totally decimates Stockfish 8, and Stockfish 13 (which uses neural net) totally decimates Stockfish 11. Stockfish 14 just got 30 elo points stronger than 13.
As it stands now, Stockfish 14 is the strongest chess entity humanity has ever seen.
AZ notably got completely tilted once it started losing, doesn't necessarily recognize strange positions you can't normally get into, and doesn't care about its win margin at all.
Claiming that something that uses a neural net trained on hundreds of gigs of data isn't deep learning .. I mean it's possible, I don't know the details.
What is it about now, open vs closed source? Different methods of deep learning and big data fighting? (Both of these are also interesting ofc)
Though definitely not directly comparable, dataset of GPT2-xl is 8 million web-pages. What I mean to say is that this is clearly deep learning.
My point is that having such a huge dataset would not be extremely useful without using a deep neural net (of at least one hidden layer)
It's just over 82,000 parameters.[1] That's a very shallow, small NN - by comparison something like EfficientNet-B1[2] is 7.8M parameters, and that's considered a small network.
[1] https://www.chessprogramming.org/Stockfish_NNUE#NNUE_Structu...
This isn't true. The size of the training data doesn't imply anything about the size of the neural network.
In the case of Stockfish, the NN is quite shallow, and implemented using a custom framework designed to to run fast on CPUs.
See https://news.ycombinator.com/item?id=26746160 for previous commentary on this.
> Though definitely not directly comparable, dataset of GPT2-xl is 8 million web-pages.
This is irrelevant. You can train GPT3 on a smaller dataset, or a smaller model on the same dataset as GPT3.
> What I mean to say is that this is clearly deep learning.
It's been clear that neural network models are superior since Alpha Go. There's not "Deep Learning vs <something else>" anymore because the <something else> isn't competitive and no one is really working on it.
https://github.com/glinscott/nnue-pytorch/blob/master/docs/n...
It's pretty interesting read.