How some common material imbalances affect your win-rate
lichess.org
lichess.org
I’ve always been taught the two rooks are better and who can argue with 5 + 5 > 9? But also, I’ve also lost almost every game where I’ve had the rooks. I always thought that I just needed to be a better player to take advantage of it. Glad to know it wasn’t just me.
Just goes to show that these shortcuts, like the point system, are only heuristics, and pretty shallow ones at that. Knowing that being up a bishop gives you slightly more of an advantage than a knight is better than nothing. But learning in which sorts of positions a knight is actually better than a bishop will give you a much deeper understanding (and correspondingly more wins).
The queen can be good as a single piece whereas the rooks have to coordinate, which is hard when there's a queen on the board ready to fork. Often the rooks end up having to defend eachother, and if this happens suboptimally they can become very immobilised on a useless file or rank whereas the queen can fly around the board attacking stuff at will.
Like you said, material imbalance is very difficult to evaluate and understand. I generally recommend not to incorporate them into one's play until ~1700 FIDE elo. Though sometimes you're forced into it of course.
It's also a matter of style. I personally love materially imbalanced positions, and I'm pretty good at them, so I incorporate a fair bit of exchange sacrifices into my play. Because then I often get positions I understand better than my opponent and that makes up for the material on its own, I find.
But other players just prefer a different style of play and maybe shouldn't go for it even if objectively it's the best move, because they'll end up misplaying the position.
And in fact most of the wins I remember the most clearly involved some kind of well-found sacrifice. The others tend to blend together too much in my mind. And most of the losses I remember most clearly involved horrible blunders.
2) Use StockFish to evaluate all the positions
3) Pick a position at random, show it to the user, and ask for their evaluation
4) User gets a score based on the difference b/t their eval and StockFish's.
The idea being that this could allow you to rapidly hone your position-evaluation skill. That might be a faster way to improve at chess than just grinding through many games.
But it takes far more to build something useful for humans than just asking stockfish to give you the best move.
For one thing you don't want to learn to play like machines. Often the moves that machines make are terrible human moves. Machines walk into extremely dangerous positions all the time; positions where you need to make a series of only moves to gain a small advantage. Humans shouldn't be doing this.
Machines also rely on tactics heavily. Unless you understand the tactics in play it will make your play much worse. It might appear to you take a piece can't be taken, but stockfish says it can. Ok? Unless you understand the reason why, you aren't going to gain anything from this exercise.
Now if we assume there is often a slightly better player in these games, the better player will more likely get into an advantageous position early in the game and win more often, but not entirely because they got into the more advantageous position. What I'm trying to say is that someone who is up a rook will win partially because they're up a rook and partially because they are likely a better player in the first place and will continue to play better than their opponent.
I think to do this study correctly you'd need to place new players in a random position drawn from the dataset and actually evaluate the win rate without the confounding factor of having gotten into that position in the first place.
I think that's true at 2000+ but at my rating (500) it's just a comedy of blunders
Interesting hypothesis, but my intuition is that this effect is small. It should be easy to measure. Just take random matches, and study winrate when a player has beaten their opponent once before.
The problem is that this doesn't really average out with more samples.
But I don't see how that wouldn't average out over lots of games?
If I have time later I'll try to whip up a small coin-tossing demo in Python that demonstrates what I mean, I'm definitely not the best explainer!
My gut feel is that the effect is minimal at best, but I don't have an argument for that.
[1] - https://en.wikipedia.org/wiki/King_and_pawn_versus_king_endg...
Nice example of more extreme imbalance is that sometimes King+Queen vs King+Pawn is drawn - that is when pawn is on seventh rank of bishop or rook file with defending king in front of it, while attacking king is not close enough to force checkmate.
See: https://en.wikipedia.org/wiki/Queen_versus_pawn_endgame#Quee...
I gained a lot of rating once I started really studying the ins and outs of all the different rook and pawn vs rook endings. Because no matter what the tablebase might say, you have to be very precise with or without the pawn.
So I started getting a lot of wins out of equal positions, and got a lot better at saving a draw in bad positions.
Pawn and minor piece vs minor piece is also fraught with drawing chances, especially with bishops. Because you can just sac the piece for the pawn and there's insufficient material for mate.
Queen and pawn vs queen is basically a 3 result game at the club level. So many ways to blunder your queen. And at GM level it tends towards a draw I think. Though it's so stupidly complicated I never even tried to properly study it.
Oh, so it's not just me.
Of course, this isn't actually true. But it's a useful rule of thumb to just assume it's a draw unless you're sure it's a win, because it probably is, fucking somehow.
Two rooks turn out to be significantly worse than a queen instead of slightly better. But the most surprising thing to me is that having the two bishops seems to be worth almost nothing -- close to 50% odds across skill level. A single pawn advantage is more valuable! (The article says "4-5% more likely to win" at Elo 1200-1400, but that doesn't match the graph.) These surprising results were also more consistent across skill level, while the well-known advantages are worth significantly more to skilled players.
I don’t think so, and I would like to see the author look in the data for 2 knights vs 2 bishops. My gut feeling is that knights are hard to block and can attack every square, while bishops can only attack their color, so half the board. They can’t protect each other and rely on pawns for this.
I’m happy to trade my knights for the pawns, but I would like to see if I’m wrong.
Generally the trade off is a bishop pair, where you do have all the squares covered, vs the slower knights, who are also the only pieces that can threaten a piece without being threatened back.
The advantages are highly position dependent, with closed games largely going to the knights and open games going to the bishops pair.
But most games turn into open games at some point, so statistically the bishops end up better.
Though
"Even at lower Elo ranges like 1200-1400, you're 4-5% more likely to win if you have the only bishop pair. 1200 Elo players don't know how to take advantage of having the bishop pair, and yet it helps them win all the same [...] A 1000-1200 Elo player is only about 8% more likely to win when up two minors for a rook."
But this is not what you're testing. You're not determining the effect of imbalance alone, but imbalance + particular position of pieces. To determine the effect of imbalance you should study how does not having a certain piece since the beginning affect the result of the game.
Also, being a pawn up in the opening is much more different than being a pawn up in the end game. It would be interesting to redo this analysis but splitting by opening, middle game, and end game.
How exactly does one take advantage of having a bishop pair? Why is that not obvious to players under 1200?
(I used to play computer correspondence chess.)
Edit: I should add that +/-1.0 is roughly a pawn's worth of material. So a computer up a knight or bishop -- nominally 3 points of material -- should virtually always win.
You can of course give up material for positional advantages as well, but from the article I don't think the author analyzed that. It would be difficult to accurately measure that anyway, some positions can be up +5 points according to the computer but only if you find 20 perfect moves in a row. Needless to say, most humans would not find those moves especially in the lower ELO brackets that the article analyzed.
Knowing how the word is used in chess is not enough to know where the line was drawn in the data here, which is the basis for the entire analysis shown.
In simple mathematical terms a clean material advantage is:
Overall eval >= material eval
If overall eval < material eval, then compensation(other player) > 0 = "the opponent has some compensation for the pawn"
I really doubt if there was any attempt to check if there's some "strategic compensation" - how would you do that at a large scale? I doubt that even running a solid chess engine evaluation on all these positions is feasible, you need something where you can simply/cheaply filter positions from the database and then just count the winrate.
Hence it should be clear to anyone that without engine assistance it is totally impossible to determine for 10^6 games if a certain material difference was "clean" or not. Most likely he just checked for a stability over n consecutive moves, at least this is the usual way. Others have noted more problems, that might help put his highly deviating, washed out results in order.
> The Elo rating system is a method for calculating the relative skill levels of players in zero-sum games such as chess. It is named after its creator Arpad Elo, a Hungarian-American physics professor.
https://en.m.wikipedia.org/wiki/Elo_rating_system
There are other rating systems that are similar, but people commonly call any similar rating system "Elo". In this case, Lichess uses https://en.m.wikipedia.org/wiki/Glicko_rating_system, so the GP is arguing against calling it "Elo". I think it's kind of a pedantic point given that more people will understand "Elo" than "glicko-2", but I suppose "rating" would be clearer than either term.
Also, it's "Elo" not "ELO".
Thanks for giving me a good start to the morning!