What are Magnus Carlsen's chances of reaching 2900?
new.chess24.com
new.chess24.com
Thus, Carlsen's best chances largely depend on elo inflation to widen the elo distribution itself.
One example of this: Louis de Bourdonnais which was #1 in the world in the 1800s skyrocketed his elo rating from 2600 to 2650 mostly as a result of the following 10 players moving their average elo from 2350 to 2450.
His winrate against them never changed. And what is required for 2350 players to go higher? That the players below they skill also raise their elos giving the top 10 more points when beating them. The more players will be fide rated the more Carslen's chances will increase but without a sufficient supply of 2700s, those at 2750 won't raise in elo, etc, etc
A player with a rating of 2750 today might have beat Kasparov in his prime, simply because chess has evolved.
In a way, it's similar to sports. Plenty of people today might have beat Carl Lewis at 100m 30 years ago, but they would probably not be considered greater athletes.
So in order to beat 2900, it would help if several other players got rankings of 2800-2850, but only if they did so without actually playing much better chess.
Or he could just have a temporary lucky/confident streak, and gain 40 rating points in a few tournaments.
Same with Lance Armstrong...
The fact that chess has evolved is irrelevant because it applies to all players from carlsen to beginners everyone has the same tools.
Why are you assuming that Carlsen would always win against other players if the other players are that much stronger? Maybe if the following 25 players were all 2800+ then Carlsen would take more losses.
This case is somewhat strange because he's not the usual mathematical crank; he's actually a quite well-recognized economics researcher who is clearly mathematically competent. But for some reason he persists. (If you search his name, you can find his YouTube channel where he lectures on his RH work...)
"A Hadamard factorization of the Riemann Xi-function is constructed to characterize the zeros of the zeta function."
Like, dude, you are claiming to solve a Millennium problem, you're gonna have to put a little more effort into explaining, ya know, the breakthrough...
We care about the Riemann Hypothesis not for the theorem itself, but because the methods used to solve it will surely be revolutionary and it's those methods that are important, not the theorem.
That's what makes it a Millennium problem. So a paper like this immediately shows the author doesn't even understand why the problem is important, let alone its solution.
In the abstract? If papers needed to cover everything in the abstract they would need to have an abstract for the abstract and they you’d just end up complaining about the abstract abstract not being complete…
Or you can add, oh I dunno, two or three more sentences to maybe hint at what is novel and important about the approach?
Sorry, there's nothing holy about the Riemann Hypothesis; anyone can add their input to the problem, especially on ArXiV and/or YouTube. You know, that's how discussions can start, and sometimes people want to discuss things with others.
It seems like you might think only certain people should work on these special problems, and only if they do it correctly according to you.
Your comment is just an ad hominem, so why don't you tell us why you're actually mentioning this? Do you think the math is wrong in the relevant material from the article because of his RH-related musings on ArXiV, or do you have an anti-RH-researcher bias that you want to share with the world? Are you just hating on someone?
You have an interesting idea? Go ahead, please share it! But it's not productive to just start with "Look everybody, I solved it!"
It just wastes everyone's time. For example, this paper. It would take a lot of time to sit down and work through the math until I find a specific error in it. But the barren abstract and reference sections strongly suggest that would be a huge waste of my time and energy.
Mochizuki pulled this stunt with the ABC conjecture. Wasted YEARS of mathematicians' time just to arrive at the conclusion it was all elaborate mathematical smoke and mirrors. The guy is an egotistical asshole with a messiah complex, and managed to burn a lot of PhD students by chasing his red herring. Not cool.
This guy though, knows what the right way to present the idea would be, but for some reason (ego? knowing it's weak/broken? pique?) chose not to.
You've effectively said that people can't post things on ArXiV unless they're up to your unstated standards; otherwise, they're just "cranks".
Also, no one is "bypassing peer review" by posting on ArXiV and/or YouTube, nor have I seen anyone claim "a virtue" of any sort. Where are you getting all this? From the abstracts?
> It just wastes everyone's time. For example, this paper. It would take a lot of time to sit down and work through the math until I find a specific error in it. But the barren abstract and reference sections strongly suggest that would be a huge waste of my time and energy.
Well, since you've now admitted that you haven't read his work, it seems like everything you said earlier really must be coming from the abstracts alone.
> It just wastes everyone's time. For example, this paper. It would take a lot of time to sit down and work through the math until I find a specific error in it. But the barren abstract and reference sections strongly suggest that would be a huge waste of my time and energy.
Who's time is wasted? The random people who volunteer to read his ArXiV submissions and/or YouTube videos? Really, the only waste of time I've seen is your ad hominem comment in a HN post that isn't even about the author you're blatantly criticizing.
> Mochizuki pulled this stunt with the ABC conjecture. Wasted YEARS of mathematicians' time just to arrive at the conclusion it was all elaborate mathematical smoke and mirrors. The guy is an egotistical asshole with a messiah complex, and managed to burn a lot of PhD students by chasing his red herring. Not cool.
And that's why we must all attack Mochizuki whenever we see his name, right? Do you know Nick? Is he egotistical? Does he have a messiah complex? Are you just going on a tangent now? Are you just math trolling?
It is bad practice to post incorrect results and not retract them when this is pointed out.
And it's good practice to take shots at people whenever you see their name?
Also, no, there's no "practice" that says you can't keep a mistake posted on ArXiV or YouTube. That just sounds like something you've made up to justify attacking someone.
And who's pointing this out to him?
Not trolling, I am serious. It took Peter freakin Scholze to finally settle the debate over Mochizuki's work. What came out of it? What else could Scholze have been working on instead of spending time finding the incredibly subtle gaps in logic that he used to construct his false theory?
What Mochizuki did was wrong is because he was utterly uncooperative with the mathematics community, he would answer questions only with more papers that never addressed the concerns being raised, and he was offended that so much scrutiny was applied.
Has Nick made an effort to educate people on the incredible breakthrough he has made? How many lectures has he given on it? Any conference videos on YouTube? Because I'll watch them. Mochi is an extreme example, so I'm sure he's not trying to do anything wrong. He should just retract the paper and keep working on it. An interested volunteer can tell him what is specifically wrong, and that's all that needs to happen.
[0] Not sure if 'types' is the right word? Using it colloquially. Maybe 'classes' would be more accurate?
As an aside, this is also perhaps also one of the best motivating examples of the idea of "oracles" in math. ("Ok, assume we can do $IMPOSSIBLE_THING... now what?" :) )
Unfortunately, if you aren't already part of the in-crowd, getting peer review at all can be very difficult, if you have any type of unorthodox idea. Because, you guessed it, you get labelled as a "crank", so why bother wasting time reviewing your paper?
If you have valid criticisms of something, you can make them in public and help everyone involved grow.
I certainly wouldn't like it if after submitting a blog post, the top comment on HN is about me being wrong 5 years ago.
We all make mistakes, and no one would really care if he admitted this. What makes the case notable is that he has enough mathematical training that he should be able recognize the proof is wrong. Especially after the errors are uncovered by others and communicated to him. Continuing to assert the proof is correct after errors have been found is bizarre.
You know that ArXiV and YouTube aren't considered "journals" and don't have a similar requirement for "retractions". It sounds like you think he should hide his shame; otherwise he deserves your attacks.
> We all make mistakes, and no one would really care if he admitted this. What makes the case notable is that he has enough mathematical training that he should really be able recognize his proof is wrong, especially after the errors are uncovered by others. Continuing to assert the proof is correct after errors have been found is bizarre.
People probably shouldn't, and don't, care because he's not affecting them in any way with his RH-related ArXiV and YouTube musing.
More importantly, where are there people showing him how his proof is wrong, and where is he outright denying their points/proofs? It seems like I'm missing a link or two, because I haven't seen any of these things, yet you're referring to them as though they're apparent and damning.
My guess is that he simply doesn't get feedback about this stuff, and he probably doesn't even care that much, because this is just a set of ideas with which he likes to work. I've seen no signs of a charged or "high stakes" math community engagement, and definitely not the kind that deserves these unwarranted personal criticisms.
At the very least, you're raising ArXiV and YouTube to a standard they openly do not meet.
Also, did you raise these points to him and he denied your evidence/proofs? Do you know someone who did and you're speaking for them?
It against the prevailing norms of scholarly communication to publish results with serious errors known to the author.
I agree YouTube does not have this policy. I still find incorrectly claiming a proof of RH on YouTube distasteful, for similar reasons.
I personally know a mathematician who has pointed out the mistakes to him. But also, Polson good enough at math himself that he should be aware of these points. I would criticize him just the same if he knowingly published false economics results.
If you're implying that he violated ArXiV's submission policy, then you're going to need to stretch the definitions of those words a bit, as you attempted to do.
> I agree YouTube does not have this policy. I still find incorrectly claiming a proof of RH on YouTube distasteful, for similar reasons.
I get that you have all these personal opinions/takes, but ArXiV and YouTube don't appear to be justifying them.
> I personally know a mathematician who has pointed out the mistakes to him. But also, Polson good enough at math himself that he should be aware of these points. I would criticize him just the same if he knowingly published false economics results.
OK, so where did they post these discussions with Nick so that we can all clearly see that he's a bad or stupid person, as you're implying? Your comments assume that this is all common knowledge or apparent, but it isn't.
Are we just supposed to take your word for it? Well, I happen to know Nick, and, from my experiences with him, I have no reason to believe any of the things you're saying and/or implying.
Norms of scholarly communication are not a personal opinion.
You have also curiously avoided my question about whether it is improper to knowingly disseminate incorrect results.
My guess is that he simply doesn't get feedback about this stuff, and he probably doesn't even care that much, because this is just a set of ideas with which he likes to work.
Having every right to be a jackass doesn't make you not a jackass.
I view the whole thing more as a criticism of the sorry state of FIDE and the world championship that an actual serious goal.
Pretty much every record in every sport is arbitrary, yet people like to set goals for motivation.
I wonder if it is a valid strategy if he drops out of World Championship title contests to allow others to increase their score by not losing against him. I would have thought that wouldn't move the needle much but I guess he needs everything there is to get to 2900.
There are going to be team compositions that didn't win an NBA championship that are stronger than ones that did win a team composition simply because at the time the competition was stronger and so they weren't the best. So, if you use #ofChampionships to generate a ranking of all team compositions there's going to be a lot of debate that the teams aren't actually ranked by strength since some of the teams that racked up championships did so when other teams sucked (I know less about NBA but for NHL's original 6 this becomes the case, the competition was so weak that some people didn't know they were drafted because it wasn't something you cared about).
To get back to the original point, your ELO rating is in relative to your peers. The idea is that for a given point differential you have a x% chance of beating them. So the point being made is that you can't use ELO to compare anything besides two players who can physically can play each other as taking a playing in the past with n elo and having them play a current player with n elo won't match the x% chance.
To counter some of the other points (bowling and the mile). These are precise rankings not relative. If somebody runs a 4min mile in the past you can expect them to run a 4min mile in the future. This doesn't mean that they'll win the same %of races though as they did which is what elo is more about. Same thing with bowling, if somebody constantly bowled a 280 in the past you can expect them (ignoring lane/oil changes) to bowl 280 in the future but if the competition now bowls a 290 that player is going to suck instead of being a legend. So the bowler of the past may have had a very high ELO but now when taken into the future will (after many losses) have a lower ELO despite having the same peformance.
Federer getting grand slams - all against local in time competitors - is impressive. Had he played 50 years ago, he'd have likely gotten more. Had he played in 50 years, likely less, expecting competition to improve.
Same for NBA records - all against a moving, relative backdrop.
Same for wrestling, baseball, hockey, and on and on and on.
Just because your score is relative to others does not make reaching some milestone, especially this one, less impressive.
Also, it's easier to have a higher ELO if others do too, not the other way around.
This doesn't seem necessarily true to me. If Magnus was in his prime in say 1980s it's not clear if he could ever get so close to 2900 since every other super-GM had lower ELO which would make his progress slower than now. Not only do we not know how an average 2022-super-GM compares to a 1980-super-GM, we also don't know how Magnus compares to an 1980-super-GM. E.g. Tal achieved his peak ELO 2705 in 1980, but at that time Korchnoi had an ELO of 2695 (diff of 10), whereas, Magnus currently has 2864 and #2 Ding Liren has 2808 (diff of 56). But Magnus' peak rating of 2872 is 30 less than Caruana's peak rating of 2842 during the same time.
I suspect this is false. Chess is one of the few endeavors where we have decent historical records. Combine this with the fact that computers can now analyze lines a very long way, and I suspect that you can do things like compare how often GMs missed a superior line of play.
By most measures, the GMs of the past played weaker chess than our current GMs. The current GMs benefit from having started at younger ages with a larger body of chess theory and super-strong computers to work with and against. I suspect that the GMs of the past would be as strong as the GMs of the present given those. However, they didn't have those and, consequently, the chess, itself, was weaker.
Side note: I adore Mikhail Tal's "This position looks interesting and complex. Banzai!" games. However, most of them would get him absolutely creamed against the current top players. And to be fair--even the top players of his own time tended to refute them. However, Tal's games are often way more exciting than everybody else's.
It's still impressive as heck.
One tempting comparison is Usain Bolt, but it feels a little different to me because he dominates both objectively (WR) and versus his competitors.
The decathlon works the same way- the points scoring there isn't relative like Elo is, it's done by performance in an individual event. You don't get any extra points for winning a given decathlon event, there's a scoring scale for your marks regardless of position.
The exact same performance will differ in how it affects two players' Elo depending on both players ratings. If all players' ratings were to be 100 points higher with all else equal, the top players ratings would rise as well. If everyone in my Mile runs 10 seconds faster, my time doesn't also get 10 seconds faster.
Basically, one record is an individual performance mark. The other is a ranking, which is a different kind of metric entirely.
With all that in mind, chasing a world record Elo is still a worthy goal IMO. The competitive chess scene is old and robust enough that it remains a meaningful benchmark, even if it is a relative measure.
And the ambition to reach unprecedented strength in a sport is emphatically not an 'unserious' or 'poor' goal: it's the motivation of all sporting champions.
2900 might seem arbitrary, but it's a symbolic level under the rating system we have.
But it seems ratings have been deflating lately (https://en.chessbase.com/post/the-elo-ratings-inflation-or-d... is a fun read), so hitting 2900 would be even more impressive!
Far far below the lofty heights of Carlsen &co., I play a fair bit on Lichess. Sometimes I'll play someone with a similar ranking to myself and beat them and will earn 3 or 4 points. Then we'll have a re-match which I lose and they'll gain 12 or 15 points.
I just can't fathom how them beating me earns them more points than I earn from beating them --if we were on the same or similar rankings to start with. But it almost always seems to pan out this way.
[And I'm not talking about opponents with a '?' after their ranking, which signifies they've not played enough games for the ranking to be accurate yet. I know that, in those circumstances, the rankings can move hugely in either direction].
I've also noticed that there doesn't seem to be any relation between the circumstances of a win and the points awarded. For example, shouldn't a win or draw when playing black earn more points than when playing white --given that white had the advantage of moving first?
And what about the manner of victory? It seems there's no more ranking points to be gained from; snatching a victory where both players ended up down to a single pawn each and the winner just managed to queen one square ahead of the loser... than there are from completely decimating your opponent.
So it could be that your opponent has played significantly less games than you. However if he loses he also should lose more points than you would.
The color you play has no influence on the rating numbers. And also the type of win/draw/loss does not matter at all. Ratings are just statistics of your results.
Lichess works on a different system than professional ratings. It takes into account the RD (sort of a standard deviation on your rating). Playing against a player with a more stable rating will cause the rating to change more and vice-versa. Professional ratings are always assumed to be stable.
> I've also noticed that there doesn't seem to be any relation between the circumstances of a win and the points awarded. For example, shouldn't a win or draw when playing black earn more points than when playing white --given that white had the advantage of moving first?
This is something that is sometimes used as a tiebreaker in tournaments. Sort of an away goals rule, if you're familiar with soccer. But as far as ratings go, in the long run you're bound to play ~50% of the games with white and ~50% with black, so it smooths out itself.
> And what about the manner of victory? It seems there's no more ranking points to be gained from; snatching a victory where both players ended up down to a single pawn each and the winner just managed to queen one square ahead of the loser... than there are from completely decimating your opponent.
I'm sorry but this makes absolutely no sense. There are three results: Win - Loss - Draw, everything else is just discourse around the game.
Endgames are something players trade into, reducing to a won endgame is part of the skill involved in the game. Why would that count as less of a win?
Also, material is a pretty stupid way to determine how close a game was. Should I get less rating points because I won with a queen sac?
About white versus black, there is some consensus that after about 20 moves, white's advantage is gone. I doubt that statement is supported by statistics. I guess the rating just assumes everyone plays white and black just as often, so it will even out in the end. (edit) By the way, if you are strong with white and weak with black, how should your ELO rating be calculated? It is only one number, not two. Assuming an average seems fine with me.
And the way a win is made, those are for the beauty contests :)
A simple calculation for ELO is just to see it as both players placing a bet, the K factor is the amount of points in the jar. If the K factor is 25, you have 1600 ELO and I have 1400 ELO, you might put in the jar 16 points, while I put in 9 points. Winner takes the jar, on a draw it is plit in half. Difference between win and draw is always half the K factor. When both players have a different K factor, the calculation for both players will be different, it's not zero sum.
That statement is supported by statistics, however, the problem with that statement is that the average chess game is 25 moves.
except in top-level classical play.
But it's not relevant, because virtually all experts agree that chess is a theoretical draw. It's extremely difficult to prove it, however.
depends on the position, from not a lot, to 100+ moves.
but don't forget that there are forced draws in chess, namely, repeated position 3 times (doesn't have to be consecutive), 50 move rule (50 moves without a pawn move or a capture), and insufficient material.
So trading off material would shorten the game.
That explains some odditites I've observed, like different strength with the same rating at different hours of the day or different days.
Because if you count to be losing by five points you could either resign (I'm assuming a game between strong players) or start playing more adventurous moves to turn the game around. Adventurous moves usually make for a larger win, let's say fifty points. Should players be punished for that and only play games that minimize the difference between the scores of the players? How boring. I guess this applies also to chess. Sacrifice, sacrifice, win. Or sacrifice, sacrifice, mistake, loss. I won't punish a mistake so much more than any other one.
It is likely that the person with 12/15 point gains has less games, thus his delta variation is higher than yours. It takes tens of games before you take "normal" points.
It's probably a sign that they have a new account. If a player has played very few games, then his rating is only approximate. It's likely his true rating is far from his current rating. The rating added or subtracted due to a win/loss is on purpose made large, so he quickly can get to his true rating. When he gets to his true rating, he will presumably win and lose the same amount of games and stay at the rating.
Lichess and Chess.com both use (variants of) Glicko (instead of ELO). The first step is to calculate the rating deviance. A high deviance means that the player hasn't settled on a given rating. This factors into the calculation of the new rating.
The sites have chosen this system so that higher (or lower) rated players that make a new account quickly get close to their "true" rating.
In Elo, the rating gain for one player equals the rating loss for another, because each player's rating is the only independent variable.
Glicko calculates a kind of 'certainty', aka 'rating deviation' which is based on the number of games played in some period of time. The idea is that the ratings of players who have not played in a while are more likely to be wrong.
Quote from the article: "...won every tournament he took part in, scored 32 wins, 47 draws and not a single loss, with a rating performance of 2889."
Also, "the same performance" has a different meaning from TPR performance I guess. With a TPR of 2889 you will not reach 2900, afaik. Unless 1 tournament is played with a TPR of 2920 and the next one with a TPR of 2840. But then it's a fleeting 2900 :) I guess that will do for Carlsen :)
It seems kind of unfair that this is the strategy needed to increase ELO. If the safest strategy to guarantee a win in a tournament is winning a single game and drawing the rest, then you shouldn't be penalized for that in your ELO, versus playing risky and winning more but also losing some.
Maximizing your rating on the other hand will require you to play more and more risky games.
Every game was drawn until Kramnik randomly blundered a trivial mate-in-1 in a simple position. That mistake was undoubtedly driven by psychological reasons. Playing against an engine is nothing like playing against a human and it's difficult if not impossible to put oneself in the proper mindset to play well.
The following games were then also drawn until Kramnik decided to do the chess equivalent (when playing against a computer) of going for a hail mary, with black, in the final game to try to even the match score. Suffice to say, that failed and the match ended as a generally disappointing, for everybody, 4-2.
There have been a variety of gimmicky matches since (fast time controls, various odds matches, etc), but nothing serious.
Stockfish: 3585 Elo
"Elo suggested scaling ratings so that a difference of 200 rating points in chess would mean that the stronger player has an expected score (which basically is an expected average score) of approximately 0.75, and the USCF initially aimed for an average club player to have a rating of 1500."
I guess that means that Magnus has expected score of roughly 0.25^((3585 - 2864)/200) = 0.00675 against Stockfish 15, which is basically 1 in 200 games?
Computers are definitely much stronger than humans, but not 3600 better. Magnus would certainly be able to eek out plenty of draws, if not only because white can create "simplified" (as a euphemism for dead) positions in just about any variation if he really wants. And Magnus regularly plays these sort of positions literally at the level of supercomputers.
I'd also add that much of the dominance of computers is not based just on raw ability alone, but more psychological issues. Humans can become tilted, intimidated, frustrated, tired, and so on. One of the last major human vs computer events was Kramnik vs Fritz. Kramnik, in a relatively simple position, ended up blundering mate in 1 with plenty of time on his clock. It's unlikely he would have ever made the same mistake against a human. It's just very difficult to get in the same mindset when playing against a human as when playing against a computer. Chess, in spite of being a game of complete information, is still extremely influenced by psychology.
Instead, you should convert the 0.25 to "odds" form. 0.25 is 1:3 odds, represented by the number 1/3. (1/3)^((3585 - 2864)/200) is about 0.01905 (still in odds form). To convert this back to an expected score you would take 0.01905 / (1 + 0.01905) = 0.0187. So Magnus Carlsen's expected score is 0.0187.
Applying the same method to Stockfish, we have 3:1 odds, which is represented by the number 3. 3^((3585 - 2864)/200) is about 52.48. Converting back to expected score we get 52.48 / (1 + 52.48) = 0.9813. So Stockfish's expected score is 0.9813.
Our sanity check is to add 0.0187 + 0.9813. The result is 1.0, as it should be.
This seems like a big issue. There are currently a bunch of really promising young players coming up the ranks who will likely be continuing to improve and be underrated for a while. That's going to make it even harder.
There is a cringey announcement of the takeover on YouTube, including a TikTok video of Danny Rentsch dancing (to put it mildly ...).
The plan seems to be to draw Carlsen into that toxic sphere of influence. Chess streaming and the culture surrounding chess.com seem to be for the lowest common denominator.
Carlsen fit much better into the European culture of chess24, I hope he can ignore the circus and it won't affect his play.
If the top players are not winning their games, they will not stay at the top for long.
Edit : He can try to game the ratings system by looking for high-rated opponents who he knows will not play as well as their ratings suggest for some reason (like illness or personal issues). But he doesn't seem the sort to enjoy this
It will also not affect your FIDE rating if the tournament is not organised to the required standard (which is basically nonexistent for under-1800 rated players, and going all the way to the required presence of International Arbiters, metal detectors and blood tests for top level play)
What are they looking for? Communicators and ... Adderall? Xanzolam?
Probably not Xanzolam, because there's no way it'd enhance your performance, but yeah, you get the idea.
what else do they check for? caffeine is probably helpful and not banned. LSD (microdosing)? Galantamine? If there are nootropics that can legitimately enhance high level chess play why aren't we using them on our top scientists and engineers?
Tl;dr: They use the world anti-doping agency list.
Does this mean we've basically solved chess at this point?
TCEC shows opening lines which allow for AIs to best each other, maybe tournaments need to move to that eventually. It's what they've done in checkers
Could he game the system by doing 1000 rated matches against poor players until he hits 2900?
In theory he plays one 2500 player every month, and wins, for 39 months in a row and we're done? I'd be impressed if he won all of them (he doesn't usually do that.)
His most recent game on 2700chess is a draw vs 2490 rated player Schitco.