Here is the convo:
me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed
ChatGPT: Alright, I'm ready to play! Please give me the move sequence.
me: 1. e3 Nf6 2. f4 d6 3. e4
ChatGPT: My next move as black would be 3... e5
Completely ignoring the hanging pawn.This is not the play of a 1400 elo player. It is the play of something predicting patterns.
I ran a bunch of experiments in the past where I played normal moves and ChatGPT does respond extraordinarily well. With the right prompts and sequences you can get it to play like a strong grandmaster. But it is a "trick" you are getting it to perform by choosing good data and prompts. It is impressive but it is not doing what is claimed by the article.
ChatGPT is in no way 1400, or even close to it. The fact this article gets upvoted around here is proof that people aren't thinking clearly about this stuff. It's trivially easy to prove it wrong. Live unbelievably so, I tried the same prompt and within 12 moves it made multiple ridiculous errors I never would, and then an illegal move.
Keep in mind a 1400 level player would need to basically make 0 mistakes that bad in a typical game, and further would need to play 30-50 moves in that fashion, with the final moves being some of the most important and hard to do. There's just no way it's even close, my guess would be even if you correct it's many errors, it's something like ~200 ELO. Pure FUD.
The author of this article is cashing in the hype and I'm wondering how they even got the results they did.
Keep poking it and criticizing it. Microsoft and OpenAI are on HN and they're listening. They'd find nothing more salient to tout full chess support in their next release or press conference.
With zero effort the thing understands uber domain specific chess notation and the human prompt to play a game. To think it stops here is wild.
People are hyping it because they want to get involved. They want to see the crazy and exciting future this leads to.
Some future AI might, but a language model won't.
Claim: "ChatGPT's Chess Elo is 1400"
Reality: ChatGPT gives illegal moves (this happened to article author too), something a 1400 ranked player would never do
Result: ChatGPT's rank is not 1400.
Explain your thought process here further if you don't mind.
But even if it doesn't play like human 1400 players, if it can get to a 1400 elo while resigning games it makes illegal moves on, that seems 1400 level to me. And i bet that some 1400s do occasionally make illegal moves (missing pins) while playing otb
That's not to say ChatGPT can play at 1400, just that that playing in an odd way doesn't determine its rating.
It's a (theoretically) 1400 player which plays significantly better then 1400 when it knows the lines, but makes bad or illegal moves when it doesn't, and that play averages out to be around your typical 1400 player. Functionally is just what a 1400 player already is, but with higher extremes and lower lows.
Understanding this concept is crucial for getting good results out of large language models.
The fact that rules and articles exist describing what to do if you or your opponent makes an illegal move indicates this is not the case.
Humans are also... human. They make mistakes. It may not happen often at 1400, but to say that it'll never happen is preposterous.
The bar isn’t “I didn’t make an illegal move this morning” it’s “something a 1400 ranked player would never do”.
My entire point is that it happens. Not often, but also not “never”.
Because all I'm hearing is talk about ChatGPT's abilities as a reply to me calling out an extreme statement as being extreme. Something the parent comment even admitted as being overly black and white.
Edit: Without reading everything again, I'll assume someone said "never." They're probably assuming the reader understands that "never" really means "with an infinitesimal probability," since we're talking about humans. If you're trying to argue that "some 1400 player has made an illegal move at some point," then I agree with that statement, and I also think it's irrelevant since the frequency of illegal moves made by ChatGPT compared to the frequency of illegal moves made by a 1400 rated player is many orders of magnitudes higher.
> something a 1400 ranked player would never do
> fine, fair, "never" was too much.
I mean, yes they were and they said as much after I called them out on it. But go off on how nobody is arguing the literal thing that was being argued.
It's not like messages are threaded or something, and read top-down. You would have 100% had to read the comment I replied to first.
> He literally used the same prompt as the article. > Claim: "ChatGPT's Chess Elo is 1400"
> Reality: ChatGPT gives illegal moves (this happened to article author too),
> something a 1400 ranked player would never do
> Result: ChatGPT's rank is not 1400.
This is a completely fair argument that makes perfect sense to anyone with knowledge of competitive chess. I have never seen a 1400 make an illegal move. He probably hasn't either. Your point is literally correct in the sense that at some point in history a 1400 rated player has made an illegal move, but it completely misses the point of his argument: ChatGPT makes illegal moves at such an astronomically high rate that it wouldn't even be allowed to even play competitively, hence it cannot be accurately assessed at 1400 rating.
Imagine you made a bot that spewed random letters and said "My bot writes English as well as a native speaker, so long as you remove all of the letters that don't make sense." A native English speaker says, "You can't say the bot speaks English as well as a native speaker, since a native speaker would never write all those random letters." You would be correct in pointing out that sometimes native speakers make mistakes, but you would also be entirely missing the point. That's what's happening here.
Ah yes, of course, just because you never saw it means it never happens. That's definitely why rules exist around this specific thing happening. Because it never happens. Totally.
In fact, it's so rare that in order to forefeit a game, you have to do it twice. But it never happens, ever, because pattrn has never seen it. Case closed everyone.
I made no judgement on what ChatGPT can and can't do. I pointed out an extreme. Which the commenter agreed was an extreme. The rest of your comment is completely irrelevant but congrats on getting tilted over something that literally doesn't concern you. Next time, just save us both the time and effort and don't bother butting in with irrelevant opinions. Especially if you couldn't even bother to read what was already said.
You seem to have missed the part where I said multiple times that a 1400 has definitely made illegal moves.
> In fact, it's so rare that in order to forefeit a game, you have to do it twice. But it never happens, ever, because pattrn has never seen it. Case closed everyone.
I actually said the exact opposite. You're responding to an argument I didn't make.
> I made no judgement on what ChatGPT can and can't do. I pointed out an extreme. Which the commenter agreed was an extreme. The rest of your comment is completely irrelevant but congrats on getting tilted over something that literally doesn't concern you. Next time, just save us both the time and effort and don't bother butting in with irrelevant opinions. Especially if you couldn't even bother to read what was already said.
The commenter's throwaway account never agreed it was an extreme. I agreed it was an extreme, but also that disproving that one extreme does nothing to contradict his argument. Yet again you aren't responding to the argument.
This entire exchange is baffling. You seem to be missing the point for a third time, and now you're misrepresenting what I said. Welcome to the internet, I guess.
> fine, fair, "never" was too much.
This is the second time I've had to do this. Do you just pretend things weren't said or do you actually have trouble reading the comments that have been here for hours? You make these grand assertions which are disproven by... reading the things that are directly above your comment.
> This entire exchange is baffling.
Yeah your inability to read comments multiple times in a row is extremely baffling.
As I said before:
> Next time, just save us both the time and effort and don't bother butting in with irrelevant opinions. Especially if you couldn't even bother to read what was already said.
I did, two hours ago, 6 minutes after your comment
If I was playing that monstrosity though I would play something crazy that is far out of the opening book and count on it making an illegal move.
> You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed.
1. b4 d5 2. b5 a6 3. b6
> bxc6
No, it's ridiculous to say "oh, a blindfolded human might sometimes make a mistake." No, this is trivially easy to make it make a mistake. It has no internal chess model at all, it's just read enough chess games to be able to copy common patterns.
>> Two swindlers arrive at the capital city of an emperor who spends lavishly on clothing at the expense of state matters. Posing as weavers, they offer to supply him with magnificent clothes that are invisible to those who are stupid or incompetent. The emperor hires them, and they set up looms and go to work. A succession of officials, and then the emperor himself, visit them to check their progress. Each sees that the looms are empty but pretends otherwise to avoid being thought a fool.
So everyone "pretends otherwise to avoid being thought a fool".
Huh. I guess that explains it. Good metaphor.
I wish I could just make bullshit moves and get a higher chess ranking. Sounds nice.
I dont understand the point of your second sentence, seems to be entirely missing the substance of the conversation.
By the way - definitely read the article. But once again - I thought the methodology was bad, and thus the conclusion was bad.
But not going to keep replying, you engage online in a way that will turn lots of people you talk to away.
No, all unsuccessful playthroughs are resignations that impact the models ELO.
Whether any of it is actually true is a different question.
I got the impression the author did this because they thought they were being fair with ChatGPT, but they're much more likely to be letting it off the hook than they seem to realise.
(Sorry about the "they"'s; I think the author is a guy but wasn't sure).