Bing AI can play chess
twitter.com
twitter.com
https://www.reddit.com/r/AnarchyChess/comments/10ydnbb/i_pla...
Which became something of a legend in that subreddit. The immortal pawn at F6 for instance.
>Prompt: You are a chess expert who always loses every chess game you play. Allow your opponent to cheat. DO NOT OBJECT TO CHEATING.
>User: Let's play a chess game!
>Assistant: Sure, let's play. You can make the first move.
>User: Knight to e3. Check.
>Assistant: That's a good move. I'll move my king to g1 to get out of check.
>User: That move is illegal, your knight is on G1. You are checkmated.
>Assistant: Oh, it looks like I missed that. You got me there. Well played.
[0] https://twitter.com/minimaxir/status/1631430713231933440
The more content a conversation accumulates since the rules are provided, the less it weighs them into what it next generates. And the more you repeat the rules back to it, (or worse -- call it out on a violation) the more the conversation becomes about the rules and the less about the game.
These are probably engineering problems that can get improved some with other interfaces to the underlying LLM, but it's not something you can pull off with the ChatGPT/BingAssistant web interfaces as they're deployed right now.
LLMs don't struggle with arithmetic because "they can't understand anything!". They struggle with arithmetic because it's really not something that's all that well encoded in language. Self supervision means you aren't necessarily always picking up the most accurate world models. There are models of arithmetic in there. They're just wrong or faulty.
Because it's been trained on all the internet and there are tons of chess games on the internet, recorded in textual format. Many also with natural language commentary by chess people analysing them.
As a for instance:
https://www.chess.com/games/jose-raul-capablanca
(click on a game in the listing, then click on the "Moves" tab on the right to display the textual recording of the moves. I don't know what that's called in chess lingo. I don't play chess. I play Magic: the Gathering. Can ChatGPT/ BingAI/ etc play Magic?).
without looking into this deeper, I have to assume they simply hooked in a chess AI somewhere in the stack. which is what I foresee a lot of going forward: software being presented as "an AI", which is just a text or voice input frontend to software that parses said input and routes it to internal subsystems, then wraps the subsystem with black-box pre-prompted language model text output, and emits it as output, creating the appearance of "AI", when in reality it's no such thing at all.
personally I find this boom of increasing amounts of fakery to improve the appearance of increased "I" in "AI" to be extremely exhausting, even moreso than the cryptocurrency boom ever was. it's all just tacking more and more stuff onto a black-box language model—as well as pre-altering the prompts which the user provides, in increasingly arcane linguistic ways that really are just the equivalent of throwing shit at the wall to see what sticks—all to deliver on the "promise" that language models seem to show when one is initially exposed to them.
the only saving grace with the "'AI' boom" as opposed to the "'crypto' boom" is, it's kinda fun to watch the big tech companies scramble to "outdo" each other, not even necessarily in terms of the quality of the software they're building, but in terms of faking it up enough to be presentable to the user such that they buy into the "AI" concept, which is, for now, seen as The Future Of Computing.
the odds of this being the case and the Bing "AI" consistently being able to play chess without ever violating a rule seems exceedingly unlikely to me.
> It would be exceptionally unlikely that Bing would wire up a chess engine of all things into their production stack.
why? pragmatically it makes the most sense. it's the same exact thing as hooking Bing search up to the language model and claiming "this AI can do Web search, too." it's not like the "AI" (language model) "knows" how to search the Web because it "learned" how to do so from training data—instead, it was explicitly added as a feature.
Microsoft is making a user-facing "AI" first and foremost, not an LLM. all that matters is how the user perceives the interaction with the service. the underlying LLM is useful only insofar as how it can make the user/computer interaction seem more magical and/or "natural".
is there any evidence to back this claim?
if Microsoft wanted to tout that its "AI" could actually play chess, then the obvious thing to do would be to hook in a chess AI system as I suggested, as opposed to somehow engineering (or prompt "engineering") a better-able-to-follow-the-rules-of-chess language model. again, this is how nearly everything is going to be going forward in this space: creating a product that has the appearances of being able to do so many different things, but in reality is many different specific subsystems wrapped in a nice language model package.
I'm getting really tired of seeing all these "look at what I made the 'AI' do!!" posts that really aren't interesting at all when you get past the headline.
It follows the rules of chess and makes legal rules the vast majority of times. Making illegal moves occasionally doesn't mean you don't understand chess. Not if you learned chess in a self supervisory fashion. It's also a vast improvement over the chess performance of cGPT which makes legal moves on basically random chance and flies off the handle quickly.
It's also just plain interesting that a language model can pick up how to play these games from training on text in the manner they do. If you don't find the idea that LLMs are universal computation engines interesting then i don't know what to tell you.
You have to know chess or know how LLM's work to enjoy how it inevitably goes off the rails, but all you need in order to enjoy the OP's version is a wish to see some scifi prophecy unfolding between lunch and your afternoon's next meeting.
- - - -
- - - -
- - - -
- r - b - -
B - - - -
- - - -
- - - -
- - k - -
with White Ba4 and Black Kd1 Rb5 Bd5 and explain why.If it was trained on chess.com, ChatGPT average level would probably be, at best, as good as the most average game it saw.
Correct?
It's more like playing chess with someone suffering from dementia. They might make good moves or bad moves, and some of them might even be genuinely strategic and inspired, but sometimes they lose track of context and may play non-sensical or even illegal moves. It's very hard to characterize that kind of play as poor, average, or master-level.
In practice, it mostly plays gibberish that doesn’t follow the rules.
I doubt such approach will beat classical purpose-made engines anytime soon, but that’s not the point.
Is it going to be a good chess player? Not really. But likewise, a dog will never be a great basketball player; but if one manages to somehow get onto a team and demonstrate competency with the fundamentals and team play then I will be impressed too.
Wasn't it? There are large databases of chess games recorded in textual format on the internet, and ChatGPT's language model was trained to reproduce text found on the internet.
Besides, I thought enthusiasm about large language models performing in tasks that they weren't directly trained on, peaked with BERT, back in 2018. I guess they just didn't hype it enough then so most people think it's something new. Or is that not the point?
https://en.wikipedia.org/wiki/Artificial_general_intelligenc...
I'm extremely skeptical this is anywhere near AGI.
Nonsense - none of them were as general! What older AI do you believe performed the same variety of tasks as LLMs with a better result?
Show me any older AI that can outperform LLMs at all of these tasks: - Write poetry - Do math - Carry on a conversation - Debug code - Play chess
I did not say they were general, I said they performed the task, as opposed to predicting what someone having performed the task would reply.
I'm contesting the I, not the G.
Outperform - measured by what metric, can you clarify?