Yes, that's how you can really tell if the model is doing real thinking and not recombinating things. If it can correctly play a novel game, then it's doing more than that.
"Chess but white and black swap their knights" for example?
None of these changes are explained to the LLM, so if it can tell it's still chess, it must deduce this on its own.
Would any LLM be able to play at a decent level?
These LLM's just exhibited agency.
Swallow your pride.
>'thinking' vs 'just recombinating things
If there is a difference, and LLM's can do one but not the other... >By that standard (and it is a good standard), none of these "AI" things are doing any thinking
>"Does it generalize past the training data" has been a pre-registered goalpost since before the attention transformer architecture came on the scene.
Then what the fuck are they doing.Learning is thinking, reasoning, what have you.
Move goalposts, re-define words, it won't matter.
We are pawns, hoping to be maybe a Rook to the King by endgame.
Some think we can promote our pawns to Queens to match.
Luckily, the Jester muses!