I would say this did a really good job of playing chess. It moved the pieces consistently and traded pieces when required.
This is worlds away from the frontier ~1 year ago where models would hallucinate pieces into existence.
However, the pawn was defended by the queen and it took a forced queen trade to unlock the move.
I have seen much worse blunders from human players. And, I have made much worse blunders.
The fairest apples-to-apples comparison of an LLM whose training data included chess games would be a trained human such as Magnus Carlson, who can quite happily play a dozen or more simultaneous blindfold chess games.
Is this because the context is being saturated? How did you set it up?
Was the prompt something like "Here's the state of the board, you're white, your move, what do you do?" and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (thanks for sharing!) to my mental model and understanding.
I'd be curious how it would work if it started fresh each time. My guess is it would never make an illegal move, although it may not actually play all that well.