In a sample of >1000 games, GPT-3.5-turbo-instruct plays chess with ~1800 elo
github.com
github.com
I checked a random sample of games manually, and all of them quickly diverged from each other and from the lichess database (of ~5B games), so it is inconceivable that the training data contained a significant portion of these games.
Hopefully OpenAI releases an instruct version of GPT-4 so we can see if scaling continues to bear fruit.
EDIT: I played against it once. It got a winning position and then blundered mate in 1. Considering the chat model can barely find legal moves this is still very impressive IMO.