You picked a piss poor metric and you are getting piss poor results.
You can absolutely train an LLM-based model to be superhuman at chess. We just don't care enough to do so.
LLMs being even as good at implicitly tracking chess board states as they currently are is an emergent capability. The difficulty wasn't in "understanding rules of chess" in a long while now.