Then why do Chess AI perform much better than LLMs trying to play chess.
Small ai: likely makes up a function or suggests a single function which isn't sufficient. Refuses to budge from its answer or apologies and gets confused
Large LLM: able to actually understand the question, combine several functions. If it doesn't work you can tell it why and it fixes it
An LLM could have picked up some chess patterns through osmosis, but it can not reason explicitly in the domain.