Yup, LLMs are not "artificial intelligence" - they just generate most probable token, until their authors hardcode functionality for specific community tests.
Fundamentally though there is missing but implied information here that the LLM can’t seem to surface, no matter how many times it’s asked to check itself. I wonder what other questions like this could be asked with similar results.