Stupid question: What is the "tricky" part about it? How can someone who isn't completely retarded get this even wrong?
This puzzle is trivial, and especially doesn't require any kind of "thinking around the corner". So what is this about?
That language models can't do math (or logic) and can't even reliably tell which of two given written out numbers is the bigger one is imho a different story.
Language models are great for text. But when you need a tool for math and/or logic you should use an adequate tool. We have for example algebra systems for that. Or prove assistants. Or just good old Prolog. It makes no sense to use the wrong tool and than wonder that the results are terribly wrong.
But for the above "puzzle" you don't even need a calculator. So I really don't get the point.