For example, try to ask (better in Russian), how many letters "а" are there in Russian word "банан". It seems all models answer with "3". Playing with it reveals that apparently LLMs confuse Russian "банан" with English "banana" (same meaning). Trying to get LLMs to produce a correct answer results is some hilarity.
I wonder if each "failure" of this kind deserves an academic article, though. Well, perhaps it does, when different models exhibit the same behaviour...