Also: it works much better in English than in other languages I know. English grammar is simple, and many of the function words carry a strong prediction for the next tokens.
Which languages are harder for LLMs? Is there any writeup or analysis of this?
If I had to guess, from my time working with ASR (Advanced Speech Recognition) languages and its difficulty with the Indian (country of) accent; I'd guess Indian languages.
what's advanced speech recognition?
Any with inflections and particularly with grammatical gender.
Idk. First, the obvious candidates: those with small, low quality text corpora. Second, I expect LLMs to perform less on long distance dependencies. English has a pretty strict word order, and the constituents are close to each other. There's also a lot of input available, so there are more large gaps to learn from. But (some) Germanic languages have some weird rules which allows constituents to move far away, so it's likely LLMs will make more errors there. Also, free word order languages might be a problem, although I've always suspected that those languages aren't that free in reality: speakers probably have strong preferences.