Do they need to learn rules though, or could they memorize enough to learn the probabilities of the most likely continuation? To my mind the blurry-jpeg-metaphor[1] is very much spot-on. While it is speculation, it would seem to me personally that LLMs de-facto doing a (lossy) memorization and neural networks in general being able to do well in "local consistency" seems a way to think about them that is consistent with most of the model behaviour we observe.
1. https://www.newyorker.com/tech/annals-of-technology/chatgpt-...