LLMs are compression and prediction. The most efficient way to (lossfully) compress most things is by actually understanding them. Not saying LLMs are doing a good job of that, but that is the fundamental mechanism here.
If you're interested in why compression is like understanding in many ways, I'd suggest reading through the wikipedia article on Kolmogorov complexity.
This is a case where it's going to be next to impossible to provide proof that no counterexamples exist. Conversely, if what I've written there is wrong then a single counterexample will likely suffice to blow the entire thing out of the water.