This is similar to other brain functions that aren't present at birth and require stimulation, such as sight. That is, if your eyes are forced closed for the first few months of your life, you will never be able to see, even if later they are uncovered, and keep working perfectly. The brain functions responsible for interpreting visual signals can only develop if they get visual signals in a (quite short) developmental window - and we know this with quite a bit of certainty from quite cruel animal studies.
Language acquisition is not proven to be the same, as the required studies would be deeply unethical, but the few experiences with feral children are highly suggestive that the same applies.
A good enough formula for a task isn't a solution for every task. Yes Newtonian mechanics work, but Einstein is a better reflection of reality.
Newton: Do you need more than that to describe the speed of a thrown baseball on a train? No. DO you you need more than newton to get to the moon? No. Is it going to be accurate at high speed in a large scale system (anything traveling near C)? NO, it fails spectacularly.
NN's are great at simulation, language, weather... But what people using them for weather seem to understand and the ML folks (screaming about AI and AGI) dont is that simulation is not a path to emulation. Lorenz showed that there were limits in weather, that most other disciplines have embraced these limits.
The "emergent property" aspect is when LLMs are good at a task at scale X*3 but were incompetent at scale X.
You could sort of represent the deterministic contents of an LLM by compiling all the algorithms and training data in some form, or maybe a visual mosaic of the weights and tokens, or what have you...but that still doesn't really explain the outcome when a model is presented with novel strings. The patterns are emergent properties that converge on familiar language--they're something deeper than the individual words that result.