Ah good. Well, then if you're interested in a half-empirical sincere attempt to characterise "why certain weights obtain certain values under optimisation" then i'm much more inclined to be, say, more humble on these matters.
The reason CNN weights obtain 'recursive-hierarchical representations' of pixel-pattern geometry in the training data follow from the recusrive-heirachical relationship of their weight matrices and from the geometry of the 'pixel space' from which the training data is drawn.
This is certainly interesting; and there's something magical feeling about 'principles of least action' at work. Indeed, many new physicists have a kind of schizophrenic reaction to discovering action principles -- since it imparts to nature a strange apparent conspiracy.
Of course, the job of any good physicist is to be sceptical of this conspiracy, and to get to the heart of how 'accounting tricks' performed by moving objects over time create this illusion.
Likewise this is the job of any good ML researcher; yet they do the oppoiste. Rather than get to the heart of this apparent conspiracy, they call it 'emergence' -- this offends my sense of what the virtues of a scientist ought be.
In any case, on the matter of the LLMs obtaining 'useful' weights for any given task here the job of the researcher is circumstantial, empirical, and sceptical: go and find those 'accounting tricks' within the training data that give rise to this apparent conspiracy of the system to acquire a useful state.
There is no emergence: there is just a set of weights which compress the structure of a target space. At some point this set is large enough, and 'lies across the space like a mental chain does a gate'.
Emergence is an ontological relation between parts and wholes whereby wholes arent reducible to their parts because of ontologically-relevant interaction properties between their parts which aren't intrinsic properties of them.
The fluidity of water emerges out of hydrogen bonding which does not occur when you isolate H20 alone. There is no such relationship here.
This ontologising of the formal, this language which gives a causal-physical semantics to purely formal properties of abstract models -- this is pseudoscience. It's done as part of a computational-idealist worldview in vogue because it's a helpful language for VC investment were-changing-the-world hype.
The formal properties of NNs cannot be described in these terms, because they do not have ontological relationship -- they have formal (mathematical, statistica, etc.) ones.