As for weak emergence[2] that is de facto.
As for strong emergence[3] well... do we have an example of the phenomena anywhere?
To be clear, ML people have a different definition which I do not think is entirely useful. Let's look at the wording used in [4]
Scaling up language models has been shown to predictably improve performance and sample
efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable
phenomenon that we refer to as emergent abilities of large language models. We consider an
ability to be emergent if it is not present in smaller models but is present in larger models.
Thus, emergent abilities cannot be predicted simply by extrapolating the performance of
smaller models. The existence of such emergence raises the question of whether additional
scaling could potentially further expand the range of capabilities of language models Emergence is when quantitative changes in a system result in qualitative changes in behavior.
An ability is emergent if it is not present in smaller models but is present in larger models.
I'm not sure anything here is actually meaningful. We have very little knowledge about the interpretation of ML models, so it seems to be jumping the gun to say that phenomena cannot be predicted. The aspects are highly coupled with things like training techniques, optimization objectives (or loss functions), optimization methods, not just architectural designs.We've also after more research found smaller capable of performing things that large models can do. So the goal post moves and makes the condition of emergent abilities. [RTFA]
So how do we define emergent abilities? Everything static except for the number of parameters? Does this make sense? Changing the number of parameters is an abstracted way of changing the optimization technique or training method. Clearly these are also different by nature of different batch sizes and/or a LR scheduler. So it is hard to be consistent.
Not to mention that we are not great at predicting much of what networks can do in general. There's plenty of people working on the subject (though this is a small proportion of total ML researchers, even if we exclude those that are highly application focused). So are we gonna call something emergent if it is just something we don't know about? And who gets to be the one to decide? I've seen abilities called emergent that are clearly a result of training methods like KL Divergence but surprising to individuals who do not have a rigorous statistics and/or metric theory background.
It just all seems to be jumping the gun. It's a new field and a new science. I think it is okay if we're willing to admit that our understanding just isn't great yet. There's no need to embellish or over attribute. These models are without a doubt powerful and useful, but at the same time I feel we are happy to greatly exaggerate all aspects about them. And for the life of me, I can't figure out why critiques are interpreted as dismissive. You can call an LLM a stochastic parrot and still think it is a great achievement and useful tool. Moreso, criticism is essential. Criticism gives us direction in research. Hype gives us motivation. But the two have to be in balance. Motivation without direction is an undirected Monte Carlo search (can still work, but we can do better than a drunken man). Overly criticizing (turning into dismissing, such as calling LLMs useless) is just pulling wool over one's eyes. It is the equivalent of lights being turned off and deciding to sit down and give up. When the lights get turned off you should find how to turn them back on. And we have more clues than a pure random process. After all, aren't we trying to make these better? I know I am. It is why I got interested in researching these things in the first place. I just wish we could have moderate hype rather than this ridiculously excessive one (which I think operates in a feedback loop with excessive criticism as it drives polarization).
[0] Sabine Hossenfelder (I know HN loves her): https://www.youtube.com/watch?v=bJE6-VTdbjw
[1] Sean Carrol: https://www.youtube.com/watch?v=0_PdLja-eGQ
Yes, Sabine and Sean are both controversial characters but they are well known.
[2] Weak emergence is where some larger phenomena forms out of smaller phenomena and results in something that the smaller thing can't do. The result requires the interactions of parts of a system. Temperature is a common example because it is generally easier to discuss the aggregated value of all the particles' "jigglyness" rather than each individual particle. The resultant property can be derived from the individual components. Conway's game of life is another common example. But in game theory we'd call a coalition an emergent property since utility is higher in the collective than the sum of each individual. Clearly this is true for ANNs since they are composed of neurons and a single neuron cannot perform these tasks.
[3] Strong convergence is about behavior that CANNOT be derived from the individual constituents.