Reading anything by major researchers in AI feels like an adversarial battle where they're trying to misuse as much technical scientific and philosophical language as possible and we adjacent are trying to hold the line.
In philosophy and esp. the philosophy of science, emergence is a relation between a whole and its parts such that a property of the whole does not obtain just in virtue of properties of its parts taken in isolation. "Emergence" has this prior positive, semi-magical, scientific association which confuses the issue in this case.
No properties of the LLM obtain from its parts differently as parameters scale, the mechanism is the same. The performance differs not due to emergence, but due to the "modelling gap" present between the statistical structure of free text and that of mathematics. With enough examples, the gap closes... indeed, you can model the addition function (add(x, y) = x + y) just by an infinite sample of its domains.
A better technical term here might be "scale-dependent capabilities". For LLM, simple arithmetic is extremely scale dependent, whereas basic text generation is less-so. The reason for this seems obvious, as given above... so the use of the term "emergence" here I interpert as more PRish mystification.