> nothing in this paper should be interpreted as claiming that large language models cannot display emergent abilities; rather, our message is that some previously claimed emergent abilities appear to be mirages induced by researcher analyses
imo their paper's position is
- downstream task metric jumps in accuracy don't imply sudden changes in pretrain loss
- you can cherry pick metrics to show the sudden appearance of emergent abilities
- model scale is not enough to explain emergent abilities because you can scale up one model family and not get them
Their follow-up paper [0] has an appendix that touches upon how:
- it's not impossible for continuous metrics to undergo sudden increases of performance
- therefore not all sudden increases of performance are a consequence of using discontinuous metrics
- therefore emergent abilities could still actually exist
[0] Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive? https://arxiv.org/abs/2406.04391
Also the size of the model itself isn't the only factor that determines performance, LLama 3 70B outperforms LLama 2 70B even though they have the same size.