The paper is goal-post shifting by measuring another thing. At the level of yes/no the behavior is emergent.
Would you rather be able to evaluate a model on it's demonstrated capabilities (multistep reasoning, question and answer, instruction following, theory of mind, etc) or some nebulous metric along an axis that may as well not correspond to practice.
We only care about how good AI is at things that matter to us as humans. Why not test for these directly?
If some perfect metric is discovered that shows the phenomenon of emergence is actually continuous, then that would be helpful.