Is the binary distinction helpful? It doesn't appear to be a more helpful way of evaluating how capable a model is.
Would you rather be able to evaluate a model on it's demonstrated capabilities (multistep reasoning, question and answer, instruction following, theory of mind, etc) or some nebulous metric along an axis that may as well not correspond to practice.
We only care about how good AI is at things that matter to us as humans. Why not test for these directly?
If some perfect metric is discovered that shows the phenomenon of emergence is actually continuous, then that would be helpful.