I don't think this paper is dismissing the importance of correct yes/no tests and reaching an accuracy threshold making it generally useful to humans, but that you should use more than correct yes/no tests before declaring some behavior is emergent.
Would you rather be able to evaluate a model on it's demonstrated capabilities (multistep reasoning, question and answer, instruction following, theory of mind, etc) or some nebulous metric along an axis that may as well not correspond to practice.
We only care about how good AI is at things that matter to us as humans. Why not test for these directly?
If some perfect metric is discovered that shows the phenomenon of emergence is actually continuous, then that would be helpful.