> I don’t think this rebuts the significance of emergence, since metrics like exact match are what we ultimately want to optimize for many tasks. Consider asking ChatGPT what 15 + 23 is—you want the answer to be 38, and nothing else. Maybe 37 is closer to 38 than -2.591, but assigning some partial credit to that answer seems unhelpful for testing ability to do that task, and how to assign it would be arbitrary.
Not sure I can agree with this. Let's for the sake of the argument say that normal calculators don't exist. I can choose to run my calculation thought ChatGPT or do it myself. It is true that I would vastly prefer the answer to always be correct, but it's not true that there is never value in bounding the error.
Put another way: would you rather want a calculator that is at most 5% off 100% of the time. Or a calculator that is 100% correct 95% of the time, but may output garbage the remaining 5%?
Upon further reflection: it seems mighty ambitious to expect a statistical model to be always correct. If always correct was possible you probably didn't need a statistical model to begin with. Given that the model will fall from time to time it seems especially useful that you can bound the error somehow. +1 for smooth metrics i guess.
I guess all of this is slightly off a tangent as Wei seems to be just arguing that "there might be things LLM's can do or will learn to do in the future that we didn't train them for. And we won't necessarily be able to predict it from capabilities of smaller models". I agree with this, just not the way he arrives at the conclusion.