The fact that there are reasonable metrics that make the emergence completely predictable as a simple threshold effect is massive.
One alternative way to phrase this that I haven't seen in either thread yet (but I am on mobile and haven't looked thoroughly) would be: Human observers/users dramatically underestimate how much small language models already have learned, because we are very good at spotting the errors. Scaling up looks like emergence occurs because the last few errors get eliminated.