Why R the Critical Value and Emergent Behavior of Large Language Models Fake?
cacm.acm.org
cacm.acm.org
Hope this helps anyone else who thinks this is about the R Programming Language or the correlation coefficient symbolized by r
The whole point is that below a threshold, you can increase training by a factor of 10 or 100 and there's barely any change. But at the threshold, the same relative increase suddenly produces big improvements.
How do you know accuracy didn't increase by a factor of 100 from 0.001% to 0.1%? Another factor of 100 for both and you're at 10% with "only" 10⁴ as many training FLOPS!
If you want to use a nonlinear transformation to show very small and very large inputs in the same graph, surely a similar effort should be made to make changes in accuracy near the limits of 0% and 100% more visible.
Logarithmic axes are the norm for showing scaling, power law effects etc. Once you get used to them there's no way back.