3 * N < 42 * N
42 - 3 < N * (42 - 3)
It helps to know layer you're working on.
People seem to make mistake of thinking how good LLMs are around tasks that they are familiar with and extrapolating it to whole population.
It's good mental exercise to think about how little you can do compared to expert on tasks you never thought of working.
Ie. if you're programmer or know something about finance, don't think how much it enables you to do better coding or investing, think instead how much it doesn't enable you to work on something you don't know like maybe molecular biology or visual special effects – it's all there but it's much better multiplier for people who do know their shit.
Knowing layer you're working on helps a lot, it gets multiplied.
Knowing programming is becoming more fundamental skill than ever before as it lies at the foundation of almost everything else.
I’ve heard people say older models can’t do X, when I used that way etc. I suspect people are applying their own learning curve as part of their assessment of progress, you get better at writing prompts and it feels like the model improved.
Which is why I’m saying we need some objective metrics to judge predictions of actual capacity.
These companies aren’t just making stuff up, they really do want to improve the models, and the models really are improving.
I’m aware of multiple cases of benchmark cheating/“optimization”. So, taking benchmarks a face value seems laughable.