But it is not the MBA's view of winning, it's just one potential conclusion you could draw from the bitter lesson of Machine Learning. As long as the need for more intelligence outpaces the economics of using intelligence, you'll get bigger models. This idea that small, fine-tuned models can outperform bigger models capabilities wise is mostly misinformed. They are genuinely good at other metrics, but sadly more actually means better in ML-land (most of the time at least).
There’s evidence that model size and representational capacity are not exactly the same, and that scale is maybe more important for learning than it is for representation (past a point). Consider the early work from the current neural scaling paradigm. The Chinchilla scaling study shows that smaller models can match the performance of larger models by training longer.
To GP’s point, if everyone is exploiting the scaling lever, few resources are being allocated to finding more efficient training algorithms that could let us work with right-sized models instead of pulling the scaling lever as hard as we can afford to.
I’ll end with a dramatic example from my field of materials science (which admittedly might not strictly generalize to LLMs). A lot of the field is pursuing the model scaling strategy, and it’s still paying off. But [0] recently reported competitive accuracy with much smaller models that run faster and can address much larger problems. The model architecture is pretty much the same, but they use a different training strategy and really focus on data quality
Or put differently a wide exploration of long chains of obvious insights might be more valuable than a more narrow exploration of shorter chains of deeper insights. But perhaps that is a wrong sense of the difference between a small model and a large model.
Based on a quick cursory glance at your example: Better data + better technique led to a better result with less parameters. Would you assume that then scaling both the dataset and model once again would lead to even better results? If you haven't fully encompassed the underlying distribution with datapoints, intuition says yes.
What I wanted to initially highlight was actually something slightly different: Specifically that people keep trying to "outsmart" optimizers by either fully hand-crafting solutions or skewing existing machine learning algorithms via additional tricks that are supposed to encode "human intuition" or something similar (to be fair there are ways to do it correctly). These all tend to fall short in a few years simply due to "line go up" being stupidly effective (compute getting cheaper, more training data being available, better optimization strategies, better architectures) [0]
Specifically this idea of small fine tuned LoRA models falls into the trap quite often: People assume you can beat the big, slow, general purpose LLMs with a small highly specialized model that has been fine tuned on the "good" human intuition of your special inhouse dataset.
LoRA can do great things, but it is often misunderstood what LoRA actually does.
0: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
What I’m pushing back on is what I think is a sort of one-dimensional view of Sutton’s bitter lesson. People seem to equate it with model scaling, but there are lots of general ways to leverage computation that don’t involve just scaling models and supervised training datasets up. For example Sutton’s first example is straight up search, no parameters at all.
The point of the force field example is that it seems you don’t need billions of parameters to represent the functions we’re interested in, but with small models it’s harder to find those functions by pushing harder on the standard training algorithms, and that maybe some different algorithm that leverages computation more effectively could do so.
But that's exactly my claim: Smaller models can win on the computational efficiency front, but not the overall capability front. And as long as compute is getting cheaper, investing in more efficiency while there are still major capability on the table isn't a good business strategy.
Smaller, efficient models could lead to some really interesting things though especially considering it could lead to some Jevon's paradox like moment. To be honest, I feel like the biggest issue in regards to LLM usage in practice is that the patterns of use aren't really well developed. Yes we have agents, but it's still somewhat unclear what an agent can "do" - People seem to mostly focus on replacing some kind of existing process with an agent driven one, but actually coming up with AI-native processes is way harder.
A deeply ironic comment which associates <THING YOU DON'T LIKE> with <GROUP YOU DON'T LIKE> due to complete ignorance about the group. An MBA would never approve a technique with basically unlimited capex. So I hate to break it to you but "bigger weights" is 100% the computer scientist's view of winning because everything is an "abstraction".