If you can make the existing model faster, you can then save your inference budget to then make your model bigger, which then makes it smarter.
A lot of how smart the models can be comes down to budget. If you can make your existing thing cheaper, you can instead make it bigger for the same price.
(Not trying to flame bait or anything. I just wouldn’t call LLM as exhibiting intelligence. It is great at making connections based on probability but doesn’t have a semantic understanding of what it is doing)
> doesn’t have a semantic understanding of what it is doing
I hope you realize this is an area of open, active research.
In general we (humans) need to be humble about the limitations of our knowledge about how we function, it's an insanely complicated problem.
We do.
Which is why we shouldn't be assuming we're more than just probability engines, or be assuming we have more consciousness than a neural network.
There's diminishing returns and at some point making a model bigger makes it dumber.