Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.
Based on my experience with running models locally, there is a threshold of intelligence required to be useful. But it’s possible there is also a ceiling where smarter isn’t necessarily better. If you ask a 4B parameter model to fix a bug, it might e.g. fix the bug but fail to fix a compilation error created by the fix. If you ask a frontier model, it might fix the bug, re-write your unit tests, and update the readme. Maybe you wanted those things but maybe you didn’t. “Smarter” is often shorthand for more proactive, and guessing more about your intent. Which is great when it gets it right, and annoying when it gets it wrong.
I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.