Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.
Based on my experience with running models locally, there is a threshold of intelligence required to be useful. But it’s possible there is also a ceiling where smarter isn’t necessarily better. If you ask a 4B parameter model to fix a bug, it might e.g. fix the bug but fail to fix a compilation error created by the fix. If you ask a frontier model, it might fix the bug, re-write your unit tests, and update the readme. Maybe you wanted those things but maybe you didn’t. “Smarter” is often shorthand for more proactive, and guessing more about your intent. Which is great when it gets it right, and annoying when it gets it wrong.
I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.
Do you want fable for one-shotting a game or website? Probably! The whole thing is mostly existing examples with small modifications that it will definitely get right. Do you want fable to just go nuts on a large, custom, unusual code base built around domain-specific problem solutions? Absolutely not, it will fix every problem it's presented with while creating lots of new ones.
Past 10k lines on something custom and with real-world complexity, you have to start thinking about which model should design, which should implement, which should review, and the appropriate effort-settings for each. Even then.. the answers aren't static because it depends on the task. And all this is assuming the starting place actually inherited reasonable due diligence on architecture/design. The idea of releasing the most generally intelligent models on 10k lines that were themselves the product of agents is yet another matter.
Part of what's at work here is that, like humans, every model can very easily create working code that it is completely incapable of maintaining. So realistically using multiple strengths tactically to avoid "excess creativity" needs to be SOP already, even if granular experts and specialists aren't in the usual workflow yet.
If we are talking about human labor, how many people hack their way through their work day?
Define hack...
Not doing their work, copying other peoples work, putting off work till later, taking credit for other peoples work, literal law violations.
Actually humans do this quite a lot and there are just massive numbers of business and regulatory processes and checks to ensure they are not doing it. With humans every human that is good enough to hire and do you work you want also have the ability to steal everything in sight and run away if they so choose.