Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o's capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn't be replicated on any of the newer "long" term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.
Combing back to coding agents, where they're releasing new models and harness tweaks multiple times per month, and vibes are the only - let's not say sensible, maybe realistic - thing a developer reasonably has to go on.
Hah, welcome to what it can feel like to manage people.
you send an email to get something done and you get an unreliable answer from the peons... is it truly that different from an unreliable AI?
In practice, I think that kind of effort is mostly a hindrance at the moment, because of how fast things are moving, and because so much about this is subjective and everyone has different preferences.
On these things, yes. But in the case of LLMs, this is impossible. It's not possible to understand what the models are doing, and for several different reasons. You can at best evaluate them, like you do with your fellow engineers when hiring. But that's not a guarantee of anything.
And that's why I don't think this is an engineering renaissance. Engineering progress goes with better understanding of our tools, and adoption of more rigorous practices. LLMs go in the opposite direction.
At the time of the Renaissance, its trajectory wasn't some well-signed garden path, it was chaos.
If there is a field that should recognize that major improvements come with creative destruction, and during the period of churn nobody can credibly guarantee what part of the destruction will be successful and therefore be knighted the "creative" part, it would be computer/software engineering.
If you enjoy being an expert in which LLM is best, then I'm all for it. However that isn't an interesting problem for me or my boss. I'm happy someone else figured it out so I can work on interesting problems.
The above likely scares all the model providers: they are a commodity with easially substitute competition. There is a minimum quality standard, but once you meet that there is nothing to differentiate you and so price matters and it becomes a race to the bottom.
"But models will get better". Maybe. But there has been a decline in the speed of development
That's not really possible with frontier models changing every 3-6 months, just like javascript frameworks there's not enough time to learn the ins and outs and they mostly expose the same interface, so unless you have rigorous evals you are working on vibes, and the amount of meetings I have had in the last 3 years of engineers confidently reporting on their vibes is SO TIRING.