whereas my experience describing my problem and actually asking the AI is much, much smoother.
I'm not convinced the "LLM+scaffolding" paradigm will work all that well. sanity degrades with context length, and even the models with huge context windows don't seem to use it all that effectively. RAG searches often give lackluster results. the models fundamentally seem to do poorly with using commands to accomplish tasks.
I think fundamental model advances are needed to make most things more than superficially automatable: better planning/goal-directed behavior, a more organic connection to RAG context, automatic gym synthesis, and RL-based fine tuning (that holds up to distribution shift.)
I think that will come, but I think if LLMs plateau here they won't have much more impact than Google Search did in the '90s.
I’d give building with sonnet 4 a fair shot. It’s really good, not accurate all the time but pretty good.
Given that Google IPOd in 99, and is one of the biggest tech companies in the world, I'm not sure what you mean by that.
e.g. if OpenAI is responsible for any damages caused by ChatGPT then the service shuts down until you waive liability and then it's back up. Similarly if companies are responsible for the chat bots they deploy then they can buy insurance or put up guard rails around the chat bot, or not use it.
In a reality with perfect knowledge, complete laws always applied, and populated by un-bankrupt-able immortals with infinite lines of credit, yes. :P
I've said the same thing as you, that there is a LOT left to be done with current AI capabilities, and we've barely scratched the surface.