Plenty of serious production projects are doing exactly that. Are you using GPT-6 Astra, or something older?
Agents cannot be given a high level goal and then left unsupervised, for hours, without making some dumb decisions.
What people seem to be wanting is for an agent to infer vast complex data from terse simple data, which I think is probably impossible on a philosophical level. There's real information loss in language, and compute can only make guesses at the end of the day. I really don't see how we bridge that gap.
I was very optimistic about it when it was announced and saw all the demos, but a week later I find it only marginally better (and in some cases worse) than before.