Except that this isn't objectively an improvement, even in a perfect world where AI is substantially better than it currently is.
> you’d now write “why did the roman empire fall”. Because the AI likes that phrasing and produces better results.
"the AI likes that phrasing" is exactly the problem here. How is anyone supposed to know what the AI "likes", other than painstaking trial-and-error in the unbounded and arbitrarily high-dimensional search space of human language?
Even the people who built the model probably don't know. Language models (and deep NNs in general) are extraordinarily complicated things, and there are problems with pretty much every technique that purports to provide visibility into their inner workings. There are just too many parameters and too many "information paths" in such a thing for regular people to wrap their heads around it. The ability to incorporate a high amount of complexity is a big part of why those models are so effective to begin with, but it also makes them really hard to reason about.
"AI" is currently in a weird spot where it's starting to kinda-sorta behave like an intelligent human in some limited settings, but in general is nowhere near as smart as a human. Most models still have a very shallow conceptual understanding of anything, even if they're becoming uncanny in their ability to match sophisticated patterns. It might not even be possible to teach some concepts to language models as they currently exist today, if only because there is only limited conceptual understanding available to be learned from corpora of text and images, even huge ones. Humans are still tremendously more effective than our best language models at understanding meaning and intent. Can an AI ever learn about love, regret, fear, or bliss, by reading millions of news articles and books and looking at millions of images?
Thus AI right now is in a kind of "worst of both worlds" situation, where it is complicated enough to be hard to reason about precisely, but still mostly unsophisticated and therefore highly sensitive to how inputs are crafted. Therefore it's hard to formulate inputs that provide useful outputs. It's still alpha-level technology at best, and there might be one or several conceptual innovations remaining between what we have today and something resembling general intelligence.
Consider also that "AI assistance" is complementary to keyword search, not a replacement for it. Google search AI is becoming something like a "digital librarian", a creature that can understand your queries and guide you to a starting place in the relevant literature. But much like in a real library, the digital librarian is going to be most useful as a starting point. At some point, if you already know what you're looking for, you still are going to want to search on "structured" criteria, as well as, yes, keywords embedded in text.
And finally, do you really want to type a 300-word description in order to get good search results? I was already getting good results with 3 keywords. I have already done the sophisticated pattern-matching and concept-graphing in my own brain, and now I know exactly what terms I want to look for. Why should I be forced to coach an AI on how to redo all that work for itself, instead of just letting me do a damn keyword search? Not to mention wasting my time and giving me carpal tunnel typing it all out.