Geoffrey Hinton recently discussed how neural networks, even human ones, may be "generative" in how they recall information. Our memories of events are hazy, and change every time we recall them. Memorized rote facts can also be hazy in this way, and are subject to mix-ups and confabulations. That we "generate" memories from neural weights in our brain can also help explain why it seems like our brains store so much information. Perhaps they don't, they instead store lossy neural weights built from our past sensory experience, and we generate the rest and use metacognitive reasoning and a form of attentive reasoning upon recent data/context to sort through the potential errors in memory.
As you point out, LLMs work much better when you ask them to operate on objects within its context window, especially artifacts it knows how to work with, like code and text. But I think people are so trained to ask questions to the oracle and expect answers (e.g. Google), and who can blame them, that is the UX built into people's muscle memory for open-ended text input boxes. The launch of ChatGPT Search is recognition of this. Plus, most people are being told to treat these chat boxes as strong AI rather than as text/code-processing programs with specific strengths and weaknesses.