OP highlights application problems, and RAG specifically. But that is not an LLM problem.
Chat is such a “leaky” abstraction for LLMs
I think most people share the same negative experience as they only interact with LLMs through the chat UI by OpenAI and Anthropic. The real magic moment for me was still the autocompletion moment from the gh copilot.