Hallucination is unfortunately inevitable when it comes to any autoregressive model, even with RAG. You can minimize hallucination by prompting, but you'll still see some factually incorrect responses here and there (https://zilliz.com/blog/ChatGPT-VectorDB-Prompt-as-code).
I unfortunately don't think we'll be able to solve hallucination anytime soon. Maybe with the successor to the transformer architecture?