So if I ask Bing about me it says "Rory McCune is a Cloud Native Security Advocate at Aqua Security." without any ref.
The problem is, that's not correct, that's a job I had two years ago, but someone reading that could be forgiven for thinking that's a fact, given how it was presented.
In this case that's harmless, but I could easily see cases where it would not be harmless.
(I don’t know if that’s how Bing AI works)
So the gaps are the only areas where the LLM can hallucinate on and if your search query is easily available information on the internet, then hallucinations will be less or none.
Edit: I have used RAG with a project that I am working on and it's quite hard to ascertain if the LLM used the information provided as part of the RAG documents or just made up information on it's own, since even without RAG, we were getting similar responses 7 times out of 10.
After all, if you've trained an LLM on a masses of unchecked data you've scraped from the internet, your training data probably includes "Joe Biden is the president" and "Donald Trump is the president" and "Barrack Obama is the president" and "Emmanuel Macron est le président" and so on. It would be understandable if an LLM was confused about who the president was.
These people think handing an LLM the contents of https://en.wikipedia.org/wiki/President_of_the_United_States then asking who the president is sounds a lot more feasible.
Personally I'm not so sure - I've never seen a RAG implementation that impressed me.