ChatGPT is not up to date unless you start using the plugins. This sort of indexing is based on vector databases and various intermediate prompting. If you want to get technical, the academic term is "Retrieval Augmented Generation".
I unfortunately don't think we'll be able to solve hallucination anytime soon. Maybe with the successor to the transformer architecture?
We’ve been testing LLM responses with a CLI, we’re using it to generate accuracy statistics, which is especially useful when the use-case Q/A is limited.
If ‘confidence’ can be returned to the user, then at least they can have an indication if there is a higher quality-risk with a given response.