I agree in that a perfectly consistent dataset won't completely stop statistical language models from hallucinating but it will reduce it. I think it is established that data quality is more important than quantity. Bullshit in -> bullshit out, so a focus on data quality is good and needed IMO.
I am also saying LMs output should cite sources and give confidence scores (which reflects how much the output is in or out of the training distrtibution).