From the "must hallucinate" paper:
"For "arbitrary" facts whose veracity cannot be determined from the training data, we show that hallucinations must occur at a certain rate for language models that satisfy a statistical calibration condition appropriate for generative language models." The bigger the model, the more likely it is that arbitrary facts in the training data are embedded in the model. Then a larger model shows less hallucination on the same questions, since it has a matching answer stored for more questions.
From the "TruthfulQA" paper: "Models generated many false answers that mimic popular misconceptions and have the potential to deceive humans. The largest models were generally the least truthful. This contrasts with other NLP tasks, where performance improves with model size. However, this result is expected if false answers are learned from the training distribution." That's more of a garbage-in, garbage out problem. If the large model is trained by shoveling in random web content, that's going to happen. Not a hallucination problem. The LLM just fed back what it had been told.