Big tech created a problem for themselves by allowing people to believe the things their products generate using LLMs are facts.
We are only reaching the obvious conclusion of where this leads.
On the other hand, training a small model to hallucinate less would be a significant development. Perhaps with post-training fine-tuning, after getting a sense of what depth of factual knowledge the model has actually absorbed, adding a chunk of training samples with a question that goes beyond the model's fact knowledge limitations, and the model responding "Sorry, I'm a small language model and that question is out of my depth." I know we all hate refusals but surely there's room to improve them.
I always thought that was a correct and useful observation.
But all of their output it literally "made up". If they didn't make things up, they wouldn't have a chat interface. Making things up is quite literally the core of this technology. If you want a query engine that doesn't make things up, use some sort of SQL.
Obviously it's asking for a lot to try to cram more "self awareness" into small models, but I doubt the current state of the art is a hard ceiling.
This has already been tried, llama pioneered it (as far as I can infer from public knowledge, maybe openai did it years ago I don't know).
They looped through a bunch of wikipedia pages, made questions out of the info given there, posed them to the LLM and then whenever the answer did not match what was in wikipedia, they went ahead and finetuned on "that question: Sorry I don't know ...".
Then, we went one step ahead, and finetuned it to use search in these cases instead of saying I don't know. Finetune it on the answer toolCall("search", "that question", ...) or whatever.
Something close to the above is how all models with search tool capability are fine tuned.
All these hallucinations are despite those efforts, it was much worse before.
This whole method depends on the assumption that there is actually a path in the internal representation that fires when it's gonna hallucinate. The results so far tell us that it is partially true. No way to quantify it of course.
To fix that properly we likely need training objective functions that incorporate some notion of correctness of information. But that's easier said than done.