"Making stuff up and being confidently wrong are well known side-effects of LLMs and there are many techniques to change this behavior."
I didn't know there are many techniques to mitigate this
I didn't know there are many techniques to mitigate this
More generally, if you teach the model to reject nonsense questions and admit if it doesn't know something it's more likely to do that
A trivial idea - you can use GPT-3 to inject bullshit/hallucinations into real text. Then train the model to solve the reverse task, of detecting bullshit in input text.