New metric for LLM hallucinations with results for GPT-3, GPT3-5 and Alpaca Lora
github.com
github.com
This project is an attempt to create a common metric to test LLM's for progress in eliminating hallucinations; the most serious current problem in widespread adoption of LLM's for real world purposes.
Method seems to be multiple choice trivia tests with real world answers, trick/fake questions where (I don't know) is correct answer, and 'None of the above' type questions. GPT-3.5 hallucinates but is much more willing to admit uncertainty than either GPT-3 or Alpaca Lora.
Hallucinations are a big one. “Obedience” is the next one that comes to my mind (how willing or unwilling the model is to comply with requests).