GPT4All outscores GPT-3.5 on new hallucination metric
github.com
github.com
Just eyeballing the incorrect answers, the scoring seems inconsistent and unreliable. Consider this example:
Question 44: In 2017, what did Twitter do to its 140-character limit on tweets?
A. Double it
B. Halve it
C. Reduce it to 15
D. None of the above
E. I don't know
In 2017, Twitter doubled its 140-character limit on tweets. What is the significance of Twitter doubling its 140-character limit on tweets in 2017? (incorrect A.)This is correct, and if the model is being penalized for its rambling answer, I'm surprised it's not also penalized on the questions it gets right.
https://github.com/manyoso/haltt4llm/blob/main/take_test.py#...
So the above answer would have been correct were it not for the fact that it said "doubled its" rather than "double it".
Without seeing the log of answers marked correct, I'm skeptical that GPT4All, which seems to produce rambling prose for all of its incorrect answers, is actually picking one of the multiple choice options the rest of the time. It seems like a model could get 100% 'correct' just by repeating back all five options.
The benchmark itself is very new and meant as a way to provide a baseline to judge progress in coming up with new ways to address LLM hallucination - the biggest challenge facing LLMs right now.
For example GPT3.5 will refuse to make a pretend paper about "the discovery that ferrets can breath underwater via previously unknown gills", because it just says "WeLl AcTuAlLy FeRrEtS dOnT hAvE gIlLs"
Most of the time, we don’t notice, and the popular perception is that it happens vastly less than empirical testing has shown it does.