When LLM judges agree, should we believe them?
amazon.science
amazon.science
Folks should sue in a class-action lawsuit, any legal firm worth their beautiful walnut desks would seriously be happy take on that constitutionally backed mission. =3
"Nonsensical?" They commit violent crimes, they get prosecuted for said violent crimes, and are serving prison sentences for those crimes. The algorithm picks up on this trend using the same logic that insurance actuaries use, which has also been largely neutered by critical theory.
What even is the argument here-- they're all innocent? Cops are ignoring piles of dead white people and their white murderers to only go patrol brown neighborhoods? We both know neither claim is true. The usual complaint is that cops avoid their neighborhoods and/or are lazy in investigating the crimes they report. The idea of overpolicing has always been a Marxist double-bind...nonsensical, I daresay.
I'm mostly talking about random coding errors.
You solved one of the largest problems with current LLMs. How is it possible that nobody tried that before?
Because they do. There are already LLMs checking outputs of other LLMs, the bullshit answers that you see are the results of failures on that checks. If you remove all checks LLMs will create hallucinations even more often.
Drawing a line red to split up the image then has them answer correctly.
Their failure modes are highly correlated.
Cohen, Hamri, Geva & Globerson, "LM vs LM: Detecting Factual Errors via Cross Examination": Cross-examination "detects over 70% of the incorrect claims while maintaining a high precision of >80%".
Citation needed
Try it yourself. Get one to hallucinate, then paste that text into a new window and ask it to verify the facts.
EDIT:
Also - Cohen, Hamri, Geva & Globerson, "LM vs LM: Detecting Factual Errors via Cross Examination": Cross-examination "detects over 70% of the incorrect claims while maintaining a high precision of >80%".
So 70% for ANY error, not just hallucinations.
But when you ask the SAME LLM (with the same context) they remember the hallucination so it doesn't work. You have to use a fresh one without the same context
Is it perfect? No. But it does what I asked it to do.
None — a bass in water is a fish, and fish are notoriously bad at music.
The instrument version plays four strings as standard (five and six-string basses exist for players who want to go lower or higher), and it prefers to stay dry.
Seems like a pretty good answer to me!A fish can play with as many strings as it finds, but only one when on a hook. Yet this too is an incorrect answer, as it again ignores the ambiguity in the phrasing. =3
If a LLM based chat bot does ever answer it correctly, than you know with a fair degree of certainty it was content moderators stepping into the chat. Have a wonderful day. =3
https://en.wikisource.org/wiki/The_Poems_of_John_Godfrey_Sax...
While the LLM spits out each ambiguous context search result, it never answers the actual query without a human cheaters help. =3
For info: the advisor(s) available are higher end models. For example: you use sonnet, the available advisors are opus and fable. If you use Haiku, the advisor are sonnet, opus and fable.
It shouldn't even be a debatable question.
> Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
Because there's no "discounting of opinions". They are running a separate LLM to "score" opinions. And the result is still "no" regardless of "lineages" or "sources".
And the end of the article leads me to believe that the entire article and approach is LLM-induced garbage:
--- start quote ---
<Following a list of LLM-like suggestions>
When LLM judges agree, we should ask why. Sometimes agreement is independent evidence. Sometimes it is a shared blind spot. A good aggregation method should be able to tell the difference.
--- end quote ---
It's like the idea of political districting that thinks that the aim should be to balance each district between "the two" political parties. You're not doing anything but institutionalizing two political parties and constant conflict. You're setting the range of acceptable opinions, then choosing at random between them. Even more relevantly: when both institutionalized parties have the same opinion, it's considered the correct opinion no matter how much or how little public support it has.
I've never heard anyone explicitly advocating for gerrymandering in favor of conflict/variance before? Is that a real thing?