Secondly, it doesn't fix stupidity. A participant who earnestly takes the quality goals of the system to heart instead of focusing on maximizing their take (thus, obviously stupid) will still make bad classifications due to that reason.
1. I would expect any paid arrangement to include a quality-control mechanism. With the possible exception of if it was designed from scratch by complete ignoramuses.
2. Do you have a proposal for a better incentive?
2. Criticism of a method does not require that there is a viable alternative. Perhaps the better idea is just to not incentivize people to do tasks they are not qualified for
Agreed, and would add that it doesn’t fix other things like lack of skill, focus, time, etc.
An example is the output of the Amazon Turk “Sheep Market” experiment:
https://docubase.mit.edu/project/the-sheep-market/
Some of those sheep were really ba-aaa-ad.
edit: ugh. it's even worse, lmarena itself is a proprietary system, so the users presumably don't even get the benefit of an open dataset out of all this
I'm being (mostly) serious, suppose you're a stuffed ahort trying to boost your valuation, how can you work out who's smart enough to train your LLM? (Never mind how to get them to work for you!)
Then for LMArena there is the host of other biases / construct validity: people are easily fooled, even PhD experts; in many cases it’s easier for a model to learn how to persuade than actually learn the right answers.
But a lot of dismissive comments as if frontier labs don’t know this, they have some of the best talent in the world. They aren’t perfect but they in a large sene know what they’re doing and what the tradeoffs of various approaches are.
Human annotations are an absolute nightmare for quality which is why coding agents are so nice: they’re verifiable and so you can train them in a way closer to e.g. alphago without the ceiling of human performance
So we should expect the models to eventually tend toward the same behaviors that politicians exhibit?
Isn’t it fascinating how it comes down to quality of judgement (and the descriptions thereof)?
We need an LMArena rated by experts.
But at least the two examples of judging AI provided in the article can be solved by any moron by expending enough effort. Any moron can tell you what Dorothy says to Toto when entering Oz by just watching the first thirty minutes of the movie. And while validating answer B in the pan question takes some ninth-grade math (or a short trip to wikipedia), figuring out that a nine inch diameter circle is in fact not the same area as a 9x13 inch square is not rocket science. And with a bit of craft paper you could evaluate both answers even without math knowledge
So the short answer is: with effort. You spend lots of effort on finding a good evaluator, so the evaluator can judge the LLM for you. Or take "average humans" and force them to spend more effort on evaluating each answer
By being closed, they'll never be optimal.
Instead, finance bros are convinced by the argument that number goes up.
def is_it_true(question):
return profit_if_true(question) > profit_if_false(question)
AI will make it cheaper, faster, better, no problem. You can eat the cake now and save it for later.People hold falsehoods to be true, and cannot calculate a 10% tip.
We give them WAY too much credit by watching mostly the things they have been trained specifically to do and pretending this indicates a general mental competence that just doesn't exist.