Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
Does no one else hate it when this happens (especially when on a handheld device)?
> If so, where do they indicate they failed to randomize/blind the raters?
Win rate if user is under time constraint
This is hard to read tbh. Is it STEM? Non-STEM? If it is STEM then this shows there is a bias. If it is Non-STEM then this shows a bias. If it is a mix, well we can't know anything without understanding the split.Note that Non-STEM is still within error. STEM is less than 2 sigma variance, so our confidence still shouldn't be that high.
If 10% of people just click based on how fast the response was because they don't want to read both outputs, your p-value for the latter hypothesis will be atrocious, no matter how large the sample is.
But I'm glad you pointed that out, I now suspect that is responsible for a large part of the disagreement between "huh? a statistically significant blind evaluation is a statistically significant blind evaluation" vs "oh, this was obviously a terrible study" repliers is due to different interpretations of that post. Thanks. I genuinely didn't consider the alternative interpretation before.
Couldn't this be considered a form of preference?
Whether it's the type of preference OpenAI was testing for, or the type of preference you care about, is another matter.