https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro...
3.5 (what is used here) is better than crowd workers https://arxiv.org/abs/2303.15056
https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro...
3.5 (what is used here) is better than crowd workers https://arxiv.org/abs/2303.15056
Per the article: "outperformed the most skilled crowdworkers" on nuanced (but not highly technical) tasks like sentiment labeling.
By definition, it can't outperform the expert ensemble because that's where the gold labels come from.
>By definition, it can't outperform the expert ensemble because that's where the gold labels come from.
The ensemble no but it can outperform an expert trying to solve it. But yes the benchmarks are biased to the experts here.
That said -- it looks like not only does model+ do worse than experts on the other 12/18 (not 11/18 by my counting), but when it does, it does so by a significantly wider margin (2x-3x on average). For example, the maximum model+ outperformance is on label 'Stealing' (0.11) while there are 6 labels for which the expert outperforms (by margins ranging from 0.12 to 0.29).
In other words: distinctly sub-par compared with the average expert. Which is probably why they didn't claim it as a result in the paper :)
While missing also the part about when it outperforms, it does so "at a significantly wider margin (2x-3x)". Which is why, no, it's not "mostly on par".
Just look at the data.