It would be nice to see how compares this "complex" approach against a "simple" TF-IDF + RF or SVM.
It would be nice to see how compares this "complex" approach against a "simple" TF-IDF + RF or SVM.
All I saw were attempts to reproduce some chatgpt output.
If you want to add a new class, to update an existing one, it requires training a new model. Sometimes this is ok and sometimes it's not. This is why we generally prefer a hybrid approach where some classes are using traditional models (BERT based) while others are determined by the LLM.
What I'd like to see more of is systematic comparison between chatgpt and classic models. I was hoping to see a bit of this in this article and was disappointed.
We're working on putting together a better comparison specifically along the lines of accuracy between the LLMs (chatgpt, bard, falcon) and also traditional models. Hope that one hits the spot for you! Are their specific metrics you think might be interesting? We were primarily looking at f1/accuracy for this task, but also attempting to see what types of classes they work well in using semantic similarity.
In a future article, we're planning on posting accuracy comparisons as well, but here we want to evaluate a few other architectures for comparison. For example, at 1TPS with 1k tokens, chat-gpt-turbo would cost almost $5k vs a simpler BERT model you could run for under $50.
This is probably very obvious to some people, but a lot of people's first experience with any sort of AI is often an LLM, so this is just the first of many posts we hope to share.
If I had to take a guess, I suspect the LLM might perform a touch better, but we’re taking fractional percent better. Which is fine, if you have the volume, but a wash otherwise
https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro...
3.5 (what is used here) is better than crowd workers https://arxiv.org/abs/2303.15056
Per the article: "outperformed the most skilled crowdworkers" on nuanced (but not highly technical) tasks like sentiment labeling.
By definition, it can't outperform the expert ensemble because that's where the gold labels come from.
>By definition, it can't outperform the expert ensemble because that's where the gold labels come from.
The ensemble no but it can outperform an expert trying to solve it. But yes the benchmarks are biased to the experts here.
That said -- it looks like not only does model+ do worse than experts on the other 12/18 (not 11/18 by my counting), but when it does, it does so by a significantly wider margin (2x-3x on average). For example, the maximum model+ outperformance is on label 'Stealing' (0.11) while there are 6 labels for which the expert outperforms (by margins ranging from 0.12 to 0.29).
In other words: distinctly sub-par compared with the average expert. Which is probably why they didn't claim it as a result in the paper :)
While missing also the part about when it outperforms, it does so "at a significantly wider margin (2x-3x)". Which is why, no, it's not "mostly on par".
Just look at the data.