LLMs can label data as well as human annotators, but 20 times faster
refuel.ai
refuel.ai
1) it will work well, at first, and only become low-quality after they (and their budgets) have become accustomed to paying 1/20th as much for the service
2) even if they pay for "human" labeling, they will go for the low cost bid, in a far-away country, which will subcontract to an LLM service without telling them
3) "hey, we should pay more for this input, in order to avoid not-yet-seen quality problems in the future", has practically never won an argument in any large corporation ever. I won't say absolutely 0 times, but pretty close.
Long story short, the use of LLM's by Big Tech may be doomed. Much like how "SEO optimization" turns quickly into clickbait and link farms if there is not high-urgency and high-priority efforts to combat it, LLM's (and other trendy forms of AI that require lots of labeled input) will quickly turn sour and produce even less impressive results than they already do.
The current wave of "AI" hype looks set to succeed about as well as IBM Watson.
For us human labeling is suprisingly cheap, the main advantage of GPT-4 would be that it would be much faster, since scams are always changing we could general new labels regularly and be continuously retraining our model.
In the end we didn't go down that route, there were several problems:
- GPT-4 accuracy wasn't as good as human labelers. I believe this is because scam messages are intentionally tricky, and require a much more general understanding of the world compared to the datasets used in this article which feature simpler labeling problems. Also, I don't trust that there was no funny business going on in generating the results for this blog, since there is clear conflict of interest with the business that owns it.
- GPT-4 would be consistently fooled by certain types of scams whereas human annotators work off a consensus procedure. This could probably be solved in the future when there's a larger pool of other high-quality LLMs available, and we can pool them for consensus.
- Concern that some PII information gets accidentally sent to OpenAI, of course nobody trusts that those guys will treat our customers data with any level of appropriate ethics.
All the datasets and labeling configs used for these experiments are available in our Github repo (https://github.com/refuel-ai/autolabel) as mentioned in the report. Hope these are useful!
From benchmarking, we've been positively surprised by how effective few-shot learning and PEFT are, at closing the domain gap.
"When it encounters novel data (value) it will likely perform poorly" -- is that not true of human annotators too? :)
Some humans have intelligence and reasoning abilities. No LLMs do :)
And with the topical depth say ChatGPT4 has, I would think these labels have more value, although just as with humans some validation and verification steps are required.
I’m not sold this has directional value.
The need for labeled data for any kind of training is a constant though :)
For each of these datasets, we specify task guidelines/prompts for the LLM and human annotators, and compare each of their performance against ground truth labels.
Fixed it for you.
Is there some noise in these labels? Sure! But the relative performance with respect to these is still a valid evaluation
The cost, hovewer, goes down to 1/1000 for cheaper models (and they can still 100% some of the datasets), meaning that for the same price as a human annotator you could have 1000 parallel realtime LLM annotators