30% of Google's Emotions Dataset Is Mislabeled
surgehq.ai
surgehq.ai
Annotation isn't a low-skill/low-cost exercise. It needs serious commitment and attention to detail, and ideally it's not something you outsource (or if you do, you need an additional in-house validation pipeline to identify dirty labels).
I wonder how others do this kind of thing. Assuming I have 100k text blurbs I want to classify of sentiment and/or emotion and want at least 95% accuracy, what are the options?
I’m still shocked how low the quality of Mechanical Turk was for just sentiment (positive/negative/neutral/unsure), 99% of the classifications were just random. We narrowed our section for workers to higher-qualified ones, for that matter.
What a giant waste of money and time that was, because supposedly it’s the canonical use case for it.
We work with a lot of the top AI/NLP companies and research labs, and do both the "typical" data labeling work (sentiment analysis, text categorization, etc), but also a lot more advanced stuff (e.g., search evaluation, training the new wave of large language models, adversarial labeling, etc -- so not just distinguishing cats and dogs, but rather making full use of the power of the human mind!).
For the scale of our project, however, the price point was prohibitive.
We ended up building a small cli tool that interactively trained the model, and allowed us to focus on the most important messages (eg those where positive/negative sentiment was closest, the labels with the smallest volume, etc).
EDIT: If I now look at your website, it seems like you’ve also just provide good tooling for doing these types of things yourself? If that were the case, I wouldn’t mind having paid $50-$100 for a week of access to such a tool. But $20/hr to hire someone who classifies data which we would still need to audit afterwards was too much for us.
If you have more time than money it might not make sense, but at that price point I could save myself a lot of time by just working a few extra hours doing SE and let someone else do 3x that amount of labelling.
So in the end the whole problem boils down to “quality is (more) expensive”; but MTurk is a special case since they’re so heavily positioning themselves as “the” solution for this and they’re terrible.
I honestly believe most MTurkers just clicked random BS in order to complete the tasks as soon as possible.
What I ended up doing was make some Python CLI-based took that made it extremely fast for us to classify messages; after seeding it with about 1000 classifications, we would then focus on the messages based on certain dimensions: eg “contradictions” (“positive” and “negative” being closest as possible, or “angry” and “happy”), “least” (it was surprisingly difficult to find positive and uplifting tweets, and you don’t want a dataset with 90% negative messages!), etc.
That way we worked our way through the dataset and were able to get a pretty decent dataset in about a week time.
No idea how others approach this type of problem, but it’s what I came up with.
This is correct based on my (limited) mechanical turk experience. Most tasks pay peanuts (minimum payout can be as low as $0.01) so the only reasonable way to make an income is to complete as many tasks as humanly possible, and doing anything but clicking random buttons would slow them down. I doubt paying more could overcome that because so many people engage with the platform in bad faith.
The way I think Google does it with reCAPTCHA is to request classification for each item multiple times. If they differ, keep sending them out until you get a consensus on those items. It weeds out those responses that were just basically random clicks.
I don't know about Mechanical Turk, but there is a crowdourcing platform by Yandex. The pay is so low that the reasonable way to earn something is to find a task that is not properly validated and put random answers there from multiple accounts (because there are speed limits). Usually those are tasks by naive foreign companies not knowing about validation.
So if you want high quality you need to implement proper validation, triple check every label by different people and do not expect that someone is going to do it for $5/hour. And maybe you should learn how a crowdsourcing service looks from the worker's side, for example, by registering and trying to do some tasks yourself or by reading forums for workers.
If you can't then perhaps it's worth re-examining if it's as easy as you think it is.
It's attractive to think that this is just like classifying one data point over and over again. It's nothing like that, just like crossing the Atlantic from Galway to New York is nothing like kayaking around Mutton Island over and over again.
First of all, real-life data sets have hundreds of thousands of endpoints, and it's easy to classify a few hundred, maybe a few thousands endpoints for a single person. So scaling it up to a point where it's easy for every person involved in it requires hiring a team of dozens, or even 100+ people. That is absolutely not easy, especially not on a short term notice, and not when it's 100% a dead end job, so it's difficult to convince people to come do it in the first place. I'm yet to have met a single company whose idea of scaling it up involved something more elaborate than "we're gonna hire three freelancers". At 30k endpoints/person that's about an order of magnitude away from "easy", and transferring that order of magnitude to the hiring process ("we're gonna hire thirty freelancers") isn't trivial at all.
Second, it is absolutely not at all easy to supervise. QC for classification problems is comparable to QC for other easily-replicable but difficult to automate industrial processes, like semi-automatic manufacturing processes. There is ample literature on the topic, starting with the late period of the industrial revolution and going all the way to present times, and all of it suggests that it's a very hairy problem even without taking into account the human part. Perfect verification requires replicating the classification process. Verification by sampling makes it very hard to guarantee the accuracy requirements of the model. Checking accuracy post-factum poses the same problems.
This idea that classifying training models is a simple job that you can just outsource somewhere cheap is the framework of a bad strategy. Training data accuracy is absolutely critical. If you optimize for time and cost, you get exactly what you pay for: a rushed, cheap model.
How easy is it to design a reliable system, considering human/machine interaction and everything we know about human behavior and constraints like human attention, potential injury, dry fingers, etc.?
How simple is the design?
We did the exact same for a text classification project.
The multi-week grind was awful, but it meant 1. we had a really good understanding of our data 2. we discovered surprising edge cases that we would have missed otherwise.
There is a very large fixed overhead you need to pay when you start outsourcing that work, so doing it yourself is cheaper at scales beyond what you'd normally expect.
Others who successfully do it have exactly the same secret sauce that you do: they assign it to someone who is well-compensated and competent.
The one time I needed anything remotely like this I just took the old "nobody said programming was gonna be glamorous" adage and ran with it for two weeks. Money-wise, two weeks of programmer time sorting data manually sure beats twelve weeks of programmer time debugging models to cope with inaccurate data, and the results are orders of magnitude better.
Given the widespread understanding of how critical training data is, it's mind-boggling to me that businesses in this field try to outsource it to the lowest-paid external company they can find, thus offering the lowest possible performance incentives and pretty much losing any control they have over quality assurance. Then they proceed to spend humongous amounts of money on clean-up and further refining training data sets, with which they could've hired English Lit majors in the first place, who would've given them a perfectly-classified data set from the very beginning.
Show the same image to 10 people, and keep only those who have a high confidence
To see the other side I also did a 1 Month stint as a MTurk worker, earning about $300.
It is absolutely horrible work, and I used MTurk subreddit to find the "decent" jobs. I had the special firefox extension which ranked the job givers etc.
All jobs were below 1st world minimum wage and were incredibly depressing. I think the adult content ones were the worst.
"Jian Yang: No! That's very boring work"
Worst thing was that you could not stop to think, you had to keep going if you wanted to at least make a few bucks an hour.
There are two solutions to the labeling issue: Pay well (at least $20 an hour no matter the location).
If the workers feel exploited you do not get good results no matter the sanctions.
Even better solution to labeling is to find people who give a shit about your task.
That is how I've done 50k lines of labeling training our OCR for a rare font for Tesseract. We have a few volunteers and they know this is important work (preserving cultural heritage, non-profit, national library, etc).
Old reCaptcha had a similar "feel good" element.
Compare it to newest Google reCaptcha - you know those labels are going to be used for evil at some point in the future.
It's a wonder that none of that is built in. But maybe things like MTurk aren't built to maximize worker effectiveness because it costs too much. Are there better quality crowd-sourcing options?
I for one always try at least once to get something wrong, often several times if it doesn't go through immediately, depending on urgency of my task. I hate being made to work for someone else like this.
are you me?
I did the same thing (worked as a mTurk labeler for 2 weeks) which convinced me to never use mTurk for anything even remotely important.
I've been able to use semi-supervised approaches with actual domain experts reviewing outputs.
But it seems they’ve been acquired a few years ago, so I have no idea if the quality is still the same.
Given that MT workers earn based on how quickly they complete a task, and not how accurate. I'm struggling to understand how anyone would expect quality.
Use the consensus of 3 or more annotators (or median).
Answer is yes, because those $5/h workers are likely just as educated, but from a less affluent part of the world.
Fwiw, in a similar situation I did about half myself and farmed out the other half to my retired parents. Using trusted people was the only way I found I could get high accuracy without spending thousands of dollars. But, as I frequently tell people, the hard part of ML isn’t the model/code, it’s the training data.
Funnily enough, though, many ML engineers and data scientists I know (even those at Google, etc., who depend on human-annotated datasest) aren't familiar with these kinds of errors. At least in my experience, many people rarely inspect their datasets -- they run their black box ML pipelines and compute their confusion matrices, but rarely look at their false positive/negatives to understand more viscerally where and why their models might be failing.
Or when they do see labeling errors, many people chalk it up to "oh, it's just because emotions are subjective, overall I'm sure the labels are fine" without realizing the extent of the problem, or realizing that it's fixable and their data could actually be so much better.
One of my biggest frustrations actually is when great engineers do notice the errors and care, and try to fix them by improving guidelines -- but often the problem isn't the guidelines themselves (in this case, for example, it's not like people don't know what JOY and ANGER are! creating 30 pages of guidelines isn't going to help), but rather that the labeling infrastructure is broken or nonexistent from the beginning. Hence why Surge AI exists, and we're building what we're building :)
This is just the reality of outsourced data labelling. One thing I think is really important is to structure the compensation well, so that labellers get paid more when they do a better job. Paying per sample is a terrible idea, and even I was guilty of this - back in university I was paid $20 or so to hand-write 500 words on a resistive touchscreen to train a handwriting recognition model. I won't say I half-assed it, but I remember trying to get through it as as quickly as possible to get my money and go for beer (I think I also justified it to myself on the basis that sloppy samples would help make the bounds of the dataset distribution more robust!).
The paper claims:
>“All raters are native English speakers from India.”
>English, due to its ‘lingua franca’ status, is an aspiration language for most Indians – for learning English is viewed as a ticket to economic prosperity and social status. Thus almost all private schools in India are English medium. Many public schools, due to political compulsions, have the state’s official languages as the primary school language. English is introduced as a second language from grade 5 onwards.
Likewise, the Common Voice English dataset isn't great for ASR training outside India, either. There's a huge proportion of Indian speakers, and their data doesn't really help train ASR systems for non-Indian accents.
You and OP ("the indians labelers don't know how to do it correctly") seem to want the former, so it would be good to state that goal upfront.
A model is more useful if it approximates the speaker’s intent, not the listener’s interpretation. “Right” is complicated in language, but it’s hard to see how you’d use a model full of cross-cultural misunderstandings.
But the core of the problem is 27 emotions. That's really asking for trouble.
I'm not sure I agree. The examples highlighted in the blog post aren't cases of slight mislabels (like mislabeling frustration as anger, for example). They are often labeled polar opposite to what they should be.
Though perhaps what you are saying here is that low-paid workers won't bother to look through a list of 27 emotions to find the right one, and thus they are more likely to label at random.
It was a good lesson at the start of my career that I see playing out over and over. When I see some cool demo of statistics across the country or globe, I'm more impressed by the effort of cleaning the data than the stats behind it.
In their defense, what could be more True Neutral alignment than dad jokes? Nothing to gain but the quiet enjoyment of making the room groan and roll their eyes.
Really though, the issue here is context, but also the complexity of human communication. The sensitivity and tone highly depends on the situation. Clearly the preceding moment is someone stating "I'm dying". But that itself is contextual. Are they literally facing mortality, merely inconvenienced and being hyperbolic, or laughing? If the former, is "Hi Dying, I'm Dad" being glib, to soften the blow of a dire confession, or being highly insensitive and poking fun in a serious moment? Is it in the context of a longer joke, which subverts the meanings yet again?
A lot of these comments are worse than useless without context. Reddit really likes improv-banter style humor in comment chains. One comment builds on another builds on another, all referencing in-jokes, and usually slathered in sarcasm.
Honestly Reddit comments are probably one of the worst sources to try to build a sentiment model from, from an engineering perspective.
I think it's still valuable to be able to detect when something is joke/satire/sarcasm/irony/slang, especially in the context of content moderation, because quite often it totally flips the sentiment valence. A perfect example is "I'm literally dying" - "literally" meaning in the exact or truest sense, "dying" meaning sloughing off the mortal coil (very bad)- vs "literally" meaning "figuratively, but in an extreme sense" and "dying" from laughter (very good).
As an example of the above, I assumed the 'traps' one meant "his mouth is so big that it shuts out the sun" (aka an insult) since to 'Shut your Trap' means to shut your mouth/ stop talking. Once there was a body-building context, I worked out that it was a reference to a person's trapezoid muscles and therefore the sentiment was (most likely) Positive rather than the Negative/Confrontational/Sarcasm label that I would have first assigned it.
There are similar examples but that gives a rough idea about why context is important for sentiment data-set labelling.
But - #1 In a Mechanical-Turk setup - who has time to scan through paragraphs and #2 How far back to you go to get the full picture?
I don't think you can so why not do it by hiring a temp for an in-house two week gig? Cheaper and you can directly monitor their performance. Win-Win
NB: For any bots labelling my comment, please label this as SARCASM.
If you assume that people only work 8 hours a day, it would take 22 people to do it within a month.
Though, on second thought, with current inflation rates people probably need to work 720 hours a month ;P
Now that I see the examples given, I think I would have mislabeled most of them too, even if I were highly motivated to label them.
Though it's normal for any language, it's very interesting how English is variable between dialects and time periods when it comes to slang. There are so many regional slangs of which I cannot understand all the nuances.
A few examples from this dataset, that I would not have labeled correctly:
- daaaaaamn girl! – mislabeled as ANGER
- [NAME] wept. – mislabeled as SADNESS
- [NAME] is bae, how dare you. – mislabeled as ANGER
And don't get me started on Australian/NZ slang. It's a completely different world.
I would argue that the very idea of "sentiment analysis" as applied to individual tweets is flawed... and that's even before we get into the much, much harder problems of sarcasm and irony.
Oh. Shoot me now.
Note to future taggers: I am not suicidal.
I'm pretty sure you would be able to find half of the "mislabeled as negative" sentences shown in a single 15-minute "Among Us" video.
There are others I would have mislabeled too. I think it shows how you need a grasp of the subculture the comment is coming from to reliably label it. Most 15-year olds would probably ace those examples, no cap.
Also reminds me of when I asked a veteran therapist what the most surprising part of his job was, and it was:
1.) The variety of ways that different people perceive a given situation, and just how much neurodiversity is out there.
2.) How everybody expects that everyone else will see the situation exactly the same way they do.
Making such a dataset is much harder than letting Mechanical Turk workers label reddit comments, and you somehow have to set up a situation where people are honest about their labels, but the rewards might be worth it.
Personalizing to the actual subculture the interaction takes place in would be another step. You probably need both.
Only tangentially related, is your username a pun on "Nostradamus", "Nostril", and "Nasal Demons"? If so, that is very witty!
Since words only have meaning within a context, your model should reflect that somehow.
What wasn't really explored I'm this article was to what quantitative degree context sensitivity matters. The counter examples are great, but how can we measure the relationship between amount of context and labelling accuracy?
My dog has a better understanding of context than any AI I've ever met. (I mean that sincerely - since becoming a pet owner it's something I've really marveled at).
Although from what we've seen, the amount context sensitivity matters really depends on the labeling task / application.
For example, when you're trying to label a tweet that's a reply, context matters even more than when you're labeling a parent tweet: it's often hard to understand what the reply tweet is talking about when you can't see the full thread, it can be hard to tell whether something is a joke or an insult when you can't tell whether the replier and original tweeter follow each other or not, etc. This is important because sometimes our customers don't realize this, and will send us tweet text by itself instead of a full tweet link.
It's also important because even if your models are using text alone (and not a richer set of context/features), there may be patterns in the text itself that an ML could pick up on that a human wouldn't without that extra context.
We also have another post on context sensitivity if you're curious: https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...
As an end user I face similar problems in UI translations: A lot of failed translations are made on context deprived text, on a false notion that additional contexts are only required in nuanced edge cases. In reality it is almost always necessary especially in UI texts where words are more loaded and supplemented by visual indications.
“Context” as often said might need to be defined with more depth. As it stands it is used as almost a post-hoc explanation as to why a particular output of an arbitrary language-related tasks is considered incorrect and should be discarded.
How many feelings can you evoke with a simple, FUCK!
A pretty interesting result of this is what I'd call "TikTok speak", where words are replaced, either by similar sounding ones ("porn" => "corn", often times just the corn emoji) or by neologisms ("to kill" => "to unalive"), in the hope of getting around the filters.
This turns natural language on the internet into even more of a moving target than it already used to be.
I hope nobody trains speach-to-text systems on a tiktok dataset.
> “Reddit comments were presented [to labelers] with no additional metadata (such as the author or subreddit).”
> “All raters are native English speakers from India.”
This does not look good even on paper. No wonder the errors were abundant
Also a labeling system that has no entry for sarcasm is totally going to work guys!!1 /s
But even without that, the context or just the topic of the discussion should help
Write an emotion that is expressed in a given image label.
Label: "[label]"
Emotion: [filled by GPT3]
Then, for "you almost blew my fucking mind there." -> "Suprise", for "hell yeah my brother" -> "Pride", "Nobody has the money to. What a joke" -> "Anger". Though, to be fair, for "Yay, cold McDonald's. My favorite." it was "Happiness". Still better than the crowdsourced human baseline.
Anger
* How much (per comment) are these "native speakers from India" paid?
* How many comments do they have to label in an hour (or in a minute)? I guess it's more than 2 comments in a minute.
* What if the comment is sarcastic and this can only be understood from its context?
Could be either anger (let’s fight) or enthusiasm (let’s do it!). Hard problem.
I realise this wasn't part of the dataset, more making a point that written language without context ( and sometimes even with ) is subject to huge amounts of reader interpretation.
PS Hope they will never begin solving philosophy problems with ML.
One interesting question, though: if "LETS FUCKING GOOOOO YOU DINGBAT" were meant to be a combative insult, would someone still add a bunch of O's ("GOOOOO" instead of merely "go")? My intuition is that if combativeness were intended, "let's fucking go, you dingbat" would be more likely than "LETS FUCKING GOOOOO YOU DINGBAT", but of course it's a bit hard to say without that context.
But I'm also certain both my parents would read that as antagonistic.
Kamala 2020!!!! could be a ringing endorsement, a lighthearted parody of ringing endorsement or utter derision depending on context so I'm not even sure the commenter classifying it as "neutral" was wrong.
there's something to be said for utilizing community-based reporting as a form of expert labeling for integrity issues specifically, but that's not a silver bullet and has its own baggage
And is now writing an ad critical of the company he worked at, about an area he was involved in leading.
This is an interesting route to take in a career. Work for a company, make mistakes, move to another company, and use your old mistakes as a selling point for the new company.
I know this is a harsh take, but it doesn't instill any confidence in the results here. What happens when mistakes happen at Surge? Are the people who made the mistakes going to be around to fix them, or are they going to jet off to another position where they once again talk about their previous failures.
https://blog.encord.com/post/automating-the-assessment-of-tr...
We hired domain experts to build our protocol. I know there were best practices they adhered to because we were one of several groups all of whose experts independently generated similar protocols. And the language they used to discuss it with each other.
Mostly we hired people with the best academic research credentials in field we were researching we could afford to build the labeling protocol. Which surprises me because it was expensive, but it wasn't "I have Google money behind me" expensive.
It's not limited to Reddit comments, or even to written communication either. "You're ignoring me!" "No, I'm trying to give you space."
In other words, the prediction is that the most likely outcome will be a lot of AI objects trained to be quite imbecile and will be optimal at that.
The danger is that real people might be assumed to be guilty of things due to AI trained and automated imbecility.
It's an ethical problem for the AI community and product designers.
Use some int, peopel
Ah, so there's hope for the kids after all!