Happy to answer any questions.
1,299 karma · joined August 1, 2012
Ex Google, Facebook, Twitter, Dropbox, Microsoft Research, MIT.
edwin[æ]surgehq.ai
https://www.surgehq.ai/blog
https://blog.echen.me
https://twitter.com/echen
https://github.com/echen
Happy to answer any questions.
One interesting question, though: if "LETS FUCKING GOOOOO YOU DINGBAT" were meant to be a combative insult, would someone still add a bunch of O's ("GOOOOO" instead of merely "go")? My intuition is that if combativeness were intended, "let's fucking go, you dingbat" would be more likely than "LETS FUCKING GOOOOO YOU DINGBAT", but of course it's a bit hard to say without that context.
Although from what we've seen, the amount context sensitivity matters really depends on the labeling task / application.
For example, when you're trying to label a tweet that's a reply, context matters even more than when you're labeling a parent tweet: it's often hard to understand what the reply tweet is talking about when you can't see the full thread, it can be hard to tell whether something is a joke or an insult when you can't tell whether the replier and original tweeter follow each other or not, etc. This is important because sometimes our customers don't realize this, and will send us tweet text by itself instead of a full tweet link.
It's also important because even if your models are using text alone (and not a richer set of context/features), there may be patterns in the text itself that an ML could pick up on that a human wouldn't without that extra context.
We also have another post on context sensitivity if you're curious: https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...
We work with a lot of the top AI/NLP companies and research labs, and do both the "typical" data labeling work (sentiment analysis, text categorization, etc), but also a lot more advanced stuff (e.g., search evaluation, training the new wave of large language models, adversarial labeling, etc -- so not just distinguishing cats and dogs, but rather making full use of the power of the human mind!).
Funnily enough, though, many ML engineers and data scientists I know (even those at Google, etc., who depend on human-annotated datasest) aren't familiar with these kinds of errors. At least in my experience, many people rarely inspect their datasets -- they run their black box ML pipelines and compute their confusion matrices, but rarely look at their false positive/negatives to understand more viscerally where and why their models might be failing.
Or when they do see labeling errors, many people chalk it up to "oh, it's just because emotions are subjective, overall I'm sure the labels are fine" without realizing the extent of the problem, or realizing that it's fixable and their data could actually be so much better.
One of my biggest frustrations actually is when great engineers do notice the errors and care, and try to fix them by improving guidelines -- but often the problem isn't the guidelines themselves (in this case, for example, it's not like people don't know what JOY and ANGER are! creating 30 pages of guidelines isn't going to help), but rather that the labeling infrastructure is broken or nonexistent from the beginning. Hence why Surge AI exists, and we're building what we're building :)
For example, a spammy email that is strangely uncaught:
Subject: ///Y0urFREEFlashlight(NeedYourAddress)///1851
Sender: UBzjwbFb-lyZSC3-noReply@gwhsi.lairpro.com
In short, one way to prevent your language models from devolving into violence (with extremely high safety guarantees) is by building "AI red teams" of labelers who try to trick it into generating something violent. Then you train your models to detect those strategies (just like other kinds of red teams find holes in your security, which you then patch). Then your "red data labeling teams" find new strategies to trick your AI into becoming violent, you train models to counter those strategies, and so on.
For example, here are some queries in my history - would the average person understand what I'm looking for and what a good search result would be?
"neovim worktrees"
"django models through"
"pandas econdb"
1. Clicking on a link is often a negative signal. If you're searching for "when was barack obama born?", hopefully you get the answer on the search results page itself, so you never have to click. Or, the caption for each search result should have a snippet containing the answer.
2. If you stay a long time on the search result you click on, is it because the information is hard to find or because the page is deep and useful?
I used to work on Search measurement at YouTube and Facebook, and run a human evaluation platform now. Here's an example what they can look like when you use high-quality raters: https://demos.surgehq.ai/google-vs-bing/
Similarly, a lot of the training data/features ML engineers use ignore context -- for example, a Reddit comment may seem hateful in isolation, until you realize the subreddit it's in changes the meaning entirely (https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...).
Regarding your point, we actually do a lot of "adversarial labeling" to try to make ML models robust to countermeasures (e.g., making sure that the ML models train on word letter substitutions), but it's pretty tricky!
When we started optimizing for watch time at YouTube, for example, our algorithms started suggesting longer videos for the sake of longer videos, and videos with racy thumbnails.
Similarly, optimizing for engagement at Facebook led to low-quality clickbait, and Hooter's appearing as the top search result when you searched for restaurants in Houston.
Experiments that increased favorites and replies at Twitter invariably increased toxic content as well.
So while watch time, engagement, and replies would always go up -- were these really the products we wanted to build? What happened to Facebook's original mission of connecting users with their friends and family? What did "favorites" have to do with being the platform for public conversation at Twitter? A lot of work at these companies is spent measuring active users, but where were the dashboards measuring progress to these broader goals? It's easy to become blinded by standard metrics and lose sight of the original product principles that made us stand out -- and I say this as a data scientist at heart!
So could we figure out a metric that was better tuned to human values and the product mission we cared about, but also fast, rigorous, and easily measurable? After all, we still need our A/B tests, ML objective functions, and OKRs! This question is particularly important today, with all the troubles that social media platforms face, and I wrote up an approach that I've often worked on.
100%. It's crazy how often context is missing from datasets! We wrote a separate blog post recently about this too: https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...
Although -- and I say this having done a lot of my graduate coursework in linguistics -- I don't think having a linguistics background is particularly needed (unless you're doing specialized annotation, like creating syntax trees or tagging phonemes in Praat), outside of you being more likely to enjoy thinking about the nuances of language.
For example, to pick a bit on the Google Emotions dataset again, it's difficult to label this message...
“Also Republicanism is a belief system. It’s taught and handed down like religion. Conservative talk radio is its evangelism.”
...unless you're familiar with US politics. Hence why it was labeled as APPROVAL by the non-US annotators, even though it's criticizing Republicans.
We deal with these hairy problems a lot too. Even before ML proper, getting the definitions right is very tricky. Should a comment that's polite and positive on its own, but supportive of a toxic parent post ("I love Nazis!" -> "I agree!"), be treated as toxic? Is "toxic counterspeech" equally toxic? What about a comedian making fun of an actor's nose, and does it depend on whether the joke is to that actor directly vs. merely referencing them in the third person? etc.
For these reasons, I actually personally like the fact that the Jigsaw annotation guidelines are very high level (as opposed to long and prescriptive) -- it lets the data capture the spectrum of "human preferences" on its own (at least, it does when you can trust that the annotators are able and trying to do a good job).
Hopefully people also don't need to be at the level of a professional linguist to label messages like "this is fucking awesome" correctly!
And great point on context. For example, the GoEmotions dataset didn't present labelers with the actual post or subreddit the message came from -- just the text itself. That makes it really difficult to label something like "his traps hide the fucking sun"! But once you see the comment in its original context https://www.reddit.com/r/nattyorjuice/comments/aee3wx/olympi..., and know that it's in the /r/nattyorjuice bodybuilding subreddit, it's much easier to realize that this is talking about someone's large muscles.
Perhaps this is okay when your datasets are high-quality and representative of the real world, but they're usually not. For example, many toxicity and hate speech datasets mistakenly flag texts like "this is fucking awesome!" as toxic, even though they're actually quite positive -- because NLP datasets are often labeled by non-fluent speakers who pattern match on profanity.
(So is 99% accuracy or 99% precision actually a good thing? Not if your test sets are inaccurate as well!)
Many of the new, massive scale language models use the Perspective API to measure their safety. But we've noticed a number of Perspective API mistakes on texts containing positive profanity, so this post was an attempt to explain the problem and quantify it.
For example, my first week at Google/YouTube, I was in a New Hire meeting with our VP. Someone asked about profitability, and he responded that Larry said we didn't have to worry about revenue yet, since the main goal was user growth/happiness, and revenue could come later. Which I thought was fascinating, considering how big YouTube already was at the time (in 2013)!
Though I think this changed a year later, and I find YouTube ads a poor experience compared to Instagram and TikTok -- which aren't merely "better than the rest", but stuff I actually enjoy watching.
I used to work on Search and Search Measurement at YouTube, Twitter, and Microsoft, so I thought it would be fun to move beyond anecdotes, grab some data, and do a quantitative analysis.
tl;dr I didn't have historical data, to see how Google Search has trended over time. But compared to Bing, Google still generally outperforms -- although some of its failures are pretty surprising!
If there are particular areas of Google Search that people are interested in digging into, give a shout -- I love running these kinds of search / human eval analyses.
[1] https://news.ycombinator.com/item?id=29772136 [2] https://news.ycombinator.com/item?id=29417061 [3] https://news.ycombinator.com/item?id=29392702