Are popular toxicity models simply profanity detectors?
surgehq.ai
surgehq.ai
Similarly something that might be a cat call such as "Get your baps out" (shows us your breasts), could also be used by a baker since a "bap" is a type of bread roll in a slightly cheeky advert as most people are aware of the pun.
How are you going to train an AI to know the context that the person might be talking about bread instead of a woman?
Has anyone realised yet that almost all of this folly? I suppose not when there is money to be made.
This is not a sustainable strategy for society. It is of course impossible for algorithms to parse perceived intent by the lowest common denominator, and attempting to do so is nothing more than defensive legal posturing.
I think that both intent and perception are important, and they are both judged in the more modern understanding of "toxicity". Malintent is a direct "strike". But, lack of malintent doesn't mean that perceived offense should be entirely ignored.
Let's imagine a scenario: as in the example above, an English person says "can I burn a fag" and an American person believes they are casually thinking about hurting a gay man. The English person meant no harm; but, the American person is not at fault for not understanding this. They can complain, and the English person should explain their intent. The American person should learn and remember the meaning of this expression in British English, and the English person will need to remember that this expression may sound offensive to American listeners, and avoid it in such scenarios.
Basically, the important aspect is that a misunderstanding is not always the fault of the listener (which the intent-only model suggests: you took offense, but I didn't mean that, so you're silly for feeling offended). It is of course also not only the fault of the speaker (which the perception-only model would suggest: I felt offended by what you said, so it doesn't matter what you meant, you owe me an apology or more). Both parties are involved in a miscommunication, if indeed there was one.
If you call someone out for "offensive speech" and your offense rests on your own misunderstanding, it is absolutely your fault.
As a white guy wearing a suit, I got out of there without being charged. If I weren’t… I can imagine it could have gone very differently.
Also, “is was part”? So much for being clear.
Unfortunately, we seem to be in a place where bad intent is taken for granted and merits an aggressive reaction rather than placing any responsibility on the offended party to be civil and open-minded in their response to be perceiving offense.
In the case of "burning a fag", our hypothetical American has the option of inquiring about the meaning of an unfamiliar idiom rather than jumping straight to calling out the speaker.
bum is not burn.
FYI, it's combination of difficult-to-discern keming with the default font on HN, and your unfamiliarity with the phrase, but it's actually "BUM A FAG" and not "BURN A FAG".
Incidentally, did you intentionally write keming instead of kerning?
Part of being an adult is understanding that generally, people are not 4chan trolls.
If someone who's not a crackpot says to you "can I burn a gay", even though that seems unambiguous, they've probably just misspoken or are joking.
In real life almost everyone seems to get this, online it's like everything is a strawman.
Well, for one, this only applies if you already know that person is not a crackpot. Secondly, burning someone would indeed be an extreme example, but not all (accidental) offenses are so obviously monstrous. If someone said "I hate fags" (intending to say they don't like cigarettes), this would be more likely to be someone's actual homophobic opinion, depending on culture.
It's all just a bit weird to me. If someone said to me "I hate fags", I'd either ask them to clarify, or change the topic. If they were persistent, then yeah, get rid. It's not a problem that requires instantaneous resolution.
I don't really want to take the discussion in this direction, but note that almost all of that "shift" comes from one side of the political spectrum.
There is no happy medium here, particularly since those seeking to offend and to wolfwhistle to signal their allegiance to other people who hate your race (for example) will try to couch their terms to hide within the shadows of plausible deniability unless they feel completely safe.
It's more a question of how many false positives and false negatives your culture is prepared to tolerate.
I certainly agree with your final conclusion: it depends on what society is willing to tolerate in terms of false positives and in terms of false negatives. And particularly, the balance between those two kinds of errors.
Not to say that there isn't anything wrong with how people have been increasingly incensed in the last few decades, but that's basically what it was like in the 90s and before. Someone was like, "whoa dude what you said wasn't cool" was met with "it was just a joke".
Your reasoning cuts both ways I think.
If you want to be a global publisher of your opinions, you’re going to receive a global response, from people across the political and social gamut — the larger your audience, and the more political/ideological your speech, the more substantial a response.
My dictionary’s definition of “Faustian” has this usage example:
> Modern celebrities enter into a Faustian pact with the general public.
What people seem to want is the ability to enforce a one-sided bargain with the public, under which their own toxicity, opinions, and any public affirmations should never be censored, as long as their opinions are the politically correct ones.
So, give hard core racists the benefit of the doubt when they wolf whistle each other?
Can't see any historical precedent for that going disastrously wrong.
>What people seem to want is ... their own toxicity, opinions, and any public affirmations should never be censored
Some people are hypocrites. News at 11.
Some people also want anti-han Chinese racism to be treated equivalently to anti-Uyghur racism because they choose to be blind to the inherent power differential.
I remember reading a story of the Hutus and the Tutsis. The genocide began with the radio - shows that very deliberate but cautious (& deniable) dehumanization of the other that ramped up over time. Nobody policed these deliberately aggravated ethnic tensions. By the time they were consistently obvious it was a little on the late side.
On the other hand, I've been falsely accused of being a racist before and it undeniably hurt my feelings a little bit.
And to shut up targets - "he/she does not mean that", "he/she is just insecure", "you know how he/she is". If you cant prove intent, you had no grounds to even talk about that that person is doing to you again and again.
E.g. people who object to a statement because it might possibly be offensive to another group, despite members of that other group being present and not being offended. Clearly they are doing this out of the same enjoyment of other people's discomfort that you describe.
That said, it's interesting this has come up during a discussion about how the intent behind what someone said is often either misunderstood or discarded.
You are assuming way too much good faith. For every one that does it in earnest there's probably 10 grifters just taking advantage of the situation for their own personal benefit:
The real reason seems to me like virtue-signalling. They don't care whether they're personally offended, as that would require burning social "points" by rocking the boat but potentially getting no outcome out of it (nobody else might agree). It's "safer" to just suck it up than sacrifice social points especially when you're on the lowest ranks.
They also don't know nor care whether the protected class they are worried about is actually going to be offended. Again, rocking the boat and burning social capital for little personal benefit.
This leaves the argument of earning social capital by virtue signaling among like-minded peers. Pick a common issue you already know is popular within your target "market" (racism, sexism, "diversity & inclusion", etc) and raise the problem. This doesn't risk any social capital, after all, people like you won't disagree (for the same reasons) and people that might be on the fence or would like to add nuance wouldn't want to risk their social capital in fear of being labelled a racist/etc. It also transcends the ranks - if anything, a higher-ranking person in the company would risk much more by speaking out than someone on your rank or lower.
This eventually results in a positive feedback loop where everything will become offensive given enough time and those who disagree either comply or get pushed out. That's how you get bullshit like GitHub's "main" branch which helps nobody while continuing to sell services to ICE which you could argue actually hurts minorities.
And therein lies the problem. A subset of people get a benefit from engaging in this behavior. So of course they do it at every opportunity.
You get what you incentivize.
Until the outcome is so frequently a net negative (however slight) in so many cases that the behavior is marginalized to a negligibly small number of people/settings the behavior will persist in enough volume to be worth caring about.
The problem is that a large part of the tech industry nowadays relies on "engagement" aka monetizing user attention for the benefit of advertisers, thus extremely vulnerable to public opinion (the advertisers themselves would engage in this behavior, thus companies that supply them must also conform to it). The entire industry is on thin ice already so it makes sense for everyone to play it safe even though it just reinforces the feedback loop even more.
This is why this behavior is rarely seen in other industries that make the bulk of their money on a tangible output, both because the output is less fungible/replaceable as well as those industries' employees typically having less time to waste engaging in bullshit.
I'm just going by conversations I've had with people in person and online, and assuming that this is broadly reflective of society overall. Don't get me wrong, I am quite sure that there are people who are cancelling people for sadistic pleasure and get a kick out of torturing someone and use hot-button issues to do so. However I cannot imagine this type of person outnumbering those with good-faith concerns at a 10:1 ratio.
Other times, you get the "well it's so very easy to make this small change, it's practically rude (or suspicious) not to" - a completely wrong sentiment for many reasons, but easy to make you look difficult for arguing about.
Free speech is not without cost, but it has been successfully utilised to everyone's benefit for hundreds of years. I can only surmise that those advocating for this have never read about or understood the importance of the Enlightenment. They don't understand the horror of living under a social regime where the powerful control language and effectively thought. They don't understand why the ability to risk offense is critical for democracy, science, and social progress. Eliminating free speech will harm the most vulnerable in society; not the most powerful.
American history itself had tons and tons of situation where people risked a lot for saying things - including death. All within the scope you talk about. The time span you talk about include runup to civil war, civil war, reconstruction - all of which were full of political violence.
This also when dueling was the thing. Where men of means were expected to shoot each other over words. If you did not, you was effectively over in that town. And the line between "duel worthy" offense and not could be very thin.
> Eliminating free speech will harm the most vulnerable in society; not the most powerful.
Moreover, we are not talking here about legal standard at all. Neither was legal standard EVER. You could legally seek to insult and offend people as much as you wanted. Nobody cared about intent. This is about when people say "you are jerk".
There is threat to free speech, especially with new proposed laws, but it has zero to do with intent vs effect.
You would think we would learn from these examples and make offensive speech more acceptable, rather than less.
I'm not sure how you conflated legality with social acceptance. My arguments center around the later.
Now do European history pre-enlightement.
It was way, way worse. Shit talk the church and the inquisition slows up. Complain about the cut of your grain that the mill takes and the local lord threatens you lest other peasants start complaining.
The acceptance of freedom of speech as a general principal has greatly increased the civility of civilization.
You are certainly correct - unfortunately the people burning the books are not the ones reading or writing them (even if they pretend to do so, with advanced credentials or padded resumes).
- all social media implemented simple word filters. A word is on YOUR list, the post or chat line doesn't make it to your page/chat. It'd be awesome if these lists were exportable, syncable, and tradeable online.
And I wish the AI revolution would benefit me personally instead of companies.
I would like an AI that censors on my behalf only if and when I want it to, and a personalized AI that I can train to flag accounts/people that I don't want to interact with based on language. I would also like some mechanism to indicate how much is getting caught in the AI-generated filters and allow me to stop or modify its behavior. Once the AI is doing mostly what I want, I would like to be able to share it. I don't know how that would work.
- Group owners should have similar mechanisms for all posts in a group.
Another example: faggots are a traditional meatball dish from parts of England and Wales-the term has nothing to do with anyone’s sexuality. And yet, American social media companies have repeatedly blocked British users for posting about the dish.
There is also a British pudding called “spotted dick” - in this context, “dick” is a dialectical term for “pudding”, no known connection with genitalia.
A friend of my wife recently had her Facebook account restricted for making a post about cooking jerk chicken for dinner. Facebook claimed she was using “abusive language”
[1]: https://news.sky.com/story/spotted-dick-renamed-spotted-rich...
Or try explaining panto to Americans. Or even Allo Allo, a show which would .. definitely not be made nowadays.
I've seen false positives referred to as the "Scunthorpe problem". Someone on the internet has a nice map of filthy-sounding real UK place names ...
There's a reason Essex University had to register sx.ac.uk as an alternative domain name.
For starters, you can't unless you actually include the context in the training set.
Then again, a lot of humans won't successfully manage that either...
It's worse than that. The context doesn't just have to be in the training set. It also has to be available on the other end, when you're using the model to make a determination. And it isn't.
Which is why even human censors can't get it right. The parties communicating can, and almost always do, have external shared state that The Decider doesn't. Which gives the same words different meaning.
Imagine hearing an inside joke you're on the outside of and then being asked to adjudicate whether it was offensive.
E.g. if judging tweets for example you'd presumably do a lot better if you evaluated the tweets in context of followers and in context of past tweets - both recent and the accounts history. E.g. an account that posts white supremacy tweets regularly is likely to mean something entirely different if RT'ing a BLM tweet with "that is fucking awesome" than what someone with #BLM in their profile is likely to mean.
Part of the problem is that they're throwing away a huge amount of the state they do in fact have.
But I absolutely agree it's in general an unsolvable issue - I pointed to a story from my childhood elsewhere showing how trivially people cause problems for "The Decider": A childhood friend being forced not to call his brother names and switching to using names of cheeses.
Is "Edam" an insult or a food preference? You can't know without context.
And it needs to be very local context, because humans very quickly pick up on when you just start using a word to mean something else, so it doesn't even need to be any shared external state about the word, just a shared understanding of where the receiver might expect an insult coupled with an unexpected response that will then easily get labelled an insult. That unexpected term might well in itself be positive if the receiver expects criticism. "Awesome" and "I love it!" are perfectly good insults when the other party has just told you something where the appropriate response would be negative, for example.
Must never use followers. Or anything else the account holder has no control over. Otherwise you'll end up with the same kind of SEO problems where competitors and foreign governments create fake accounts to follow or interact with disfavored ones and destroy their algorithmic reputation.
Guilt by association in general is malice. Someone who regularly interacts with white supremacists might be a white supremacist -- and maybe 90% of them are -- but it could also be someone criticizing, mocking or debunking them. And if it gets out that the algorithm is penalizing people for doing that, they'll stop. Which is very bad.
I specifically follow a few locals that I would often notice upset with the same politicians that I was, but for diametrically different reasons. Even though I disagree with them on most everything, I find value in having a few of those voices on my feed visibly attached to a consistent person. It helps me to resist seeing the "other side" as just an impersonal sea of voices.
100%. It's crazy how often context is missing from datasets! We wrote a separate blog post recently about this too: https://www.surgehq.ai/blog/why-context-aware-datasets-are-c...
Obviously also a technical term used all the time in its technical sense with no problem.
Context and culture is way more important than the actual words used when trying to determine the meaning of a statement.
"Obviously also a technical term used all the time in its technical sense with no problem."
Or sometimes not ... a blog I'm involved with recently had a WordPress server problem which caused commenters to receive a message like "Illegal operation: bad nonce". The readers are non technical and mostly British. Some had to be calmed down a bit!
Maybe it's an age thing? Everyone my age knows what nonce means but mostly because it was so amusingly satirized by Brass Eye in the late 90s at the height of a media induced paedophilia scare. Chris Morris got politicians and celebrities to say stupid stuff on camera by telling them it was for charity literally called Nonce Sense:
"Did you know that a child under the influence of a paedophile may smell like hammers? That's right - I'm talking Nonce Sense"
https://thesundae.net/2021/01/21/talking-nonce-sense-in-prai...
The program was a sensation at the time but probably only within a certain age group. I can't imagine that sort of satire would have appealed to older people.
There have been some unfortunate mistakes from companies not aware of this. https://www.dailymail.co.uk/news/article-9760601/Cryptocurre...
https://groceries.aldi.co.uk/en-GB/p-mr-brains-6-pork-faggot...
Before I noticed this comment, I made one of my own, making the same point - and it seems HN has marked it as dead for much the same reason: https://news.ycombinator.com/item?id=30069899
‘Tis ironic. (edit: thank you to whoever vouched my comment, it isn't dead any more)
This guy did: https://youtu.be/3-son3EJTrU
Humans are inventing double-entendres to create the ambiguity.
there are enough trigger happy people with less understanding than an AI, all eager to kick enough rages to wear down a saint…
Many business models based on user generated content wouldn't be possible if the businesses had to pay minimum wage to people for moderating that content. Using an AI model, no matter how broken, allows them to seem more concerned than if they were just relying on an old-school word filter without doing any actual due diligence.
This isn't documented anywhere, but I find it hilarious and dystopian that posts that contain the world "hate" in the description or in the comments get blocked more easily, or soft-banned (they don't show up in anyone's feed), even if the word is used in harmless contexts...
Content management done well is expensive, and requires humans with superior grokking skills. No place cares enough to try.
Sure "AI" can do a lot of impressive first-level pattern matching, and that can be the basis of many useful outputs.
But ANYTHING that requires actually understanding the context, whether it is existence of new obstacles for a 'self-driving' vehicle, understanding of actual meaning for a language model, or anything else, is a complete and utter failure.
Despite appearances, while some of it is genuinely useful, we're really no further than fancy parlor tricks. Crack even the next level of contextual understanding, and it will be an astonishing leap.
I've been following some drama in chat/user interaction in modern games. People are getting banned in Forza Horizon 5 for funny decals. One guy got an 8,000 year ban for a Kim Jong-un KFC paint job (https://i2-prod.dailystar.co.uk/incoming/article25686363.ece...).
EA is banning entire accounts - access to ALL games on an account - for swearing in Apex Legends. All of the major studios are in the process of severely restricting and even removing chat in their games.
The latest Battlefield 2042 game doesn't even have voice chat, and they removed the global scoreboard altogether because they didn't want poorly performing players to know how poorly they were performing.
The list goes on and on. To me, this is all a serious reduction in the enjoyment I have in games. I grew up playing Counter Strike, and s**-talking and competing was a big part of the experience. I just won't play games where I'm not allowed to interact with other players.
That being said: I specifically only buy entertainment products that do NOT require me to listen to some 14-year-old describing sexual activity with my female family members in great detail because his nonexistent skills make me keep winning. Different strokes for different blokes.
At some point games companies were bound to notice the ways voice chat limited their market and remove it.
Stop taking features away from the rest of us
Yes, of course everyone realizes that no filter actually works in practice. But they are not built to work, they are built to give the smallest impression of working. Tumblr doesn't need their site to be absolutely clear of nudity to attract investors, they only need to give the impression that it's absolutely clear of nudity and that they are trying to keep it so.
Perhaps this is okay when your datasets are high-quality and representative of the real world, but they're usually not. For example, many toxicity and hate speech datasets mistakenly flag texts like "this is fucking awesome!" as toxic, even though they're actually quite positive -- because NLP datasets are often labeled by non-fluent speakers who pattern match on profanity.
(So is 99% accuracy or 99% precision actually a good thing? Not if your test sets are inaccurate as well!)
Many of the new, massive scale language models use the Perspective API to measure their safety. But we've noticed a number of Perspective API mistakes on texts containing positive profanity, so this post was an attempt to explain the problem and quantify it.
For example, to pick a bit on the Google Emotions dataset again, it's difficult to label this message...
“Also Republicanism is a belief system. It’s taught and handed down like religion. Conservative talk radio is its evangelism.”
...unless you're familiar with US politics. Hence why it was labeled as APPROVAL by the non-US annotators, even though it's criticizing Republicans.
In the end you can be profane and very positive. You can also have strike language while being incredibly vile and negative.
My wife's cousin recently had a baby, and posted a photo on Facebook. My wife commented (something like) "She's so cute I'll have to kidnap her". Facebook locked her account for 24 hours for making a "violent threat". I am sure the mother wasn't threatened by the remark in the slightest. It just added to my wife's anger at Facebook for repeatedly giving her "warnings" over trivial or innocent things.
Not online, in person, but one of the teachers at our son's school sometimes "threatens" to "steal him" from us – her remarks don't worry us, because we know she would never actually do that, it is just a colloquial way of expressing affection.
This is due to years of old media digging through Facebook posts to write stories like "Child kidnapped after threat on Facebook and Facebook DID NOTHING". This led to more and more calls for FB to "do something" to "keep people safe online". So they started to run all posts through the AI prescreener and well, these are the results.
But hey, they get to at least say that they take down X billions of posts before any other human sees it.
I don’t think the mother would be as accepting then, so it also matters who posts the comment to who
Beyond this, by hiding the comment/post, the mother is now unaware of the threat and Facebook didn't contact local law enforcement in fact worse than doing nothing at all.
Which is to say: Often it's not reasonable to act without first knowing if the recipient actually took offence or felt threatened, before even considering anything else, because you can't know. A model aimed at judging utterances needs to also take into account "how close is this group? could this be in-group language?" and it needs sufficient context to judge that. Even then it'll have huge potential for error.
Incidentally, a lot of this wouldn't be nearly as much of an issue if we could trust that these kinds of tools would be used appropriately. E.g. instead of banning things, even just asking a "are you sure this will be taken the right way?" or similar is a whole lot more benign. The dating app Bumble uses a filter that checks for possibly explicit pictures, for example, but instead of blocking them it shows you a short message asking if you're sure the recipient wants to receive what you're sending, and lets you choose. Depending on use there might well be several levels of severity in messaging justified, from just a mild "are you sure?" marker to forcing you to explicitly acknowledge that you're on notice that you're responsible for the content.
If a male relative or friend had made the same remark, would my wife’s cousin have felt threatened? My wife tells me she would not have been bothered by such a remark, in that context, from one of her male relatives or friends.
And, if someone was really bothered by it, they could report it. But, it doesn’t seem like that is what has happened here, it appears to just be the work of some automated system. I’m not saying that if the mother reports the remark Facebook shouldn’t take the report seriously - cultures differ after all, in some cultures those kind of remarks are normal in that context, in others they would seem bizarre and offensive - but it shouldn’t just assume the remark is threatening because it contains the word “kidnap”
I was more interested in hypothetical "can we conclude it is ok from friend status on Facebook without actually knowing involve people" question.
I've noticed a lot of humans can't detect sarcasm. And that isn't correlated with traditional smarts either. For AI it's a hopeless task.
> You can also have strike language while being incredibly vile and negative.
I'm reminds of a group of people I know. They love love using genteel language to throw vile insults at each other. Using profane language is a automatic foul.
Now consider a large group of people upset about censorship and motivated to adjust their language to circumvent it...
The specific words don't matter, as any reaction would just shift the insult to another term. Ban cheeses? Start saying "I love you" ("mom, he told me he loves me! Tell him to stop it!"...). Or call him awesome. If anything, I think the brother missed a trick by not going straight to compliments - few things can be more cutting.
Given the right context it'd be clear to the brother what the intent was no matter the term, but increasingly impossible for the parents to distinguish "legitimate" communications from the insults. For that matter, no words are necessary to achieve that once both sides know. Just a look is enough.
We don't know how to fix this social issue with human censors, so it's much worse than just the state of the tech - it's an issue that's basically unsolvable without addressing why the speaker is intending to insult (if they are) and why the receiver takes offence (if they are).
The best you can do is address a small proportion of the most blatant cases (and that might well be ok and enough if you can do it with few enough false positives, but getting the false positives to an acceptable level is in itself a momentous challenge).
There of course isn’t a perfect solution and likely never will be.
The problem is that unless you observe the behaviour yourself as it happens, with substantial context, the idea that you can reliably know whether the behaviour was unacceptable is fundamentally flawed.
This is exactly what the choice of innocuous terms is for; to create plausible deniability. E.g. the word "camembert" thrown at someone out of the blue might be obvious because it stands out. But that's not the story you're going to get from the person who has used it as an insult. You're going to get a story about how they just talked about what they like to eat, or something like that.
You might see through that, but what quickly happens - what indeed happened with the brothers in my example - was that the person making the insults quickly calibrates to know exactly what results in a reaction and what creates enough plausible deniability that the person deciding if this is sanctionable is unable to determine with any reliability whether what they're hearing genuinely is an insult or the "target" being oversensitive and misinterpreting things. I was bullied at school at times, and saw this first hand as a target myself - the bullies quickly learned which ways they could word themselves to make the insults clear to me while creating enough plausible deniability even when there was no conflicting claims about what exactly they'd said.
Point being being that humans even with context and in the position to interrogate the people involved can't solve this with any kind of reliability. Then the notion that an AI at current state of the art without access to context or the ability to interrogate those involved doesn't stand a chance.
Yes but we're not trying to be perfect, we have to live and deal with this grey area.
> Point being being that humans even with context and in the position to interrogate the people involved can't solve this with any kind of reliability.
I disagree here, I think we can quite reliably do this a great amount of the time. There are absolutely bad actors who will get away with things in the short term but because context builds up over time it never lasts. And once you are aware they're trying to be subversive its much easier to spot earlier. It's not an attack with legs.
> Then the notion that an AI at current state of the art without access to context or the ability to interrogate those involved doesn't stand a chance.
I absolutely agree here.
And I'm sorry the people in charge were rubbish at protecting you from bullying.
"I'm going to sleep with your father and then give him a son he actually loves."
Rated not toxic. All this API is going to do is promote a renaissance of polite burns.
> "this is fucking awesome!" as toxic, even though they're actually quite positive
And this touches on the question of what we want to do. It might not be a negative sentiment, but it might still be offensive. On the other hand, there are no offensive words, and only "positive" sentences in many sarcastic utterances.
And there are a lot of sentences that can be perceived as insensitive, that are not meant as such. GCP Grey has a video on the words "Indian" vs "Native American", and apparently it's very complicated which word you can or should use.
https://www.cgpgrey.com/blog/indian-or-native-american-reser...
No idea if there is a different metric that somehow (again) takes the number of sample into account.
There are some differences in German, too. For example "wixen/wichsen" is an old word that means to wipe/shine your shoes and is still in active use in this sense in Switzerland as well as Austria, however it lost its appeal in Germany, because it is now primarily with a different meaning. The Wix company took this different understanding of its brand name to an ad: https://www.youtube.com/watch?v=IddnMutPgTI
Since we have an IT background here, same goes for "Mongo", like in MongoDB. Mongo is considered making fun of handicapped people in Germany.
Former Fraport AG changed its brand name because it was abbreviated FAG - Flughafen AG and found it difficult to expand business with that brand name.
No bad actors, if you ask me, only different context. List could go on and on...
People with Down's syndrome used to be called "mongoloids" because they were thought to have Asian-looking eyes (epicanthic folds), so calling someone "mongo" or "mong" is basically an extra racist way of calling someone mentally retarded.
It had to go? Certainly not gone from git itself, Github or most other Git hosting services.
If someone creates a bug report on your project about virtue-signalling word changes and you give it credibility, it doesn't mean the original word went away from the lexicon.
"Hold on, my dock is nearly up" "Hold on, my ** is nearly up"
Censorship can actually make things MORE offensive. (other censored words include "but" and "come")
Still, all in all, automatically replacing a (supposed) swear/offensive word with some asterisks makes less overall damage than banning someone's account (temporarily or for longer periods or forever).
It is simply intimidating when you have to think twice or thrice before saying something for fear of being automatically banned (without possibility of explanation/recourse).
It's also likely that there's a significant lack of context to the data points, not just because the posts are divorced from their parent content, but because of a culture and language divide between the labeler and the author of the data they're labeling, as well.
Hopefully people also don't need to be at the level of a professional linguist to label messages like "this is fucking awesome" correctly!
And great point on context. For example, the GoEmotions dataset didn't present labelers with the actual post or subreddit the message came from -- just the text itself. That makes it really difficult to label something like "his traps hide the fucking sun"! But once you see the comment in its original context https://www.reddit.com/r/nattyorjuice/comments/aee3wx/olympi..., and know that it's in the /r/nattyorjuice bodybuilding subreddit, it's much easier to realize that this is talking about someone's large muscles.
That's survivorship bias.
It depends on what the purpose of content moderation is.
Good if you want to accurately identify abusive behavior and protect users from harm? No.
Good enough if you want to find the most blatant examples of name-calling and insults to appease regulators and trigger-happy lawyers by appearing to use "state of the art technology"? Sure.
I had a housemate once who was an American linguistics grad and spent a year applying semantic labels for Google. I know he worked on a team, although I don't know what they did since I didn't work at a Google at the time.
Although -- and I say this having done a lot of my graduate coursework in linguistics -- I don't think having a linguistics background is particularly needed (unless you're doing specialized annotation, like creating syntax trees or tagging phonemes in Praat), outside of you being more likely to enjoy thinking about the nuances of language.
You can see that my model is basically just filtering for profanity as an indicator of "strong emotion", which makes sense. But it's interesting that postive profanity seems to be such a thorny problem, at least for Perspective.
[1] https://twitter.com/celesteperez___/status/13508599618452070...
In general I've found the results to be much better for subjects that a lot of people are currently tweeting about.
I recently started playing counter strike source again online, just for 10 minutes of fun at first (to see if it would still tick with me). I randomly picked up a server and the ambiance was cheerful and nice. I noticed the rules said "no profanity, have fun" and indeed people were mostly polite.
I tried another server at random a bit later and there was more insults, along with a lot of taunting.
I switched back to the first server and have been regularly playing an hour or two every three days and there is a difference with other servers. Some random people coming and throwing insults, even mild ones like "fuck you" or "you son of a bitch awp" get insta ban and it makes the whole session a much better experience. Maybe it's a safe place but playing with polite people is more enjoyable to me now than playing with insult gatlings.
Language is political. There are many meanings to words, depending on context but I do think it's not innocent to swear in front of people or to use swear words to look cool. These are still swear words and insults and their first original use is to provoke or taunt or display aggression. Even if it's only used for "this is album is the shit !", it's still a (childish) provocation. Reminds me of the brogrammer fad.
FWIW: I get regularly owned on this server and I am at the bottom of the ranks but it's still more fun and enjoyable than other servers I tried where I can reach the top but... it's not a nice place. I think online servers are like bars.
Side-note: I was pleasantly surprised to see that "gg" is still thrown around after rounds :). It's way better than "git gud" that came later and that I find horribly toxic.
I'm just saying, I'm not surprised that a moderated server is more, uh, moderated, compared to an unmoderated server. I'm not sure the individual word choices are the determining factor compared to having someone paying attention and removing people with undesired behavior.
Hard to tell because of the different language. They wouldn't engage someone saying something close to that but people wouldn't express their admiration or congratulations with swear words anyway (because it's not the style of the house ?). Of course they'd be using the community's lingo though to express sick moves but these aren't insulting/taunting/agressive to begin with.
> I'm not sure the individual word choices are the determining factor compared to having someone paying attention and removing people with undesired behavior.
I think too. For the comparison I should try a CS:GO or valorant or pubg or fortnite next week-end to see the difference.
You say that like having a place be safe is inherently negative. But your post is a textbook example of why people want safe places!
An immediate solution is to apply multitask methods to your target dataset and include the one proposed in OP. It's always good to have more resources like this, even though SurgeHQ overstates the size of their resources by large margin in their copy. The 1000 post instances of their dataset is far from "the largest": I have several aggression, toxicity and bullying human-annotated datasets right here with over 100k instances.
yep. The cultural imperialism is an open secret. But I've completely lost any ability to tell the difference between people pretending not to know, and those actually not noticing.
Instead of putting money into gathering quality data, tech giants instead chose to use cheap labour (here: India) with insufficient skills. You get what you pay for.
In a similar vein, there are popular "AI Mental Health" apps I've gotten to straight up instruct me to end my own life with some trivial conversation.
EDIT: Here's one, though I don't think it's ML. https://text2data.com/Demo
> It would be really nice if you'd end your own life :) Everyone would be happy.
> This document is: positive (+0.62)
For https://monkeylearn.com/sentiment-analysis-online/:
> Positive 84.1%
For http://text-processing.com/demo/sentiment/:
> Pos 0.7, Neg 0.3
For https://aidemos.microsoft.com/text-analytics:
> 100% positive
For https://komprehend.io/sentiment-analysis
> Positive
Delegating content control to an AI (that doesn't qualify for anything intelligent) is not a working solution.
> as a first-pass filter, leaving final judgments to human decision makers — marking all profanity as toxic can make perfect sense
You would need humans to look at profanity constantly.
> Our mission involves creating a safer Internet, but we don’t want to miss out on our favorite content because of AI flaws in the meantime.
There is a limited AIs that do create content, but a profanity filter always does the exact opposite.
Is it the lofty “make people be nice to each other”?
Well, the paradox is that being nice is possible with the strongest choice of words, while being very harsh can sound most fluffy bunnies on the surface. In fact, in human relationships there are degrees of mutual familiarity where being exceedingly polite and not “insulting” your counterparty would be perceived as negative—where insults are not taken at face value, but rather as signifiers of friendliness (there’s a line, of course).
Shall we unwind the perceived goal differently?
Of course, the platform’s actual customers are the advertisers (we are talking about a hypothetical platform, but where is it really different?), and by being free to the user it participates in a very limited oligopoly of big social so no, it doesn’t really care about what anyone really meant or intended, and it definitely isn’t going to hire real humans who’d make an effort at grasping the context of the conversation.
The real objective is for the platform to not have problems with law enforcement when one user complains about another user for being naughty, discriminatory, threatening, etc.—and, of course, we shouldn’t expect anything more from toxicity detectors geared towards that goal. As long as it’s not egregious enough that users leave en masse to a competitor (which can hardly exist, no honest business could reasonably compete with “free”) the platform wouldn’t care since users only matter to advertising revenue as cattle in numbers.
we need to get rid of black people
we need to get rid of black people poverty
we need to get rid of black people below poverty level
we need to get rid of black people hurdles keeping them below poverty level
the gist is that without unravelling a sentence full context, a lot of verbage can refer to a lot of different action.
focusing on profanity is the low anging fruit, so to say.
The same words can change context depending on what kind of building you're in or what time of day it is.
Even common phrases like "I'm running over" could mean I'm coming to see you or I'm taking longer than I expected; completely unrelated things.
Flow analysis may give you something, but content analysis is basically impossible unless there's also context analysis.
But since antagonizing trolls aren't usually using profanity or necessarily a different vocabulary of words but they are engaging in patterns, that kind of engagement style analysis may be all that's actually needed
It's very difficult when you blur the lines of code and ethics, as real world ethical judgements aren't necessarily consistent or well defined in a way which is easily translatable, even by a large ML model. Jigsaw is a great example of this -- right across the aisle (pre-pandemic) from Perspective is a team fighting internet censorship. Obviously Perspective's "censorship" is different in quality from the Great Firewall, but it shows the hairiness of the problems.
All this is to say the people at Jigsaw are some of the most brilliant people I've ever met and I'm glad they're out there working on difficult problems.
Given how eager Google itself is to automate all your interaction with them (and how hard it is to get access to anything resembling actual human judgment), did they really think it wouldn't be used that way?
Or did they just not care?
Also, I would have hoped the project's close ties to the US State Department would have worried some of you brilliant people a little more.
We deal with these hairy problems a lot too. Even before ML proper, getting the definitions right is very tricky. Should a comment that's polite and positive on its own, but supportive of a toxic parent post ("I love Nazis!" -> "I agree!"), be treated as toxic? Is "toxic counterspeech" equally toxic? What about a comedian making fun of an actor's nose, and does it depend on whether the joke is to that actor directly vs. merely referencing them in the third person? etc.
For these reasons, I actually personally like the fact that the Jigsaw annotation guidelines are very high level (as opposed to long and prescriptive) -- it lets the data capture the spectrum of "human preferences" on its own (at least, it does when you can trust that the annotators are able and trying to do a good job).
We need to come up with UI patterns and flows that reflect this otherwise ML solns will continue to disappoint.
I'm not convinced this solves anything or even helps - UI patterns can easily be abused by bots and just shift the problem into a different direction.
We need to develop a way to work with ML st most queries can be answered by ML but there is an exception pattern that defaults to a manual, prescriptive process, or whatever.
Also we need APIs to allow detection of that state.
- unpleasant discourse on online platforms is blamed on the platforms
- the platforms can't moderate this manually (nor objectively), so they look for a tool
- a tool can't possibly do anything useful, but it satisfies the "something must be done" media demand
- the tool will hurt conversation, and people, and potentially eventually threaten the platforms it runs on
- but those problems feel smaller than the demand that "something must be done"
Regularly now I find social media sites telling me "do you want to review this before you post it? You're bullying or being offensive". No cunt, I'm good.
Different words and phrasing have different impacts across cultures. Unfortunately Instagram gets to decide what my entire culture is allowed to say online. That's fucked.
I'm guessing it's only a matter of time until this comes to toxicity models, so they can look up who said it, and to who it was said.
For fucks sake, I saw a video of a crowd cheering after a guy imitated one of Hitler's speeches. There wasn't a single "shit" or "fuck" in it but it contained "truth doesn't matter, only victory". That's offensive as fuck but the people didn't see anything wrong with it.
Only in the context of knowing it came from Hitler. Without that information it isn't more than a strong opinion - because it isn't explicitly linked to murder/genocide.
Toxicity is poorly defined because it's an in-group euphemism for a kind of gendered disagreeableness, where its opposite or positive case is passive and agreeable, even passive aggressive. If there is such a thing as masculine aggression, there is also feminine aggression, and a lot of what we talk about as toxicity is really a criticism of masculine aggression using the lens or perspective of feminine aggression tools. I'd propose that when we say something is "toxic," we're talking about something that violates feminine norms around in-group alignment, security, reputation, reflection, impressions, "not a good look," etc. These are all things that require an imagined third party observer to potentially interpret and be offended by them, and are not codified by rules. Encoding this into an ML model is a lost cause, because you would need to reflect them through an AI that ran on pure neurotic animus to get a sense of whether something was toxic, or "not a good look." It's like assigning a sentiment score to someone saying, "Nice hair."
The example in the article of "Fuck dude, nurses are the shit," is ranked as 98%+ "toxic," because it has two frickative swear words associated with masculine aggression traits (disagreeableness, provocation, profanity, rebellion, dissonance, loudness, etc.) and easier to write rules for - except those rules would also need to incorporate whether the phrase was an expression of aggression, or using the opposite to be wry, ironic, or in the case of the example, to express awe.
I don't think we understand enough about psychology and people to really create effective moderation models with ML, and ML will necessarily create a kind of mean reversion in the discourse they monitor, which means all conversation subject to it will tend toward neutralization, which is essentially death. (Maybe we should fork an ML project for linking language with Jungian archetypes?)
We can keep trying, and applying it as a fast search scheme for prioritizing outliers to human moderators, but pleasing an ML model is a recipe for intellectual sterility. I'd even argue the inflection point in the growth of social platforms is when moderation creates this kind of mean reversion, and you are left with the bland platitudes of blue checkmark types, being cheered on as you "grind," and boomer memes that aren't as funny as Family Circus comics. It's just death.
Anyone selling AI-based sentiment analysis is a grifter.
If you get normal not-always-online not-gen-Z people to evaluate these messages and label them as Good or Bad then you will get results like this. If I got any member of my family over the age of 30 to evaluate these messages, they'd label them all as offensive.
How sure are you on that? I'm 31 and there's definitely no one in my peer group who would find those examples offensive. Hell, only 1/2 parents might be offended by that language, and my mom has been working around young people enough for the past decade that she might have softened up since I lived with her too.
It's not impolite, it's just very informal.
I, on the other hand, am a fucking amazing shit OG bitch.
But yeah this stuff is absolutely dangerous, opinions like these are destroying culture. It won't be long until the threads these automatic tools are moderating, WONT BE WORTH HAVING.
'Nuff said.
There are contexts that quote would not be offensive to anyone (more casual settings with a group that trusts each other), but also contexts where everyone would find it offensive, no matter their age or generation.
Agreed. As an example, your usage of "black" and "white" in this manner could be offensive to someone in some context.
Not a single person I know would take offense to this.
My South African wife's family, however, is a completely different story. Her mother's side (English) would be horrified by this, whereas her dad's side (Afrikaner) wouldn't be upset in the least.
Point is, it's geography and upbringing far more than age that predicts puritanism. If anything, the young people from your area are regressing to the global mean.
Not everything is the Disney Channel and not everybody shares your views on "morality"
I actually think that too saccharine language is trying to hide something, so immoral.