Minority voices ‘filtered’ out of Google Natural Language Processing models
unite.ai
unite.ai
> has been extensively ‘filtered’ to remove black and Hispanic authors, as well as material related to gay and lesbian identities
But after it is explained that this is an unintended second order effect from trying to remove offensive content from the corpus. I'm not trying to justify the outcome (it's an issue regardless of intent) I just didn't think there was any need to strongly imply it was intended.
How odd that "offensive" and "minority" turned out to overlap so much. How could that have possibly happened?
Running over a pedestrian might be a second order effect of driving a car, except that driving a car includes applying the brakes to avoid it.
Like, it's not like this filtering was done for the purpose of silencing anyone — Google (among others) really learned the hard way to not feed smut to ML models, as it _will_ get regurgitated, always as a possible PR disaster in the making:
https://www.buzzfeed.com/fionarutherford/heres-why-some-peop...
https://www.huffpost.com/entry/microsoft-tay-racist-tweets_n...
Still interesting research, but playing the devil's advocate, I think I can see why this part of the corpus was more extensively filtered away.
Sad issue, but I don't think we'd have a sane way out without massive, manual whitelisting.
thats not whats happening in ML. whats happening is "find the easiest amount of information to train a model"
https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and...
The list does seem a bit... umm... oddly specific in places, probably due to its history as being first compiled for a photo sharing site.
The authors also observed that the text of many patents are initially obtained via imperfect examples of Optical Character Recognition (OCR), with their accompanying errors in English possibly passed through to the C4 data with little or no annotation that would distinguish it from acceptable English.
While patent offices do release PDFs of all their patent docs (and it isn't just patents; it's all the back-and-forth between the examiners & the applicant, too), a huge percentage are images of paper documents. You can always tell just by double-clicking on a word -- if the word doesn't highlight, it's an image.
OCR output from these things is generally terrible. There is no way it should ever be input to any ML model.
Very disheartening, but I'm not surprised.
Same old story, history repeats. We are just cementing the abstract into code.
A word so lovely that we really really don't want to risk putting it in our report...
For most of the 90's saying black was "wrong" in favor of African American. Now we've gone full circle, and the same people who made a sour-face at black 25 years ago are capitalizing it.
That's probably the most benign, easy case. Of course ML can't keep up with loaded-terms and slurs; most people can hardly keep track of it all.
I can't tell if that is sarcasm or not. Is the 'lovely word' disclosed somewhere else?
Yes, and determining if a text is offensive is a very socially and culturally determined. Who determines the offensiveness of the context, and how? To shift gears a bit, how many rap song lyrics were filtered out from the corpus, for example? Does use of the n-word make an entire text de facto offensive? What about references to the common name of the moth Lymantria dispar dispar? What about the football team from Washington, D.C.?
A lot of this just feels like a repeat of the Net Nanny internet filter days, when keywords were idiotically filtered out.
You've gotta be living in a bubble to think this is something only 13 year-old boys say on the internet in the privacy of their bedroom.
Are they worried that a scientist is going to be exposed to a bad word?
Are they imposing some cultural brand of morality upon the AI?