A.I. Doesn't Get Black Twitter
inverse.com
inverse.com
If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices.
When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias.
Should we attempt to detect prejudices in machine learning models and actively counter them, creating a self-reinforcing virtuous cycle of balance and tolerance?
----
As an aside, for an article about language, this one has a lot of typos. Doesn't anyone proofread before posting?
Isn't the ideal system to not assume anything about the content it's analyzing, but to base the model on behavioral patterns instead? In the case of Google ranking, either a site has traffic/low bounce/backlinks/social cred or it doesn't--why is Google attempting to 'read' and analyze the content itself? (Other than to serve ads of course)
Because Google no longer just searches for literal strings of text, ranked by links.
They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results.
For more on this, read about Hinton's work on "thought vectors" [1].
So if Google can't determine the meaning of African American language as well as it can standard English, then that content is likely to be less visible in search results.
[1] https://www.theguardian.com/science/2015/may/21/google-a-ste...
Bing seems to do the same thing. I really wish there were a grep for the web because both engines get it wrong even when I try to exclude results.
It's basically grep for the web, since it uses the actual query instead of the meaning of it.
As an aside - this is a slippery slope conversation. Some people believe language requires rules. Other, more special people, think language is fluid and anarchist. The latter of those two do not make search engines.
My guess is that it's unlikely to provide sources that are in one of these dialects, as the link density will be low.
> Some aspects of communication are likely to prove more challenging, Hinton predicted. “Irony is going to be hard to get,” he said. “You have to be master of the literal first. But then, Americans don’t get irony either. Computers are going to reach the level of Americans before Brits.”
A further problem, though, is that search engines try to understand the query and the content searched at a deeper level than mere keyword matches. For example, they may try to pick up synonyms (if I search for “movies in <local area>”, perhaps results for “films” might be helpful too) and associations.
So that when you search for "cheesecake recipes", you don't get a well-sourced, high-traffic article from a respectable website describing in detail how the new business initiative by Cheesecake Factory is a recipe for disaster.
* Disclaimer: I pulled that example out of thin air. There's actually no such article.
Which is precisely why social media (e.g. your newsfeed) is such an awful source for news, especially when you consider how low the standards are on digital publishing these days...
No.
Depends what you mean by "the corpus". Let's take the example of loan applications. Maybe the training data contains credit scores, biographical information, and personal essays about the applicants from loan officers.
Let's say loan officers tend to use more negative language when talking about black people.
Any decent machine learning system given racial data will not adopt those biases. In fact, it will counteract them. It will determine that e.g. the semantic content of the essays is an indicator of loan suitability, and it will also notice if it has to e.g. add something to the essay score for black people to make the best decisions.
As long as the output data is objective, competent ML systems will account for any bias in the input data.
> Should we attempt to detect prejudices in machine learning models
We don't have to. If the output examples we are training on are true, the ML algorithm won't adopt any incorrect biases.
This is literally the point of machine learning. You want to make the best possible decision. If your algorithm has incorrect human biases, you're losing money. ML naturally accounts for this stuff.
And the bias people are probably talking about is bias in the data, not bias in the "input", as you put it.
See http://arxiv.org/abs/1607.06520 for a counterexample to your point.
Do you know what I mean? There's truth as in what the data says, and there's truth as in what's underlying before social processes " corrupt " the truth. This paper examines ways to get at the latent truth.
Do you think that's a useful endeavor? It's dismissive to say it has anything to do with what the "author likes".
Ah, so you're defining "truth" as "whatever I imagine the world would be like if society worked the way I wished it did". There's another word for that; "fantasy".
This is rarely the case when working wtih real data, and thus inspecting whether our models are biased against protected classes is probably one of the most important things an ML practitioner should do.
There is no reason an ML algorithm would be biased against a "protected class". It doesn't know what those are. It's possible that the algorithm will uncover a truth that you don't like, like that there as differences in risk profiles across race or gender, but that doesn't mean the algorithm is incorrectly biased. It just means that reality is at odds with how you might want it to be.
Which data are you referring to? In most cases, the training data isn't human-generated, and if it is, we usually want to match human behavior as close as possible.
Read anything written by Solon Barocas: http://solon.barocas.org/
The issue of prejudice stems from refusing to treat AAVE as a valid input to an AI, mirroring the racism in American culture that similarly devalues and excludes it. Output can conform to some approximation of "standard" form without being problematic, following the pattern of interaction between standard and dialect speakers. As a speaker of the standard, it's gauche and derisive to imitate or refuse or comprehend a dialect speaker, not to respond in your own form.
What is really the difference between spelling a word wrongly and using a non-dictionary word to express something. We except one as "slang" but will criticize the other because they don't comply with established rules. Isn't that what slang is? A break from established rules.
Anyways, just a thought, because I agree with you; if you publish - proofread. I just find it interesting to throw a Socrates argument at myself from time to time.
Btw and somewhat of topic.. this was the Pinker interview that was most interesting:
Programming a computer to perform a task is a powerful way to reveal hidden assumptions. What the problem of machine learning on human data reveals, is there is no objectivity. Our most 'objective' models will simply learn and then reinforce biases and prejudices and cause harm to people. We can't build objective, neutral, value-free models when it comes to human behavior, because humans change their behavior in response to the models. When we reject stereotypes we are imposing a set of values, just as much as when we embrace them. Machine learning forces us to confront the fact that this is exactly what we are doing.
There's a kind of philosophical crisis going on here. We need an entirely new way to think about the limits of objectivity in human sciences, and how to create ethical models in the presence of feedback between science and society. The language of objective physical science doesn't work.
https://mobile.twitter.com/MarkHamiIl/status/778141129564905...
I thought most people would be able to get written Scots dialect, but I did struggle with Walter Scott's Rob Roy where one of the characters speaks it phonetically.
There's also a guy doing a newspaper column in Scots in The National, causing quite a bit of controversy.
Spoken language is another matter. I've had to interpret between an Ulsterman and an Afrikaaner both of whom were nominally speaking English as a first language.
Are we headed towards a time when deeply trained networks seem to work in an acceptably large number of cases, but how and why are unknown and susceptible to fatal flaws that rarely but catastrophically reveal themselves? I suspect that is the case with the current path of machine learning but hope I am wrong. It just seems to me that the inevitable results of certain efforts in machine learning are magic boxes that "work" but no one understands why or how, which smacks of the days of medicine prior to an understanding of the germ theory of disease where certain efforts at sanitation due to the belief in "ill humors" did improve health but were based on utterly false underlying theories, and those successes tended to reinforce other false solutions that did not improve health but in fact was harmful.
This blog post summarizes the paper "Machine Learning: The High-Interest Credit Card of Technical Debt": https://blog.acolyer.org/2016/02/29/machine-learning-the-hig...
[1] http://www.theatlantic.com/technology/archive/2015/04/the-tr...
A large number of ImageNet classes are dogs and birds, so most of the representational power in CNNs trained on this set are devoted to distinguishing between similar-looking dog breeds and bird species.
Is this long term really a problem though? Regression to the mean: No one cares about the specific slang of the 60s anymore. And the hip slang of today will be out of fashion in 15 years.
In a similar vain the extreme T9 keyboard text message abbreviations of the youth in the 2000s did go out of fashion because of automatic spell checking and speech recognition of modern smartphones.
How long term? What's "a problem"? We know with certainty that given the passage of enough time, the speech of "these kids today" will be completely unintelligible to you.
Definitely. Things are clearer when you write formally and strictly adhere to grammatical rules. When I'm speaking casually I omit words and (ab)use punctuation much differently, and while that is perfectly acceptable to do, it is more complex and ambiguous to parse.
As far as I'm aware the same pattern exists in a lot of languages and dialects, and it's possible to define formal AAVE too.
>I don't think there is anything "formal" about Standard English.
"Standard" English as she is spoke? No. But a3n was specifically talking about the formal version often called "Standard written English". It is very prescriptive and less flexible, which makes it easier to parse.
> Can you give an example of what you're talking about? Which grammatical rules do you have in mind?
Let's start with "every rule involving punctuation".
Another ambiguity is parse ambiguities like " I saw the girl with the telescope." What does "formal" English say about that?
Absolutely. It follows a strict set of formal rules and you can readily find sources to learn those rules. Heck, copyeditors regularly debate the specific nuances of how to consistently write Standard English.
Other dialects (and spoken language) are far less formalized. Where is the AP Stylebook or Strunk & White for AAVE?
I would guess the same problem occurs with other dialects which are sufficiently different from the standard.
Choose a student of any colour. When they need a text corpus for research, do you think they'll reach for a different, known, large text collection that matches their background?
[1] https://en.wikipedia.org/wiki/Oakland_Ebonics_controversy
And "more diverse datasets" isn't enough. I get that you need to train an AI on a set of old data, but any truly good NLP would be able to adapt to new terms and phrases as they're invented. Not saying it's an easy problem to solve, but many people I've spoken to don't seem to even be aware that language recognition is necessarily a moving target.
It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million people read (subscribe to?) it. That does not mean that only 8 million people in the US use so-called standard English. There are also a lot of people who don't subscribe to that particular publication but still speak or write in "standard" English and that group is much larger than the group of WSJ readers.
It IS "the" language of that two million. Just not ONLY those two million.
WSJ dataset is not the only such dataset. NLP is not very good with standard English yet and usually doesn't generalize from topic to topic. Dialects and other languages - especially those without formal rules - will come when we can deal with standard English.
The point is for rules to be formal then the must be formalized somehow and codified. This could be online or in a book, but the point is that there must be some clear delineation between when the rules are followed and when they are broken. Otherwise what's the meaning of "formal" rules?
> If so, aren't these kind of antithetical to AAVE on the face of it?
Yes. My argument is that AAVE, almost by definition, is an informal dialect without formalized rules.
This is getting rather off topic, but I think it relates to the original point of why NLP might start with Standard English even if you are not biased. A large corpus of Standard English text (such as from the WSJ) will generally be very internally consistent precisely because it follows a set of formal rules codified into a style guide. As there is no such equivalent for AAVE, even gathering a large and internally consistent corpus of AAVE text seems prohibitively difficult. That being said, I do hope researchers are working on gathering text from Twitter to build up new training sets.
So every form of English is an "informal dialect" then? Because this ain't French with the Académie publishing strict rules for use of the language. Do you also say that the languages of remote Amazon tribes aren't "real languages" because they don't have a formal government body publishing written rules?
Or do you just want to bash on AAVE and are grasping at straws for reasons why?
Yes, the majority of spoken English does not follow the rules of Standard English. Pretending that such rules don't exist is willful ignorance though: the WSJ obviously write a more formalized version of English than teenagers do in text messages.
> Do you also say that the languages of remote Amazon tribes aren't "real languages" because they don't have a formal government body publishing written rules?
Nowhere did I say that AAVE is "not a real language" because it's less formalized than Standard English. Prior to spelling reforms, English itself was extremely inconsistent and informal, but I certainly don't pretend that it wasn't a language.
> Or do you just want to bash on AAVE and are grasping at straws for reasons why?
I'm not trying to bash AAVE. In fact, I'd even posit that the reason AAVE isn't more codified is perhaps because of racial bias which treated it simply as "incorrect" English instead of a separate dialect worthy of formalization. Pretending that all languages are equally formalized is simply willful ignorance though.
> NLP is not very good with standard English yet and usually doesn't generalize from topic to topic. Dialects and other languages - especially those without formal rules - will come when we can deal with standard English.
The rules you're talking about, that get printed in books and studied, are not linguistic rules. Crucially, this means they are not widely observed in printed standard English, which in turn means they can't be relevant to training a language model to understand printed standard English.
The "formality" you seem to want to talk about has no place in this discussion. It is not relevant to any language. gordonguthrie is correct to point out that the assumption lqdc13 is trying to make is false. You are wrong to contradict him using a meaning of "formal rules" that you brought to the conversation yourself. It had a meaning -- a completely unrelated meaning -- before you showed up.
I agree that they're not widely observed in written English, but they are consistently observed in the WSJ, which was the origin of this entire debate.
As lqdc13 pointed out, NLP still isn't even consistently good at understanding standard English. One could reasonably posit that that's due to the inherent ambiguity and inconsistency of most writing and that focusing on a narrower, standardized document corpus (the WSJ) you could get better initial results. What, exactly, is controversial about that? Do you really think that the language of the WSJ is no more consistent and formalized than the language of Twitter users?
http://s3.amazonaws.com/academia.edu.documents/41737546/The_...
http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.516....
EDIT:
If you want a really deep dive, you can also check out "African American English: A Linguistic Introduction" from Cambridge University press:
www.cambridge.org/us/academic/subjects/languages-linguistics/sociolinguistics/african-american-english-linguistic-introduction?format=PB&isbn=9780521891387
That being said, I think this research emphasizes that AAVE is anything but standardized. That's not meant as a pejorative statement: it's just acknowledging that, like most languages in history, AAVE has not gone through a process of codification and standardization to formalize it.
like most languages in history, AAVE has not gone through a process of codification and standardization to formalize it
I'm not sure what you're getting at - prescriptive grammars of the codified form you're describing are the products of their political and economic circumstances. There isn't some Hegelian trajectory of linguistic validity, where all variants aspire towards legalism.
Equally true, right? But somehow it seems just a little misleading.
Likewise, notice the deceptive switch between "English" and "standard English." So basically language bots trained to identify standard English are correctly identifying non-standard English dialects as not standard English.