Google Has a Striking History of Bias Against Black Girls
time.com
time.com
People make content for the web. People are racist, and make lots of racist assumptions in their writing. Society is racist. This results in Google's algorithms reflecting the corpus they scanned, and the searches that are made. But the search is just a reflection of us, and how terrible we all are.
Google gets embarrassed by the results, and exercises editorial control over search -- which leads to them actively removing racist systemic biases. Sure it's reactive and not proactive, but it's movement in a positive direction. Google likely spends more money on editorial control for issues of racism than for any other thing that doesn't make money.
And ideally, that's what it should be. Otherwise, how do researchers know what's really there, and what Google makes up? But in fact, Google is far from that ideal. Maybe it was initially unbiased. But successive waves of fighting SEO, and responding to social pressure, have taken their toll. Not to mention the increasing focus on what's happening right now, forgetting the past, and showing searchers what they likely want to find.
I disagree this is the ideal. The first implementation of PageRank makes a very big assumption: that people who write webpages and control websites can be trusted to link to worthwhile content. If webmasters and web producers skew a certain way, then PageRank/BackRub will also skew that way.
Google original mission statement was "to organize the world's information and make it universally accessible and useful." Being a mirror of web content would only partially solve that mission, and that's assuming webmasters/producers are acting in good faith. Would it really be ideal for Google to be a "mirror" of link counts for "miserable failure" [0] and "Did the Holocaust Happen?" [1]
[0] https://en.wikipedia.org/wiki/Google_bomb
[1] https://searchengineland.com/google-holocaust-denial-site-go...
Ditto for "black girls": either your looking for porn or for stock photos. What other intentions do you have with such a query? (Besides making a point?)
Google changing results for Holocaust "denial" is likewise wrong. Someone searching for such things probably wants to find all such sites and not just a PC version of history repeating what they've already heard. Otherwise they wouldn't need to be searching in the first place. Same for someone searching "Is the moon real" or "is the earth flat".
In a great market, Google might start "cleaning" up their results to the point of annoying people, who would then switch. Unfortunately no one comes close to Google search, let alone overall lock-in. So we're stuck with whatever Google decides reflects their brand, including fixing up well-publicized queries for people that want to be offended and complain about stuff.
Edit: I search for "best basketball player" and all I see are black guys -- way over-represented. Why doesn't Google bump up Larry Bird to 4th place so there's some more diversity and fight the stereotype? Even worse, Yao Ming is nowhere to be found, what's that say?
I think you misread the Holocaust denial situation. The reason why "Did the Holocaust happen?" was so problematic was because that literal string of text is something that you'd never find on a mainstream site, such as the NYT or the Holocaust museum. And yet sites like StormFront would get bonus SERP for using that literal text in a page's title, url slug, and <h1>. StormFront benefited from a blind spot in Google's heuristics that had little to do with whether or not StormFront was actually a legitimate researcher into the question of the Holocaust's historical reality.
Same problem with Google Bombs. A law student with a blog was able to get enough people to link "waffle" to 2004 presidential candidate Sen. John Kerry [0]. Presumably, the average person googling "waffle" wants the food item, not info about Sen. Kerry. But the reality was that the metrics Google had long used to measure quality and credibility were all saying "waffle => Kerry".
Staying true to the algorithm assumes that the algorithm (and its inputs) were pure and complete to begin with. This is not necessarily the case.
[0] https://en.wikipedia.org/wiki/Political_Google_bombs_in_the_...
Yes, I did. But I don't want to see the parasitic SEO-driven copypasta linkfarms. Or perhaps have them, but including such flags.
Consider the possibility that the user is a black girl, or her parent. Generally, I'm careful not to assume I can anticipate the perspectives of everyone else.
The phrase "white boys" in general seems to be mostly desexualized. You can call someone a "white boy", either as a mild insult (like describing something as "vanilla") or a term of endearment. I think that's predicated from most people (in the U.S.) realizing that white boys/men are the status quo.
I'll ignore race and just say that I think the word "girls" in general has far more of a sexual/gendered connotation in U.S. usage. "Boy/Man" is considerably more generic and all-purpose. I use "Boy/Man/Guy" in casual conversation regardless of audience gender, e.g. e.g. "Boy, that was fucked up!" or "Man oh man that's got to hurt". I can't think of a single time in normal conversation where I would ever say "Girl" or "Lady" without referring to a female.
My question was whether there is any connection between being a white boy and Googling the search term "white boys".
It was prompted by the parent post suggesting that not seeing a connection between being a black girl and Googling "black girls" might be kind of a blind spot.
I think the implication was pretty clear that if you don't imagine someone Googling the search term "black girls", it's because they occupy the position of the "other" in your mind. However, I gave a reason to doubt it - contemplate whether and how often white boys Google "white boys".
I note that somehow, I already knew the connotations of the term "white boy" that you mentioned without Googling it. So you are not bringing anything to the point I was making.
I am not a laissez faire fanatic about much of anything, including Google search results - I just think that when people criticize such results, the search terms (and the subsequent reasoning) often seem contrived.
Same for "professor style" and expecting to get something other than that stereotypical brown sweater with elbow patches look.
If I don't want to see whatever results are useless or offensive, I'll do my own filtering.
But then, I grew up on searching the scientific literature using Science Citation Index. And other stuff in LexisNexis. You don't expect stuff to be missing because it isn't cited much, hasn't been replicated, or whatever. You can filter by how well stuff has been cited, of course.
What does that mean, to accurately depict what's "out there"? That could be as nonsensical and unproductive as completely killing the spam filter in your email. Do you feel that that GMail and other modern-spam-detecting solutions have given you an inaccurate view of the "real" mail you are getting?
By the late 2000s, content farms such as Demand Media were churning out clickbait that followed the best practices with SEO. Demand Media [0] consistently beat out the NYT on SERP, thanks in part to a store of 1 million domain names used to serve up ads and links for topics ("howtomakeasandwich.com" would like to junk articles about sandwiches). By 2010, DM had 105 million unique visitors a month [0] and reached a market cap of $2B. The NYT, by comparison, had about 32M unique monthly visitors [1] -- across all of its properties (BostonGlobe.com, etc), and its market cap in 2011 was about $1.5B [1].
Demand Media beat the living shit out of NYT and every other quality media site by playing to the the known heuristics of Google's SERP. Every other media company had the opportunity and resources to play spamlord. Apparently they didn't, and by 2011 no other media company was positioned to create content (of any quality) with the scalable efficiency of Demand Media.
Google could have just accepted that "might makes right", but that is not an objective/neutral decision, nevermind a rational one.
[0] http://variety.com/2013/biz/news/epic-fail-the-rise-and-fall...
[1] https://www.theatlantic.com/business/archive/2011/12/it-cost...
I did hate the Demand Media bullshit, however. But as I've said, I'd rather that Google flagged that stuff, and let me filter it out. I mean, maybe I'm researching an article about Demand Media. How could I do it if Google is suppressing it? Could Google perhaps have an "expert" option? Analogous to "family friendly" or whatever it's called.
(Google does this with everything else as well. What do people think of the new MacBook Pro? Searching for reviews will immediately spit back a concentrated and distilled version the first mover opinions, which will itself drive peoples’ opinions.)
Examples of how the author is approaching the topic more generally are found in headings such as "Content and Creators" and phrases like "Information monopolies such as Google". The author's book is likewise entitled "Algorithms of Oppression: How Search Engines Reinforce Racism", not "How Google Reinforces Racism".
I agree that results in Bing and Duck Duck Go, for example, currently show a lot of the flaws describe in the article. But the point (as much as the headline may lead you to believe) isn't that it's only Google.
rayiner was saying that there is a "coupled system" and Google just reflects external racism, and then Google's search results amplify racism. The article did not mention anything about those interesting points, which might be worthy of bringing up and discussing, but it seems very unfair to accuse the top comment of "not getting the point" about them, when they are rayiner's insights that were only now presented.
I wouldn't assume that Google's search results reflect anything but an attempt to bring you back by showing you results you find valuable, and to monetize your visit.
> scrubbing it to better reflect a chosen set of cultural sensitivities?
Google (and any other mass market vendor) removes a lot of things that people don't want to see, including scams, X-rated content, brutality, etc. People searching for information on their dog don't want to see bestiality. These aren't arbitrarily chosen preferences or mere "sensitivities", but social norms which in many cases have a strong moral foundation.
Conversely--haven't these same people been telling us for years that we live in a racist society? So why are they surprised that a search engine of a racist society's web pages returns racist results?
> why are they surprised
Nobody said they are surprised; that's not relevant. They are trying to solve a problem, just like people who are working to resolve the problems of cancer, gun violence, and bugs in Firefox.
Just like the HP "racist" webcams [0] might have been caught had the CV trainers been more aware of what is/isn't in the training data, and/or if HP had a few more testers of darker complexion.
And perhaps Facebook's "Year in Review" rare tendency to be "cruel" [1] would have been mitigated with a broader group of devs and testers.
[0] http://www.cnn.com/2009/TECH/12/22/hp.webcams/index.html
[1] https://meyerweb.com/eric/thoughts/2014/12/24/inadvertent-al...
I think people who believe this significantly underestimate the difficulty of solving these problems.
We would deride Uber's self-driving vehicles if it turns out their systems can't tell the difference between harmless (plastic bags) and dangerous (jaywalkers, wildlife) objects on the road. Likewise, we can judge a company for releasing a general consumer webcam that fails to fully function for 10-15% of the American population.
And yes, these problems should be solved, and Google should use its editorial control to do the right thing. I believe it has done so, though perhaps not proactively enough. But these are hard problems, ones that are still not solved in other industries. Web search is still a young industry.
[0]: https://mic.com/articles/184244/keeping-insecure-lit-hbo-cin...
But then you get into an even deeper question 'are minorities/blacks 10% of said consumers of products?'. Poverty, for example, could mean that the product doesn't work for 5% of the buyers, which then begs the question 'how much effort do you put in fixing the problem for a small percentage of the buyers?', especially when the amount of effort is going to be very large/time consuming/expensive. If you are a person looking for racism, you will judge google as racist. If you are looking at it from a position of a business attempting to make a profit on a product, you will not see it as racist.
Maybe run that question by legal first! Protected classes exist for a good reason (IMO). What percentage of a restaurant's business is from people in wheelchairs, yet the law mandates ramp access.
[1] http://newsfeed.time.com/2013/12/17/delta-airlines-is-very-s...
i realize this isn’t widely appreciated or accepted in the age of critical theory, when social desirability leads people to complain about everything in the guise of “fighting oppression,” but it’s a simple real-world fact that no one solves problems from a position of ignorance
By comparison, I think most people don't even know the basic architecture of a computer, let alone the applications that run on it. Most people don't have an understanding of the network stack - I consider people ahead of the curve if they know the difference between LAN and the wider internet. Some people even think it's possible to download more RAM...
But ignoring that, no, I still would not like to see people professionally criticizing processes and products that they do not understand. Some technical knowledge is a must for people who do professional talks or presentations, write books and research papers and teach on the topic.
For algorithms and ML, I would expect at the very least a bachelors in math/CS or the equivalent in work/personal project experience.
They could have caught more cases before going to production - that's a truism. But when X number of cases get "fixed", there are other N*X cases/other communities/other minorities that won't, and there will be ad hoc public shaming that will caught enough attention to become relevant, and so on. Because the number of things that offend at least somebody is infinite.
My point is that there is no absolute way to tell what's offensive and what isn't - it's all subjective. There is no perfect way to fix them all - not even in theory, because there are conflicting priorities - and public outrage serves as an imperfect way to prioritize the most inter-subjectively blatant. Are there other ways? Sure, but they won't stop the periodic public outrage.
Is there evidence to this effect? There are a lot of things Google spends money on.
I'd add that it does make money. Making your customer experience pleasant is an essential money-maker. If people are loudly saying racist things in your restaurant, the financially essential decision is to exercise editorial control before you lose your other customers.
I used to be (well, still am, a bit) worried that the increasing automation of our social structures (credit scores, admission processes, etc) will just encode these biases in ways that cannot be altered.
But then saw how much of NIPS was about "fairness".
If these biases get encoded, they are measurable, examinable, detectable - and can then be accounted for in ways (presumably) much more effective than existing patterns.
So now, I'm not worried about the technology, although as always, I worry about the application of it.
If I don't like what I see in my mirror, will putting makeup on the mirror improve my face? Will removing naughty sites from the search results somehow change the fact that "black girls" (or "asian girls", or "czech girls", etc) is not usually used as an innocent search term?
Google's search algorithm is "fair", but society is not: It is not "fair" that these search terms are over sexualized.
This is, of course, given some particular morality and ethos; your may not find society unfair in this way. (I do)
My point is that now it is visible and you can make more informed choices.
If you don't like what you see in mirror, why is makeup the "fix" of choice? Alter diets, exercise patterns; sun exposure, fitness levels, stress levels...
> Will removing naughty sites from the search results
Potentially. We're feedback loops as well; if these terms result in what someone is after, they will use those terms: if they don't, they won't. Eventually, I would expect people to stop trying to search with them; to then stop communicating with them (as meaning that); to then stop thinking with them (as meaning that).
First of all, to go with the obvious, you better hope that what Google deems the direction they'd like to alter reality is aligned with your own, because I'm quite sure that this will mostly be untrue.
Secondly, you're attempting to patch a problem at much too high of a level; making people less racist by altering what they're able to ingest, particularly by altering unbiased reality, is a bandaid that is destined to fail hard. What if, say, we wanted to change the negative stereotype of black Americans engaging in criminal behavior at substantially higher rates, so we demanded that Google remove access to racially distributed crime statistics. Many countries already do this, such as Sweden. We will have lost a primary tool with which to assess the nature of our reality, contextualize it, and strategize about it effectively.
I'm more interested in a free society bettering reality than I am with hoping that Google's ideal version of reality is aligned with my own... particularly because I would still consider it to be a horrible solution even if it was.
The "reality" in this case is not unbiased, because it's not the reality of "cars are hard" but the reality of "red means stop".
> access to racially distributed crime statistics
That does sound like something that would backfire or fail. Seems like you hit on the delineating factor: You search for crime statistics to contextualize and understand your reality. Is that why people search for porn?
After your proposed changes it is no longer visible, and we can no longer make informed choices - people searching for "black girls" still want the same thing, but they're getting a worse service (results that don't fit their intent) just to hide them and their preferences from other parts of the society. There's also a feedback loop in the opposite direction, if there are real problems that simply are hidden, they're not going to get fixed.
On the other hand, it's worth examining the actual problem - it's worth noting that people who do want to find sexual pictures of black girls are just as valid users as black girls, and any fair solution must keep both of their preferences in mind; it would be unfair to simply damage the service for one group to appease the another. IMHO a fair solution would be based on data about what the people who use that query actually want (in some equivalence to picking equal opportunities in the "equal opportunities vs equal outcomes" tradeoff dilemma) - if some people are using the term in their daily search and some are not, then the interests of these frequent users could reasonably be considered more important. Aggressive personalization (different users having wildly different search results for the same term because of the differences in the assumed intent) is another way to do that, but it brings a bunch of other problems.
That said...
You're missing that some problems are, actually, entirely perceptive. You're right that changing the mirror to hide problems isn't addressing the underlying issue, but sometimes treatment of symptoms is the only available (or possible the most effective) option. Like everything else, it's one tool that can be applied poorly or effectively.
> it would be unfair to simply damage the service for one group to appease the another
This is ineffective as a stand-alone principle. The failure mode is described as "your freedom ends where mine begins" and is clearly visible in the exaggerated (and somewhat contrived) example of "damaging services for murdering to appease those who want to remain un-murdered is good".
Like all principles, this has both extreme-case failure modes (services for murderers) but also fuzzy-case failure modes. If the term for my identity is sexualized, when that identity has nothing to do with sexuality... well, what's fair? Why should group X define the term instead of group Y?
So, for cases where the problem is primarily perceptive (or the symptoms are primarily perceptive), seems reasonable to me to warp the mirror. It's like using glasses: you're already looking through a warped lens, so use a correspondingly warped lens to correct.
My actual point is that systems that incidentally encode biases then let you better examine those biases, and, depending on the situation, do stuff to address them.
I don't think that's news to the author, or a refutation. She says tech companies produce tools in environments of racism/sexism, without proper safeguards, and the result is damaging to people of color and women. She's saying Google can and should consider this stuff in advance, and take responsibility for bias in their products (even when it's the reflection of ugliness in third parties as you said).
So, I don't see your perspective as being in conflict with her. Except inasmuch as you excuse Google for reflecting our biases. Do you think they should feel OK if their next AI winds up with our biases? If so, why? It doesn't feel like you've explained that.
I think we all agree Google should do as much as possible to improve results. But at the same time, if Google gets it right 99 times out of 100, the 100th is what will get an article like this written on it. That doesn't excuse that, it explains why it happens.
> Your second paragraph explains why Google, the search tool, has a striking history of bias against black women
The article says "Google has a striking history of racism", which you interpret as "Google results have included things which are strikingly racist." That's one interpretation, but it seems equally valid to interpret "Google has a striking history of racism" with "Google has been very racist" (after all, if I said, "Bob has a striking history of racism" you'd be right to assume I think Bob is a racist).
I think some of the controversy around this article stems from these interpretations, one of which is a far worse accusation than the other.
Is there any substantive basis for this claim? My impression is that Google gets away with a lot.
As one anecdote, http://proceedings.mlr.press/v81/buolamwini18a/buolamwini18a... reports facial recognition algorithm performance on darker-skinned women up to 30% lower than lighter-skinned men. The preliminary results were reported to Google, IBM, and Facebook. IBM was able to tune its algorithm achieve comparable scoring across all skin hues/genders within a week (according to my notes from the talk). It could be that search on that scale is much harder than facial recognition, and its also that a quick fix would introduce other errors, I'm sure it's an NP-complete problem, yada yada.
But Google and racialized search is not a new issue https://dataprivacylab.org/projects/onlineads/1071-1.pdf.
Yes the society is racist, but then the U.S. enacted the 13th, 14th, 15th and 19th amendments to address extreme racism and gender discrimination. That is, citizens -- responsible engineers and data scientists among them -- can take it as their obligation to confront systemic racism. Yes, Klansmen are racist, Nazis are racist, but we don't have to accept genocide -- we don't even have to accept biased, discriminatory and derogatory internet services either.
I'll bet at least 50% of the searchers are looking for porn.
So even if Google had infinite resources, its choices are: (1) act as a moral authority and deny pornographic contents to those who are looking for them, or (2) show pornographic results to people searching for innocent stuff.
...Or maybe (3) built a perfect profile of every user and just show them what they want. But do we want to go there?
(You might be thinking that "black girls" is such an innocent term that Google should be able to determine that it's non-porn, but then the question merely shifts to less popular queries.)
So if the top ten results are black women (or Asian guys, or British royalties, or whatever), and if they look funny enough, then users will be satisfied and stop looking further. So whatever pictures that went up there first stays there.
Would there be any less porn? I strongly doubt it.
"Girls" is the HBO show
"Asian girls" is porn lite
"Black girls" has no porn for me. Various sites about black girls (black girls code for example)
I tried "black girls", "white girls", "asian girls" - and also "asian boys" etc. - in various accounts, some of which were used when browsing porn (i.e., home accounts vs accounts used at work), and in both google and bing.
In accounts where porn was browsed in the past, porn results were over 50% for all of the above. In others, porn was absent.
Btw, black girls no longer returns author’s screenshot so google has changed results since this article.
Even that might not be enough to get the original results, as there are also signals based on location.
In general, in our Western society, the whole breast seems to be fine, as long as the nipple is covered up. So, what they could do is show you nipples, if you actually search "nipple". But maybe you're looking for male nipples, which obviously are not pornographic at all, so that doesn't work either.
"Vagina" and "penis" have enough educational material behind them that that may actually be what you're searching for.
"Pussy" could be a cat, "dick" could be a Richard.
Which is another aspect, making this harder. Our society loves using non-sexual words to refer to sexual things.
I suppose, if you actually throw in the word "porn", there's hardly ever going to be a scenario where you were not looking for pornographic content, but it's not entirely impossible either.
And that more than 90% of all searches with "girl" are for porn. Google already killed most of its porn-lovin' demographic with the beheading of Google Videos in ~2009.
Also, I don't understand what you're trying to say here. I'm reading sentences, and I comprehend them, but there lacks an overarching idea to bring it all together.
This is well-known, and not Google's fault. Getting offended at Google for a reality they didn't create and can't control seems silly to me. If you don't like the result, work to change the incentives. Or work to change people's biases.
This kind of problem is not new. For example insurance companies have long known that where you live affects how much they are likely to pay out. If they base insurance rates on the data, the result is that your zip code becomes a bigger determinant of insurance rates than your driving record. Which results in very large effective racial discrimination. There are laws limiting that, for example prop 109 in California. However it is an eternal struggle because, in fact, competing insurance companies have a good motivation to have the cost of insurance reflect their best estimate of the expected cost of insuring you, and their cost really is a lot higher if you're in a "black neighborhood".
>> What we know about Google’s responses to racial stereotyping in its products is that it typically denies responsibility or intent to harm, but then it is able to “tweak” or “fix” these aberrations or “glitches” in its systems.
Only Google knows the fate of the code base that has workarounds for all of these edge cases. I think the author is wrong that all of these edge cases could’ve been caught beforehand. Some sure, but I think others would always slip through.
What you just observed and described is that technology _is_ unbiased and robotic. The bias is in the humans. You can't remove it with a software bandage; the software is working as intended.
'On two occasions I have been asked, — "Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?" In one case a member of the Upper, and in the other a member of the Lower, House put this question. I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.'
But I think the problem is more subtle than you give it credit for. The software may be working as written, but 'intention' rarely maps cleanly to execution. The behavior of a complex system is a multidimensional shape; stretching it one way or another along some axis according to your intention for its behavior can easily have unanticipated, and often unobservable effects along other axes that you were not minding. So it may be an intention that the system optimize for click-through rates for ads, while at the same time not consciously intended to be biased against, for example, race, but also not deliberately crafted to avoid this behavior. This may functionally be the same thing as intending it to be racist, but it has different implications for who's guilty and how things get fixed.
Modern computing systems have components that produce output that's "better" than their input, at every level. Because we have information theory now, we know it is possible for a machine to produce good output even when some of its input is bad.
Google, incidentally, has aimed to do that for its entire existence. From the first time the Google search engine was deployed, its job was to take in the messy Web, full of misdirection and spam, and show you the useful pages you were looking for. And we can keep hoping for Google to keep doing this as the threats to information become more insidious.
Just to get calibrated: what is an example-- real or imagined-- of technology that is biased?
If I'd have to make a definition, then it would probably be a technology that somehow encodes an assumption that doesn't match reality. In this context, saying that "black girls" is an innocent search term and users searching for it don't expect to get sexual results is such an assumption - it would be polite, politically correct and possibly socially desirable; but it seems likely this is simply not true in the reality we live in.
Just as our (very straight laced) CIO walked up behind, employee hit image search on "snow bunny". Hilarity ensued...
Though some people might argue that the biases software reproduces will in turn reinforce biases held by the users at large. Conversely, you may reduce human bias if you algorithmically reduce the bias in the data (which is possible). Not saying whether Google should, that's a very delicate debate, but it's thinkable.
I wouldn't want to be in the place of the person at Google trying to make guidelines for that, but it's definitely doable. For instance, you could probably easily argue that the phrase "three black teenagers" (an example from the article) is pretty neutral in itself and definitely shouldn't be interpreted as a call to produce arrest pictures. In general the term "three [ethnicity] teenagers" could probably be normalized with regards to the setting the pictures show, for instance.
Turns out you can. It just takes effort.
And you should, because amplifying shitty things about people is not the ideal state of software.
The question is, should search algorithms show what currently is (and therefore reinforce what is), or what the culture aspires to (which in fact more accurately matches the culture in terms of desire and movement)? I think this is a pretty smart example of how the technology we use every day can be passively malevolent - creating hurdles for change rather than enabling better information transfer.
I believe expecting Google to change its search results to match an ideological agenda (I'm not using "ideological" in a derogatory sense here) is a case of shooting the messenger. It's the job of humans to make society in the image they desire and the job of the search engine to efficiently and accurately give the search results that people want.
[1] https://www.theverge.com/2017/10/3/16413082/google-4chan-las...
Let's be realistic - if someone is explicitly going out and searching for "black girls", would you reasonably assume that in the majority of cases they'd be looking for clothed ones?
I understand that different results could influence behavior and make a correction - but that wouldn't be correcting a measurement bias to better reflect reality, that would be an attempt to change (improve?) the underlying objective reality, which is conceptually different than correcting bias.
radical thought (and goes for social media as well) -
how about making the results tunable based on user preferences & filters?
oooh wow.. user control! what a concept..
Client-side controls > Server-side controls
If they are willing to rethink how people navigate (both the real world and the internet), etc - if they want to dramatically change how easily people are able to access information - why should they be bashful about nudging culture as it relates to how different subpopulations are represented?
I think that implying intent on Google's part is going way too far.
Obviously search results, to a large extent, are going to reflect the society and culture of search engine users. If our society and culture are shitty, should search results pretend otherwise? I'm pretty sure that our society and culture are rife with institutionalized racism and sexism and a whole lot of other badness, and badness of search results seems to me to be a symptom rather than a cause.
In the old days it was "most linked page with those keywords," right? So if you searched "gay man" and got a bunch of gay porn, is it necessarily google's fault that the internet is more gay porn that it is resources for gay men to discuss LGBT issues? Is it google's fault that news agencies more typically report on crime by black people than white (or that black people are more likely to be convinced/arrested in the first place?)
I think Google is responsible in 2018, now that they offer extremely tailored results. Duckduckgo, maybe not. But if Google is going to say "we're going to show you listings for restaurants in SF in your simple 'restaurants' search because we know you live in SF," I think they should also say "we're going to show you results for LGBT resources when you search for 'gay men' because nothing about your query indicates a desire for pornographic content."
This isnt that different from search engines for specialized databases. I used to use meta-search engines that woukd have checkboxes or selection menus to focus search on specific topics or collections. It was very useful. I still do it for technical papers using the "site:" operator to limit search to known-good sites for them.
In general, there's the "safe search" option that's on by default; but if someone has turned it off, then it's reasonable to assume that the default, most common intent of looking for "girls" actually is sexual, and if they were looking for something else (say, "black girls support group") then that would be a comparably rare situation where extra words should be added to the query.
This is a hard problem. There are many search queries, there are many clusters of similar search queries that address partially overlapping (topic,audience segment) pairs. Deriving sentiment scores from these segments is difficult without any feedback mechanisms. Choosing which segments to focus on is difficult or counterintuitive and depends on the topic and audience and characteristics of the audience. Once a topic and audience segment is identified for remediation, the task of effecting change is again topic and audience specific (resistant to automation).
This is a hard problem. Neither a brute force "human review squad" approach nor an automated "deep learning" approach will provide 100% coverage through all time. Does that mean that they shouldn't try? No, of course not, and they almost certainly try every day. However, it's inevitable that they will be subject to sniping articles no matter what they do.
There are many search queries to watch, but 'black girls' was persistently problematic for a long time. Noticing when a cluster of queries gets that status might be extremely resource intensive if that status was highly ephemeral; it's not. It's work that can be done asynchronously, and run daily at most. Is this simple? Nothing involving software is, imo. But it's probably much simpler than many other machine learning-driven pieces of Google's product.
Also, there are tons of feedback mechanisms available and actively used by Google. Every user interaction with search results is available to Google; a lot of these stand in for quality of result: did the person refine their query after seeing crappy results? Did they click a link and then press 'back' really quick? All of these factors already feature in Google's algorithm.
They did exactly this in 2004, when searches for "jew" brought up anti-Semitic website "Jew Watch" as the top result (because, as a linguistic quirk, "Jew" tends to be used in a slur-like way and Jewish organizations tend to use "Jewish," and the algorithm at the time considered those as different words): they placed an ad at the top of the page with an explanation of the "Offensive Search Results," distancing themselves from the content and explaining the "Jew"/"Jewish" thing.
https://en.wikipedia.org/wiki/Jew_Watch#Google_Search_result...
http://sethf.com/anticensorware/google/jew-watch/jew-watch-c... (and yes, that's Netscape Navigator!)
http://web.archive.org/web/20050123081919/http://www.google....
This isn't fixed by telling Google it's their fault and they should hide things, that's just looking away and covering things up. The true fix is making the input better, which requires society being better, which it isn't (at this point).
Because they are not. And neither is Google.
TPB, as described, is in a moral gray area - "We know some of the uploaded stuff is immoral" - so they delete what's immoral (their judgement - child porn/copyright stuff/anything between), they arbitrarily decide what you see
Google is in an algorithmic gray area - "The algorithm decides what's best for you, we want the best for you, but we won't judge what's best for you" - so they do nothing and hope for the best (and least damage), the algorithm arbitrarily decides what you see
Both seem like different routes to the same hell.
Edit: Spelling, clearer wording
TPB said "We are not responsible for the content people upload." Except then they deleted CP, probably for many good as well as self-serving reasons: CP is bad, obviously, but also brings bad press (hard to get the public on the side of piracy when it leads to ease of CP access), as well as the holy wrath of pretty much every criminal federal agency in countries around the world, as opposed to just whatever branch of whatever agency is in charge of copyright protection. Doubt they could get hosting in any country if they didn't take the CP down.
But by taking CP down, they demonstrated that they do and can monitor content, that they can take it down, that they do take responsibility for at least some of it. Legally that's a shot in the foot, I thought, but they still find hosting so who knows.
Whatever proxy they are using to find relevance is broken in this regard, and it's important to recognize that there is still an engineering problem despite, or despite it looking like, a political problem.
Are you are saying that, in a scenario where there are more results for one sub-query than for another, Google has an obligation to assume someone is searching for the sub-query it finds more politically acceptable? This seems like a bad idea to me. We should want to reduce the political/social influence of Google, rather than think of ways to give them more tools of influence and assuming they will have a positive effect.
These two things are completely different
"Because you live in SF, if you search restaurants you probably want results pertaining to restaurants close to SF"
"Because subject X-A is socially unacceptable, if you search for subject X we will only show results pertaining to subject X-B"
> Are you saying that ...
vs
> You are saying that ...
If I search for "restaurants", it shouldn't give me restaurants in Shanghai because that's the city with the largest population. It should give me Northern California, since that's where I live and Google knows that.
Similarly, Google might reasonably infer that if someone is searching for their own race, sexuality, or religion, they're probably more interested in information or support groups than porn. Not that both can't be served as results... just... priorities of what people are looking for.
Sure, if they have data to back that assumption up. My interpretation was that the parent commenter wanted them to make assumptions just based on politics, which I think is a bad idea.
Don't they already do this for porn? I thought that feature was in place years ago. It's even a minor Internet joke that Bing is only good for porn (since they're more lax about letting porn into their results).
Ideally it should know nothing about me!
At some point you've got to accept that googling "hot teens" won't get you resources on the treatment of hyperthermia in young people.
In reality though, the world is a complex place and the ideal pretend version can only be implemented in a carefully controlled environment such as a movie, a theatrical presentation, a video game or a situation where all participants agree to or are pressured into behaving in a way in which that pretend world is real. When the rest of the world leaks in, it is impossible to maintain that pressure on everyone.
About the best we have so far is "safe search" since pornography and obscenity follow fairly regular patterns. Political correctness though is constantly evolving and requires a trained academic to determine if each piece of information should be censored or allowed. Ideally in the future we'll have google glass and advanced AI classifiers constantly scanning our vision and providing us with tape delayed audio of the outside world in order to present the world to us and filter any politically incorrect content 24/7.
For now we have to do it ourselves and ignore any inconvenient facts that may make maintaining the illusion difficult. This constant struggle to deal with cognitive dissonance can be tiresome and one should restrict oneself to heavily moderated news feeds and not using search features for now. To create the carefully controlled virtual world in an interface to the rest of the world such as Google is an overwhelming task because you would need a highly sophisticated trained AI agent to maintain that illusion and to even invent information, such as lists of Nobel award winners from disadvantaged peoples, in order to maintain the illusion.
This viewpoint has served me with a perfect success rate in all of my interpersonal reactions, including relatively combative ones (handing a beer can a drunk has thrown on the street back to him and asking if he could throw it away next time, random example).
Are there any specific example of things you think are incompatible with optimism without horse blinders?
That is the opposite of how it works. Affirmative action is accepting that the world has bias against certain groups, and attempting to take that into account. Political correctness is about realizing that certain groups have been marginalized or traumatized and some words perpetuate that. In a perfect world neither are necessary.
At least that's the way it works in reality. Maybe this reality is politically incorrect too and you are saying that we need to implement meta-political correctness in which the means of creating an ideal world must be hidden to instead pretend that the implementers know the exact and particular circumstance of each person they are selecting based on facts other than merit and are weighing all of them appropriately in correcting injustice. By pretending that this ideal world in which knowledgeable administrators skillfully and in each individual case correct injustices, this will somehow make it become a reality.
Yes, because implementing something like that would be near impossible from a legislative standpoint. Legislation has to be spelled out and enforceable or it will do nothing.
And affirmative action is not there just to prop up the numbers. It is there to push companies to hire minority groups.
There are two reasons for this. For example, on college campuses, the supreme court has upheld that it is lawful to use race as a factor in admissions because a well rounded or diverse student body is a desirable trait that will produce better outcomes for all involved. Including the white men.
The second is that it is trying to undo the many, many years of oppression (i.e. slavery) that carries generation to generation.
Programs like affirmative action are trying to tip the scales for these descendants to give them a chance to compete. After all, they were systematically disenfranchised as a race. It makes sense that we attempt to systematically bring them back up to speed with the rest of the population.
>we just assume .. numbers need to be fixed first and then the rest of it will follow
We don't have to assume anything. It is both empirically evident (as well as just "duh") that getting a higher percentage of minorities into jobs and higher education will result in a higher percentage of their children having the opportunity to do so without something like affirmative action in the future.
Statistics show that man are the new minority in the campus. But it is surely will be frown upon by some AA activists if a that program is extended to cover male. If AA is said to what it is trying to accomplish, then that shouldn't be controversial at all.
I think this is presumptuous. We can try to find out, why not?
Agree. It hides problem rather than fixes it. And since it is often enforced in a top-down manner, communities that outside of the PC spectrum felt being deprived of their opportunities, this very resentment is further diving the society.
What is "political correctness?" Are we referring right now to the alt-right manufactured bogeyman? That is my sense.
I am replying in the following manner only because I believe the discourse around this subject is usually poorly defined. It is not an attack of you or yours, but my attempt to help us all better understand what we're talking about:
By saying "filter through it," what do you mean by "it?" What are the rules of the filter you are applying?
When you say "expected," who is expecting this of you? What has indicated to you that there is this expectation, and what are the consequences of bucking that expectation?
What are the filters in California? What are the filters in the deep south? Is there any federally legal relevance? For example, you could face legal issues if you make sexually connotative comments to a female coworker, and though you are more likely to be pursued and charged in California than in the Deep South for it, is that how things should be? Shouldn't women feel comfortable with their coworkers regardless of local culture?
The problem is what makes a woman feel uncomfortable is culturally driven - when you have such a mix of cultures at work people are bound to make innocent mis-steps.
Californian companies have a tendency of looking at it by race and gender identity. Do you assume the best of people or do you assume the worst?
I understand where you're coming from - maybe you're referring to things like "micro-aggressions" and other things many of us would find silly? Yup, definitely something that needs tackling. I've never argued that there would be a clear algorithm for "this makes people uncomfortable to the point that it should be illegal, this does not."
Luckily, our judges and court system are pretty good at exactly this sort of thing. I think a law on the books that goes something along the lines of "nobody should be made to feel sexually uncomfortable at work" starts the definition, and then subsequent court cases sharpens it. Sorry, Tim saying "great haircut!" doesn't count, but oops yea Janice saying "quite the cucumber you're hiding away there, Bill," while waggling her eyebrows does. Etc.
What makes you think that?
Millions of Americans have been negatively impacted by bad laws on the books.
> I think a law on the books that goes something along the lines of "nobody should be made to feel sexually uncomfortable at work" starts the definition
You have just killed all chance of workplace romance - and more than 50% of people have dated coworkers.
We have sexual harassment laws on the books already.
Why would you negatively impact millions so a few people don't have to feel mildly uncomfortable (and obviously it's mild or it would be sexual harassement which is already illegal)?
It seems to me that there's an enormous contingent of people who wants to yell about the influence of Google and Facebook and then yell even louder about how they're not actively intervening to skew content towards their political preferences (ie increase their influence massively). This split-personality approach is baffling to me.
I'm getting a sense of that from your post.
I come across thoughtful people all the time who have a baseline assumption of intelligence and intellectual honesty that doesn't match almost any group you're likely to actually come across. My prior for people having obviously contradictory views is simply higher than yours.
It's pretty straightforward to understand the view.
1. It would be better if Google and Facebook didn't even have the power to skew content.
2. Since they do have that power and seem determined to use it they should be judged on their choice to use (or not use in this case) and implementation of content skewing.
It happens even more easily if you're lumping people into a single group, but the individuals that you see speaking out about one topic are different than the individuals that you see speaking out about another.
The upshot is that from your described experience you cannot conclude that any group is being self-contradictory unless you personally witness the same person contradicting themselves in different discussions.
Shouldn't helping connect poeple with what they're looking for be the bottom line? Even if you have to manipulate your algorimths natural results to do so.
Before we condemn Google for the collective actions of the broader Internet community of websites, we should take a moment to understand this point.
Now, it is possible to create a search engine curated by humans or filters that remove objectionable content, and perhaps that approach might be desirable to some - a "MyGoogle" experience with search results customized based on Google's interpretation (or explicit settings) of the visitor's demographics, political beliefs, trigger words, and purchase history.
PageRank is only a fraction of the total inputs to Google search. Once PageRank was described it was subject to being gamed at such a scale that it essentially destroyed the web, ironically making Google that much more a necessity.
> Google and its human staff does not apply human intellect to the problem of sorting, picking, and ranking each search result.
Yes they do. By tweaking the weights of certain output categories (such as evidenced in the article, a clear drop in the number of pornographic or semi pornographic results as a result of such a tweak) there is a large amount of influence exerted on the results.
> Moreover, Google has included "Safe Search" features for several years that would remove these types of results from search.
Safe search always was a weird one: You'd expect the opposite, a 'smut search' (Tom Lehrer would have a field day with that one).
> Before we condemn Google for the collective actions of the broader Internet community of websites (and the author's step of disabling Safe Search), we should take a moment to understand this point.
Before we invalidate the authors point by handwaving and purposefully injecting chaff into the conversation let's try to understand the actual point they are trying to make: That a query for an innocent term such as 'black girls' even with 'safe search off' should not result in a bunch of porn.
> Now, it is possible to create a search engine curated by humans or filters that remove objectionable content, and perhaps that approach might be desirable to some - a "MyGoogle" experience with search results customized based on Google's interpretation (or explicit settings) of the author's demographics, political beliefs, trigger words, and purchase history.
That's nothing to do with the authors point.
I take strong issue with the "obvious" conclusion that Google knows which terms are innocent. Disclaimer: In this specific case, I agree with you that "black girls" has an innocent connotation and am not suggesting otherwise. Please do not misinterpret my thoughts below.
Shall we require Google to be ever vigilant about the meaning of words in use at various times by the various communities of the world, and to entrust them with the responsibility for determining the innocence or guiltiness of words or phrases for all mankind?
If we shall require such effort of Google, ought we not elevate Google to the role of judge over other matters of humankind given their good stewardship over matters of innocence and guilt?
Taken to the limit the argument advocates the establishment of an echo chamber to wrap tightly around each one's subjective interpretation of their reality. Isn't this the exact opposite of what we should be trying to achieve?
Have we not had enough of this Brave New World?
Google aims to 'organize the worlds information', how it does that concerns all of us so yes, this makes them responsible with determining whether a certain set of words has an innocent or negative interpretation.
It's not the world I want to live in but I live in it nonetheless. That Google should not have this power to begin with is obvious but here we are.
> If we shall require such effort of Google, ought we not elevate Google to the role of judge over other matters of humankind given their good stewardship over matters of innocence and guilt?
Definitely not, why make things even worse? They already have too much power we will not make things better by compounding the problem.
> Taken to the limit the argument advocates the establishment of an echo chamber to wrap tightly around each one's subjective interpretation of their reality. Isn't this the exact opposite of what we should be trying to achieve?
It is, but again, the filter bubble is real and the more customized Google search results are for an individual the worse this gets.
> Have we not had enough of this Brave New World?
I do, and so does the author, but clearly you are either missing the point or you have not had enough yet.
Is it your argument that they should be prevented from doing this because highly personalized, low quality results might instead be the result of their efforts?
> That a query for an innocent term such as 'black girls' even with 'safe search off' should not result in a bunch of porn.
I don't know what results the search term 'black girls' should return. If you just search for it, it doesn't return any porn for me. If you're seeing porn, your search results make reflect your search history since Google will customize your results based on your search history. Interesting.
That being said, it should return whatever people are searching for. We must be careful not to cross the line from fixing mistakes to politically correct moral censorship.
The point should be to slowly break down the bubbles, make them less of an echo chamber, and/or have them merge at points where they can. To anthropomorphize the "perspective" bubbles, we are in the dark ages, where people really don't travel very far from where they are, and they don't even speak the same language from their neighbors a short ways away. Allowing a translator and free exploration of other bubbles could allow for more alliances - the zero sum troll-only interaction is something nearly everyone see's is not helpful.
If this post got widely shared, Id see responses in Facebook from people including several types of feminists, business woman arguing cold numbers/profit, males/females arguing black angles or activism (had one just like researcher in OP), male/female opponents of both (white and non-white), activists citing whoever gets ignored most (eg nativelivesmatter to BLM activists), at least one reactionary racist whose dismissals will be common, and thoughtful oddballs who I cant easily predict. Each will have reflexive comments or even articles with evidence for their view or against others.
It can be overwhelming to look at it all but I do on critical issues to ensure Im basing my views or actions on big picture rather than a filter bubble. If not a big issue, Ill just lazily click on my feed which is definitely a bubble of folks with similar interests, the most discussions/tags/messages, or most likes. I noticed Facebook ruined my scheme more over time where I have to actually click on specific names to see those kind of view points. Ive also been told Twitter would be ideal for this and know folks that do it there but still havent got on it myself for common anti-Twitter reasons. ;)
So, you can currently do that with both platforms. If you pick representative people (gotta do that right), you will see about everything they believe as a group about what's trending in their groups or nationally. You will have more good info and see more info as bad since tunnel vision the alternative approach creates is part of why bad info spreads. Lastly, it's just really interesting if you're a people watcher like me to see all the different ideas, jokes, trends, lifestyles, etc people share. I didnt originally add these people to study politics of rough topics: each was interesting and/or helpful in discussions, often face-to-face, before I added them. So, it's fun, enlightening, and sometimes challenging to have truly, diverse friends around you.
Note: Might be worthwhile to try to design some kind of meta-search engine designed to mimick what I just described. My quick guesses at how to build it suggest it would be really tricky to deal with human factors. Ill just leave it at saying the concept would be useful.
But if no one is searching it, then of course it's just going to be semi-random results.
https://trends.google.com/trends/explore?geo=US&q=%22three%2...
The top related query? "Three white teenagers". It's obvious almost no one searches this term unless they've read the article.
Under the heading "Where are black girls now?" she notes that she doesn't see porn in the results anymore and after I googled black girls I got similar results to what she shows in the picture directly under that, where the first link is black girls code.
* Taking a neutral position when building something (its just a tool, it's making decisions based on an aggregate) IS in fact making a decision. State can be -1, 0, or 1 and neutrality yields a certain result.
* The local optima of these algorithms are likely a product of local optima to the people and teams who built them. The neutral position seemed like the correct answer.
* These articles serve to reinforce the idea that the next frontier is context within algorithmic results. This global optimum can only be achieved through diversity of perspective. Perspectives are a product of one's experience, so the value of diversity can be directly connected to the goal of desiring a optimum solution towards a hard problem.
FWIW, I am a black heterosexual engineer and my experiences inform my perspective on these type of problems and blindspots affecting people not like me, especially women.
For those wanting to go deeper on the technical, scientific, and mathematical basis for what I'm referring, google 'diversity local optima'.
Males searching for “FOOBAR girls” have, on average, different intent than women. If they represent the majority of searches, algorithms will naturally weigh the results they click more highly.
Hyper personalization is seen as an answer, but comes with the downside of a reinforcing bubble and drift to extremes. Human intervention and editing are only a partial solution, bringing up further questions about censorship and which ideology is correct.
No easy answers to all of this.
As much as I hate Google, they seem to be delivering exactly what they should. Their job isn't to try to change people's minds, it is to find what they are looking for. What's next? Search for "geek guy" and it's all white. Should we be upset? Someone searching for "geek guy" is probably looking for exactly what Google returns in this case and all the others she highlights.
As far as "black girls"... What kind of results are people expecting? Who is actually searching for black girls? Probably people looking for porn... Why should Google return suboptimal results because you got offended?
Same for "three black teens". It's not Google's fault that so much news is reporting on crime committed containing that phrase. Did she actually check crime statistics? Maybe "three black teens" get mugshots and commit crimes together at a vastly higher rate than "three white teens". Outside of porn, who is searching for this phrase? If you want "clean" and friendly results, try searching "stock photo three black teens".
If anything, this article shows that Google's personalization should detect people that want to search things (that they wouldn't really search for in earnest) in order to be offended then display poor quality results to avoid this "issue".
In truth though, I would imagine that if she and other black women often searched "professor style" then didn't click any results and immediately searched for "professor style for black women" then Google would pick up on their results being suboptimal and change. My guess is that no one actually makes these searches outside of making a political point, so personalization cannot fix it.
So yes, I agree that more pressure needs to be put on the algorithm-makers, so that there isn't complacency about accepting the results of "pure" computation. But I feel too much weight is given toward the sentiment of "If Google isn’t responsible for its algorithm, then who is?", as if the problems could be wholly or even substantially fixed through "modifications to its algorithm". The other side of the equation is the data -- and diversity is as important to have in media as it is to tech, even if tech feels like the stronger, more immediate lever in changing things.
Now that it's been a long time (in tech time) since 2010/2012, it'd be interesting to hear exactly what modifications/improvements Google made to its algorithm to return such different results for "black girls". Was it a de-emphasizing of negative/outdated sources? Or an arbitrary boost for such entities as "Black Girls Code"? Or was BGC not given an arbitrary boost, but found to have been unjustly ignored/under-weighted when it came to determining SERP? I understand not wanting to reveal details of search engine tweaks the day/month/year after they've been implemented. But half a decade later, hopefully learning more details of these "tweaks" won't leave Google vulnerable to black-hat SEO.
Her point is quite valid, and also well recognized in at least some part of the tech industry: when you "blindly" optimize for a certain audience, you risk picking up the biases of that audience.
So, if you want to appeal to a broader audience than your current one, you might need to manually tune your system to remove some of the biases from your original audience.
This is particularly important, if not from a human perspective, at least from a business perspective, when the bias of your current audience is about the larger audience.
They can stop relying on algorithms because the underlying data reflects social bias, or use them and reflect society. But, they are always going to upset some group some of the time because social norms are not consistent across society.
IMO, attacking Google on this is pointless, deal with the actual problem.
Yeah, great contribution right there! If we haven't cured cancer is because we aren't thinking hard enough!
I wish Google wasn't biased against the restaurant around the corner from me when I search for "The Trumpet", but it is. That's as it should be of course, since the audience for that result is so small.
All we're really talking about is a project that will never be completed of tailoring search results for the largest audience possible -- which is inherently a moving target.
Doing it manually as you suggest is likely hopeless and more politically fraught.
Yes, but, sadly, necessary.
Anyway, that's Google's call to make.
These days we have systems that is not as smart as us trying to figure us out, with predictable results. Tricks and quirks are ever shifting. A query that used to produce the results yesterday might not function the same two weeks from now. Any skill gained in navigating the system needs to be constantly trained and updated.
I believe something similar happened to Hawking, where he had a keyboard that was not smart, but was predictable so that he was the one learning to use the system. He disliked a new version of the software that would not behave in a predictable way because it 'learned'. I can't find the link anymore sadly.
If you simply reflect and amplify the behavior of the general population, you can get some pretty ugly results — which is because the general population has some actual ugly tendencies that are being accurately depicted. So the conversation needs to be about how we should actively bias the algorithm to damp out the ugliness, which is already being done by Twitter, FB, Google, etc.
I think we've done the experiment and proven pretty conclusively that a totally unbiased algorithm or information-sharing network is not a good thing for society, unless you want society to be like 4chan. Of course, that immediately opens the can of worms — who gets to decide how to bias the algorithm?
And as an aside, I think the term of bias in algorithms is misused when they are operating as intended. You'd call a poll biased if it implied it was polling group 'x', but did not actually do so due to some inadvertent bias. You wouldn't call it biased if it polled the group it stated it was going to poll and then gave the results - even if those results were not what you personally would want to see. If people are mostly searching for interracial couples when they search for 'white man and white woman' or they are mostly searching for porn when they search for 'black girls' then a biased algorithm would be one that actually returned white couples or non-pornographic results.
Google's search results is really the oldest analog we have for the "machine learning learns racism" quandary, too. This question is just going to get bigger, and its effects much more pernicious.
Although if one actually compares the results for [un]professional hairstyles for work there doesn't seem to be any negative bias; black people are well-represented in both result sets.
The other points look also valid to me, but again, I don't see why Google should be labeled racist when their engine adapts to what news sources publish. Somewhere in the code there could be an association like 3 teens -> gang and a piece of code skimming news sources that reacts to all articles containing the word "gang"; now just mix the two and what you get? Just try searching for "gang" and look at the colors predominance. Search engines aren't easy to write. If I was an Eiffel language expert I would write a book titled "A tour on the Eiffel language" just to prove that. Try searching that phrase without getting submerged by images of the famous tower. Incidentally, a fun search for "a tour on the Eiffel schnitzel kraut" yielded about 90% food related results and only about 5% tower images although it still contained 100% of the terms for both, but having the word Eiffel not immediately following tour could mislead the search engine, so I changed it to "the tour Eiffel on a schnitzel kraut" and results slightly changed with more towers but also more unrelated stuff like Stonehenge.
To me a search engine should return results based on harsh reality if necessary, not church choir dreams; policing search engines results in pursuit of politically correctness would be very dangerous. We won't see that day, nor will our grand-grand-grand-grand-kids, but one day eventually racism will start vanishing from society by itself and search engines will follow, and/or we'll be socially mature enough to ignore anything that could resemble racism to a today's biased mind, giving it even less exposure.
It was: "Sweden".
(Most likely this was on Alta Vista, not Google, but I don't remember for sure.)
(Probably many on HN are too young to remember her... imagine Taylor Swift x10, that was her popularity back then)
Artificial Intelligence has become an alias for "shit I have no knowledge of but would like to pretend to have expertise on that is, at least tangentially, related to a computer"
When I try it in both logged in and non-logged in mode I get reasonably similar results, the differences are mostly in the ordering. When I do the same on Bing I get about 20% or so pornographic images for both 'black girls' and 'white girls'.
When google was just an excellent database index and relevance ranker, then it was ok for its results to simply reflect the underlying data. If the internet is racist, then looking up things on the internet will find racist things.
As time goes on and google purports to give answers and facts and gets more information-ey, they're kinda taking an editorial stance that the results aren't simply relevant webpages, but correct answers in some loose sense. When google tells me the Eiffel tower is 984' tall, and 1063' to the tip, they are endorsing that as The Right Answer. Endorsements come with responsibility.
As Google pushes to give correct answers, and people increasingly expect it, it's an endorsement of search results that comes with responsibility to not just reflect the huge corpus that is the internet.
A Google search for "American women scientists": https://i.imgur.com/o6mxG56.png
Can the media stop doing this thing where they p-hack reality for proof of racist/sexist conspiracies?
In fact, the article is much more supportive of its claims, because it's an excerpt from a book that cites peer-reviewed (I believe) research. The point is that there's a trend, not that some search results are weird.
I'm pretty certain that if she searched for "<blank> girls" -- white girls, asian girls, etc. -- with safe search turned off, a good number of the results would be NSFW, because the internet is 75% porn (and 25% cat videos).
At best you are supporting hers.
She made the case for the query 'black girls' and to pre-empt comments such as yours bolstered her argument by extending it to a general case. So in no way did she 'selective present information to bolster her point', she was as even handed as she could have been (but apparently not even handed enough for some).
Would you have ever tried to search for "black girls" if you hadn't looked at this article? On the other hand, someone who prefers black girls in a sexual way would actually search for them many times. I can imagine someone in advertising searching for random pictures of black girls, but that's a niche case far outnumbered by people searching with sexual intent.
It would be reasonable to assume that any fair, objective algorithm in Google's (or anyone else's) possession that actually optimizes for expected user intent (assuming that they don't have enough privacy-violating information to assume that the user belongs to a smaller and thus less likely minority of black girls) would still bring up a bunch of porn when searching for "black girls", since most people querying for that (except when the searcher has read this article in the last 5 minutes) actually do want to find porn; but they have configured an override purely for PR reasons.
The objective algorithm would show reality as it is, but the current override masks it and tries to show an imagined vision of "reality-as-it-should-be". And that's why the problem is not solvable algorithmically - any algorithm learning from real data and actual user behavior would show the ugly, "undesirable" reality in all kinds of edge cases, and it can only be solved in a non-scalable manner by involving some manual political correctness / good taste / etc censorship.
They are especially slanted on an ideological basis. Try searching "American inventors" - nearly all the people listed are African Americans.
I doubt all of this is inadvertent, and its definitely dangerous, given the power of Google. Today they are undoubtedly influenced by "left-wing" ideology, but what about tomorrow?
> My first encounter with racism in search was in 2009 when I was talking to a friend who causally mentioned one day, “You should see what happens when you Google ‘black girls.’” I did and was stunned.
but then later on:
> Although I focus mainly on the example of black girls to talk about search bias and stereotyping, black girls are not the only girls and women marginalized in search. The results retrieved two years into this study, in 2013, representing Asian girls, Asian Indian girls, Latina girls, white girls, and so forth reveal the ways in which girls’ identities are commercialized, sexualized or made curiosities within the gaze of the search engine. Women and girls do not fare well in Google Search — that is evident.
So it's not racism, it's sexism. Or is it? What would have happened if you searched for black boys, white boys etc. in 2009? I bet it wouldn't have been PG friendly...
This might be legitimate and reasonable, but you probably should be transparent about the fact that you "unbiased" your data by externally imposing a certain view of what the "fair" data should look like.
Ultimately it's a discussion for social scientists or laywers concerned with discrimination, it's not really in the realm of being fixable by engineering or computer science imho.
If it is curating the data, then how it curates matters. In this case, they've curated most of the porn from the first page of results, from the looks of it.
Not sure how that makes me feel.
This is problematic in the same way as facebook’s echo chamber effect, that it can cause a feedback loop that reinforces and heightens bias and division.
Beyond that though, I think it gets gnarly how much we want Google to fix the world by tweaking their algorithms. We obviously want something thoughtful. At the same time I don’t think we should lay all the world’s problems at google’s feet.
Should Google show black women's natural hair in the results for "unprofessional hairstyles" to accurately reflect society's unfair attitudes towards black women's hair? Or should they show actual unprofessional hairstyles like giant mohawks etc.? Or should they show a message saying "No hairstyle is unprofessional. Be who you want to be."?
Developers who train these models or algorithms should be very well aware that the training data is nothing but a download of a sexist, capitalist, racist - society and that bias if is carried forward - irrespective of anything the future isn't going to be clean!
Is introducing an anti-bias at the search engine level the appropriate solution? It seems a bit too much like Western medicine alleviating the symptoms without addressing the root cause. An anti-bias is just masking it.
Also, if the author is a feminist, why she writing for a magazine called "Bitch"?
Now when you do a search for "three white teenagers" you mostly get mug shots of black teens.
This just means that people on the internet who write about black teenagers write about the ones who get arrested. People who create pages about "black girls" create porn content, and most people who search for "black girls" most likely are looking for porn and click on links that talk about porn.
It's amazing how people can see racism everywhere.
I can see the rationale for altering the algorithm for words like "girls".
It's also interesting that, for example, the results for "swedish girls" are much tamer. Even though that's sort of a long standing phrase that was very sexual.
In other words, this is not Google's fault and is mostly just an ugly reflection of our collective problematic values.
While trying to read to article, my browser (Safari on iPhone) was redirected to a scammy looking amazon gift card website. Did anyone else have this problem? I am trying to figure out if it is a problem with Time.com or my phone.
Sometimes I miss the HN of 2011-12, where literally 80% of the comments were about the design of the site, and how it didn't work on Safari/IE/Firefox/Chrome/Lynx.
Not often, though.
"There is no such thing as a moral or an immoral book. Books are well written, or badly written. That is all."
Replace "book" with "algorithm", and you have my sentiments on the issue.
DDG's results for "black girls" is pretty different from Google present-day, though definitely a lot cleaner compared to what the author saw in 2009:
https://duckduckgo.com/?q=black+girls&t=hb&ia=web
(sidenote: it's not clear if the author disabled SafeSearch in her 2009 search or not. SafeSearch was definitely part of Google by then, and I can't imagine "sugaryblackpussy.com" getting past that filter. If you turn DDG's Safe Search to "Moderate" -- "Off" returns the same results as "Strict" -- you will get a lot of very NSFW results. I guess this sidenote raises a new set of issues about the author's methodology but will ignore it for the sake of brevity here.)
The first result is an article titled "Black Girls Only" in Ebony magazine [0]. Which isn't a bad article (or publication). Maybe people would object that the first result is actually about sexualizing black women (albeit positively). The 2nd result is a lot less promising: "Hot Black Girls (45 pics)" at acidcow.com, which has a higher Alexa ranking [1] than Ebony.com (22K vs 63K), but basically looks to be a clunky imageboard. blackgirlscode.com is #6, followed by blackgirlsrun.com (a running club). The rest of the top results are black girl image sites (photobucket, a Facebook group for Big Beautiful Black Girls). The most notable difference between DDG and Google Results, besides what's #1, is that Google results have a lot more news articles in which "black girls" are in the headline (via NPR, nytimes, and theroot). DDG is a lot more sporadic in comparison
DDG for "white girls" is not terribly better in terms of being female-friendly:
https://duckduckgo.com/?q=white+girls
The first result is to the Amazon listing of a book titled "White Girls", by a well regarded New Yorker critic [2]. But the #2 result goes to urbandictionary.com, and #3 hilariously goes to the Wayans Brothers' classic, "White Chicks". There's a bunch of results relating to "White Girl", the mediocre 2016 film (Wikipedia, rogerebert.com, rottentomatoes.com, etc). And then a bunch of results about white females with men of other races.
[a] https://github.com/duckduckgo/zeroclickinfo-goodies/blob/b7a...
[0] http://www.ebony.com/news-views/black-girls-only-503
[1] https://www.alexa.com/siteinfo/acidcow.com
[2] https://www.amazon.com/White-Girls-Hilton-Als/dp/194045025X
I do agree, though, that it is wrong to infer that the algorithms themselves are evil or malicious -- it's possible to be racist without having negative intentions.
But what does it say about a society if it continues to optimize algorithms (and related infrastructure) for one group over another?
Say you make a sensor that is supposed to tell the difference between blue, green and purple, but you only test with one shade of blue and maybe two shades of green, you are going to have trouble actually matching your design goals.
In the case of the face detection system: they didn't specify and/or test is well enough, which can be due to a number of factors, but will most likely lie with the employees of the company that did the development. If they only have the classic 'pasty white guys' to work with, then it's going to be crap at actually doing face detections for all humans. On one hand you could setup a proper test protocol, on the other hand they shouldn't have taken broad or vague terms when developing/presenting the technology. If you don't have a broad selection of faces to test with, you shouldn't claim you have 'face detection', since you merely have 'detection of faces of the people that work on the project plus anyone who looks like them'.
This would be a completely different story if someone writing the face matching code specifically programmed code or wrote configuration data that targets skin tone or geometry of specific groups of people.
Some people would like to extend this type of technological issue into the area of HRM and race/gender-bias in society in general, but that is not a technology-only discussion and hardly something the people involved are qualified to argue about.
Also: >But what does it say about a society if it continues to optimize algorithms (and related infrastructure) for one group over another?
It says that society is imperfect, and that certain levels of xenophobia, bias and true racism exist. Doesn't say much about technology though.
Yeah, but that's what happened here with HP in 2009. I'm not a huge fan of their products these days but I don't think they would intentionally be deceptive here, i.e. I think there's a lot of room to blame incompetence before malice. If HP is a company with very few black employees, this kind of consideration may be completely off their radar. It's super unfortunate, but I don't see the company as evil or maliciously racist, per se (I think we can skip retreading the hiring for diversity debates for now).
> This would be a completely different story if someone writing the face matching code specifically programmed code or wrote configuration data that targets skin tone or geometry of specific groups of people.
Why does it matter? What's the difference between an algorithm that fails to perform because of programmer incompetence, or programmer malice? What's the difference to the end-user if the programmer was plain ignorant of good testing coverage, vs. a programmer who thought "Fuck it, minorities are a minor part of our user base. Not worth the extra engineering effort!"?
Technology is an unavoidable part of the problem. Because it is the technology that allows us the power and freedom to create and apply scalable algorithms to machinery and computers. This automation allows for efficient and reliable decision-making, and we as a society decide where that automation is appropriate and worthwhile, i.e. where human agency is no longer needed.
But technology and its fundamentals are still a key factor. Creating a multi-racial face classifier is fundamentally more work and difficulty than one trained for just one race. The math and physics are unavoidable. And every engineered system and product has to make tradeoffs between production cost and feature set.
In the case of the light-skin-optimized HP web cam, I think it's important, and fine, to call it "racist" -- a black HP customer will have an inferior experience fundamentally because he is racially black. But this isn't just a way to quickly assign blame. Recognizing that tech is fundamentally limited is the first step in understanding that systemic racism (e.g., all the decisions that led to the "racist" camera) could be a contributing factor to the camera's substandard performance.
Much harder to get to that thinking if we have a mentality of, "how could the computer be wrong/flawed?"
Why are so many white men so angry?
White men must be stopped: The very future of mankind depends [...] The future of life on the planet depends on bringing the 500-year rampage of the white man to a halt. For five centuries his ever more destructive weaponry ...
My first three DDG-results on searching 'white men'.
I shall now proceed to get over it. Suggest others do likewise.
Oh.
You are allowed to say fuck on the internet.
Also, does it really count if you skip out one letter? I never understood this. You're still all but saying the word. Nobody is being protected here lol. A kid can cycle through the five vowels in about half a second to figure out what the word is, and one of them is also kinda a swear - pissy.
EDIT: "n 2012, I wrote an article for Bitch Magazine" now I'm very confused. It's ok if it's a proper noun, I guess?
Meh, likely the answer is unknown, mostly I'm just commenting
https://www.usatoday.com/story/tech/news/2016/06/30/google-d...
> Women made up 31% of Google employees in 2015, up one percentage point since 2014, according to statistics released by the Internet giant on Thursday. One in five technical hires were women in 2015, raising the number of women in technical roles to 19% from 18% in 2014 and 17% in 2013. In 2015, women held nearly a quarter of leadership posts at Google, up from 22% in 2014 and 21% in 2013.
> Google says it's also hiring more black and Hispanic workers: 4% of hires in 2015 were black and 5% were Hispanic. Hispanic employees in technical roles increased to 3% from 2%. But the increased hiring did not budge the overall percentage of underrepresented minorities in the Google workforce as total hiring rose, with Hispanics making up 3% of the work force and African Americans 2%.
Given that women are half the population, and African Americans and Hispanics each make up about 12%, we can state categorically that Google's hiring doesn't match the general US population.
But if you account for demographics within the industry do the numbers remain biased?
For example: men account for less than 10% of nurses so if you were to go to a hospital with 50% male nurses they would be way overrepresented. Statistically speaking of course, I'm in no way trying to give the impression we need to do something about the lack of diversity within the nursing profession.
--edit--
...or maybe we need to do something about the lack of diversity in the nursing profession? Don't actually know how this ball rolls?
But that's not what this is really about, it's about racial justice. It's about redistributing wealth and power from whites to everyone else in an attempt to create more equality across the board, not about "fair" hiring practices.
Seems ironic.