Amazon scraps secret AI recruiting tool that showed bias against women
reuters.com
reuters.com
At start the AI is like a baby, it doesn't know anything or have any opinions. By teaching it using a set of data, in this case a set of resumes and the outcome then it can form an opinion.
The AI becoming biased tells that the "teacher" was biased also. So actually Amazon's recruiting process seems to be a mess with the technical skills on the resume amounting to zilch, gender and the aggressiveness of the resume's language being the most important (because that's how the human recruiters actually hired people when someone put a resume).
The number of women and men in the data set shouldn't matter (algorithms learn that even if there was 1 woman, if she was hired then it will be positive about future woman candidates). What matters is the rejection rate which it learned from the data.. The hiring process is inherently biased against women.
Technically one could say that the AI was successful because it emulated the current Amazon hiring status.
This is incorrect. The key thing to keep in mind is that they are not just predicting who is a good candidate, they are also ranking by the certainty of their prediction.
Lower numbers of female candidates could plausibly lead to lower certainty for the prediction model as it would have less data on those people. I've never trained a model on resumes, but I definitely often see this "lower certainty on minorites" thing for models I do train.
The lower certainty would in turn lead to lower rankings for women even without any bias in the data.
Now, I'm not saying that Amazon's data isn't biased. I would not be surprised if it were. I'm just saying we should be careful in understanding what is evidence of bias and what is not.
I don't think that's true. "No bias" means that gender is irrelevant (i.e. its correlation with outcome is 0%). Therefore the system shouldn't even take it into account - it would evaluate both men and women just by other criteria (experience, technical skills, etc), and it would have equal amounts of data for both (because it wouldn't even see them as different).
You need bias to even separate the dataset into distinct categories.
False. If we're talking about the technical statistical definition, bias means systematic deviation from the underlying truth in the data -- see this article by Chris Stucchio with some images for clarification:
https://jacobitemag.com/2017/08/29/a-i-bias-doesnt-mean-what...
"In statistics, a “bias” is defined as a statistical predictor which makes errors that all have the same direction. A separate term — “variance” — is used to describe errors without any particular direction.
It’s important to distinguish bias (making errors with a common direction) from variance which is simply inaccuracy with no particular direction."
My point was that you should consider the meaning of the word under which the post you're replying to is correct, especially given that the author was claiming specific domain experience.
> The lower certainty would in turn lead to lower rankings for women even without any bias in the data.
your post said:
> If we're talking about the technical statistical definition, bias means systematic deviation from the underlying truth in the data
So I think my interpretation is correct, even though it's not "the technically statistically correct usage". You were referring to the bias of the algorithm (i.e. the mean divergence from the mean in the data), whereas we were referring to the "hiring bias" evident in the data. In fact, your "bias" was mentioned as "lower rankings for women" - i.e. "the algorithm would have (statistical) bias even without (sexist) bias in the data" and I was replying that I think that's false.
I'm not trying to split hairs (or argue), as much as further clarify the difference between (the common definition of) human bias and that of statistical bias.
Computers are very bad at actually discriminating against people, they will pick up a possible bias in a statistical dataset (ie, <protected class> uses certain sentence structure and is statistically less likely to get or keep the job).
Sometimes computers also pick up on statistical truths that we don't like, ie, you assign a ML to classify how likely someone is to pay back their loan and it picks up on poor people and bad neighborhoods, disproportionately affecting people of color or low income households. In theory there is nothing wrong with the data, after all, these are the people who are least likely to pay back a loan, but our moral framework usually classifies this as bad and discriminatory.
Machine Learning (AI) doesn't have moral frameworks and doesn't know what the truth is. The answers it can give us may not be answers we like or want or should have.
on a side note; human bias is usually not that different since the brain can be simplified as a bayesian filter; there are predictions on the present based on past experience, reevaluation of past experience based on current experience and prediction of future experience based on past and current experience. It's a simplification but usually most human bias is based on one of these, either explicitly social (bad experience with certain classes of people) or implicitly (tribalism).
I agree with everything else in your post, but just wanted to note that while this is true to some extent, the brain is much less rational than a pure Bayesian inference system; there are a lot of baked in heuristics designed to short-circuit the collection of data that would be required to make high-quality Bayesian inferences.
This is why excessive stereotyping and tribalism are a fundamental human trait; a pure Bayesian system wouldn't jump to conclusions as quickly as humans do, nor would it refuse to change its mind from those hastily-formed opinions.
I think I'd make the claim a bit less strongly -- we don't know if there is statistical bias or non-statistical/"gender bias" in the data; both are possible based on what we know.
However exploring the statistical bias possibility, the simple way this could happen is if the data have properties like:
1. For whatever reason, fewer women than men choose to be software engineers 2. For whatever reason, the women that choose to be software engineers are better at it than men
(Note I'm just using hypotheticals here, I'm not making claims about the truth of these, or whether it's gender bias that they are true/false).
Depending on how you've set up your classifier, you could effectively be asking "does this candidate look like software engineers I've already hired"? If so, under the first case, you'd correctly answer "not much". Or you could easily go the other way and "bias" towards women if you fit your model to the top 1% where women are better than men, in our hypothetical dataset.
This would result in "gender bias" in the results, but there's no statistical bias here, since your algorithm is correctly answering the question you asked. It's probably the wrong question though!
Figuring out if/when you're asking the right question is quite difficult, and as the sibling comment rightly pointed out, sometimes (e.g. insurance pricing) the strictly "correct" result (from a business/financial point of view) ends up being considered discriminatory under the moral lens.
This is why we can't just wash our hands of these problems and let a machine do it; until we're comfortable that machines understand our morality, they will do that part wrong.
gp: "The number of women and men in the data set shouldn't matter (algorithms learn that even if there was 1 woman, if she was hired then it will be positive about future woman candidates)."
This is false for typical models.
Seems challenging since much of AI, especially classification, is essentially a discrimination algorithm.
This isn't an insurmountable problem, but does require extra work then just "encode, throw it in and see what happens".
Amazon only scrapped the original team, but formed a new one in which diversity is a goal for the output.
Then this is completeley useless. You want this "AI" to discriminate based on a number of things. That's the whole point. You want to find people that can work for you. If a specific school or title is a bad indicator (based on what you hired now), then it just is that.
Machine learning generally doesn't have any prior opinions about things and will learn any possible correlation in the data.
It could for example discover that certain words or sentence structures used in the resume are more likely associated with bad candidates. Later you find out that <protected class> has a huge amount of people that use these certain words/structures while most other people don't.
And now the AI discriminates against them.
ML will pick up on any possible signal including noise.
This is not true.
Probabilistic-ly speaking, if we are computing P(hiring | gender); Lower certainty means there is a high variance in prior over women. However, over a large dataset, the "score" would almost certainly be equal to the mean of the distribution, and be independent of the variance.
In simpler words, if there was a frequency diagram of scores for each gender (most likely bell curves), then only the peak of the bell curve would matter. The flatness / thinness of the curve would be completely irrelevant to the final score. The peak is the mean, and the flatness is the uncertainty. Only the mean matters.
Unless they presented lots of unqualified resumes of people not in tech as part of the training, which seems like something someone might think reasonable. Then, the model would (correctly) determine that very few people coming from women's colleges are CS majors, and penalize them. However, I'd still expect a well built model to adjust so that if someone was a CS major, it would adjust accordingly and get rid of any default penalty for being at a particular college.
If the whole thing was hand-engineered, then of course all bets are off. It's hard to deal well with unbalanced classes, and as you mentioned, without knowing what their data looks like we can only speculate on what really happened.
But I will say this: this is not a general failure of ML, these sorts of problems can be avoided if you know what you're doing, unless your data is garbage.
That's exactly the issue we are talking about here. Woman's colleges would have less training data so they would get updated less. For many classes of models (such as neural networks with weight decay or common initialization schemes) this would encourage the model to be more "neutral" about women and assign predictions closer to 0.5 for them. This might not affect the overall accuracy for women (as it might not influence whether or not they go above or below 0.5), but it would cause the predictions for women to be less confident and thus have a lower ranking (closer to the middle of the pack as opposed to the top).
A class imbalance doesn't change that: if there's no gradient to follow, then the class in question will be strictly ignored unless you've somehow forced the model to pay attention to it in the architecture (which is possible, but would take some specific effort).
What I'm suggesting is that it's likely that they did (perhaps accidentally?) let a loss gradient between the classes slip into their data, because they had a whole bunch of female resumes that were from people not in tech. That would explain the difference, whereas at least with NNs, simply having imbalanced classes would not.
Specifically, the gp is pointing out that typical approaches will not pay attention to a feature that doesn't have many data points associated with it. In other words, if it hasn't seen very much of something then it won't "form an opinion" about it and thus the other features will be the ones determining the output value.
Additionally, the gp also points out that if you were to accidentally do something (say, feed in non-tech resumes) that exposed your model to an otherwise missing feature (say, predominantly female hobbies or women's colleges or whatever) in a negative light, then you will have (inadvertently) directly trained your model to treat those features as negatives.
Of course, another (hacky) hypothetical (noted elsewhere in this thread) would be to use "resume + hire/pass" as your data set. In that case, your model would simply try to emulate your current hiring practices. If your current practices exhibit a notable bias towards a given feature, then your model presumably will too.
Unless Amazon is willing to accept a) another pool of data or b) that the data will yield bias and apply a correction, the AI is almost guaranteed to be taught the bias.
This is why AI is so confusing. All "AI" does is rapidly accelerate human decisions by not involving them, so that speed and consistency are guaranteed. They are not replacements for human decision making, they are replacements for human decision making at scale.
If we can't figure out how to do unbiased interviews at the individual level, then AI will never solve this problem. Anyone that tells you otherwise is selling you snake oil.
I wonder to what extent people want to solve it and perhaps more importantly whether or not it can be solved at all...
(Serious question. Not intended as snark. Genuinely wondering if I'm missing some deeper current in your post?)
If my supposition is correct then the other parameters are at fault here from which gender and language used stick out.
Another supposition I'm going to make is that they even removed the gender from the data set so that AI didn't know it, but cross-referencing still showed "faulty" results due to hidden bias that the AI can pick up, like language used.
The data set will also have skewed heavily against people named "David". Probably only ~1% of the successful applicants.
Would you also expect the machine to be biased against candidates named David?
Hiring practices as expressed in the data get picked up by the machine and applied accordingly. As such, David is predicted to be a better hire than Denise.
This is not about "David" vs. "Denise", but how the machine learning process will aggregate and classify names. David and David-like names will come out on top while obscure names it has no idea how to deal with (0/0 historically) will probably be given no weighting at all.
Sorry "Daud!" Our algorithm says David is better.
This is most common with binary problems.
The NBA wants good basketball players. If they happen to be white, I imagine they'd draft them with equal enthusiasm as any other player. So no, it isn't.
AI is designed to get the same results as a human. How it gets to those results is often very, very different. I'm having trouble finding it, but there was an article a while back trying to do focus tracking between humans and computers for image recognition. What they found was that even when computers were relatively consistent with humans in results, they often focused on different parts of the image and relied on different correlations.
That doesn't mean that Amazon isn't biased. I mean, let's be honest, it probably is; there's no way a company this large is going to be able to perfectly filter or train every employee and on average tech bias trends against women. BUT, the point is that even if Amazon were to completely eliminate bias from every single hiring decision it used in its training data, an AI still might introduce a racial or gendered bias on its own if the data were skewed or had an unseen correlation that researchers didn't intend.
Note the title is "Amazon scraps secret AI recruiting tool that showed bias against women" not "Amazon scraps secret AI recruiting tool because it showed bias against women". But I guess the real title is less clickbaity - "Amazon scraps secret AI recruiting tool because it didn't work".
A far more reasonable way would be to take resumes of people who were hired and train the model based on their performance. For example, you could rate resumes of people who promptly quit or got fired as less attractive than resumes of people who stayed with the company for a long time. You could also factor in performance reviews.
It is entirely possible that such model would search for people who aren't usually preferred. E.g. if your recruiters are biased against Ph.D.'s, but you have some Ph.D.'s and they're highly productive, the algorithm could pick this up and rate Ph.D. resumes higher.
Now, you still wouldn't know anything about people whom you didn't hire. This means there is some possibility your employees are not representative of general population and your model would be biased because of that.
Let's say your recruiters are biased against Ph.D.'s and so they undergo extra scrutiny. You only hire candidates with a doctoral degree if they are amazing. This means within your company a doctoral degree is a good predictor of success, but in the world at large it could be a bad criteria to use.
I want to see reports of average tenure and time between promotions by gender. I suspect that the reason we don't see those published is that the numbers are damning.
It's also not hard to make the pay gap 1-2% just like it's not hard to make it 25% (both values are valid). Statistics is a fun field. Don't trust statistics you didn't fake yourself.
Amazon could easily cook the numbers to get to 1-2%, I doubt anyone checked the process of determining that number if it's unbiased and fair and accounts for other factors or not.
Aggressive behavior is considered admirable in men, and deplorable in women. Many women I know have noted comments in their performance reviews about their behavior - various words that can all be distilled to "bitchy".
But is that what we see in real life?
I don't have data or sources at hand, but I'd bet top dollar that F-M ratio among employees is much more lopsided in male favor among founders[0].
[0] Not using the word CEO, because that can be appointed for somewhat arbitrary reasons.
If you had a way to accurately predict that some company would systematically donwrate you and eventually fire you or force you to quit, would you want to interview there? If you were a recruiter in that company and could accurately predict the same, would it be ethical for you to hire the candidate anyway?
This is not to say that I approve of blindly trusting AI to filter candidates, but the overall issue isn't nearly as simple as many comments here make it out to be.
Its an interesting questions. On one hand, a practical person could argue: "Well, this is what my company looks like, and these are the types of people who fit with our culture and make it, so be it. Find me these types of candidates."
VS
"I don't like the way may company culture looks, I would rather it was more diverse. This mono-culture is potentially leaving money on the table from not being diverse enough. I'm going to take my current employees, chart their career path, composite them (maybe), tweak some of the ugly race and gender stats for those who were promoted, and feed this to my hiring algorithm."
Thatd be great, but in this case (as in most ML cases) the idea is not "follow this known, tedious process" but instead "we have inputs and results but dont know the rules that connect them, can you figure out the rules?"
> this is what my company looks like
In tech hiring, no one wants the team they have...they want more people but without regrets (including regretting the cost)
It's a fine strategy if all you're trying to do is cost-cut and replace the people that currently make these decisions (without changing the decisions).
I agree that most people with ML experience would want to do better, and could think of ways to do so with the right data, but if all the data that's available is "resume + hire/no-hire", then this might be the best they could do (or at least the limit of their assignment).
Many companies are fine with false negatives in their hiring process. Better to pass on a good candidate than hire a bad one.
That doesn’t follow.
I'll don my flack jacket for this one, but based on population statistics I believe a statistically significant number of women have children. A plausible hypothesis is that a typical female candidate is at a 9 month disadvantage against male employees and that that is a statistically significant effect detected by this Amazon tool.
Now, the article says that the results of the tool were 'nearly random', so that probably wasn't the issue. But just because the result of a machine learning process is biased does not indicate that the teacher is biased. It indicates that the data is biased, and bias always has a chance to be linked to real-world phenomenon.
Obviously I don't have much specific insight, so maybe there is a culture where they don't use leave entitlements. But if there are indicators that identify a sub-population taking a potentially 20 week contiguous break it is entirely plausible that it would turn up as a statistically significant effect in an objective performance measure. All else being equal, then a machine learning model could pick up on that.
The point isn't that it is the be-all and end all, just that the model might be picking up on something real. There are actual differences in the physical world.
Pattern recognition will learn any biases in your training data. An intelligent enough* being does much more than pattern recognition -- intelligent beings have concepts of ethics, social responsibility, value systems, dreams, ideals, and is able to know what to look for and what to ignore in the process of learning.
A dumb pattern recognition algorithm aims to maximize its correctness. Gradient descent does exactly that. It wants to be correct as much of the time as possible. An intelligent enough being, on the other hand, has at least an idea of de-prioritizing mathematical correctness and putting ethics first.
Deep learning in its current state is emphatically NOT what I would call "intelligence" in that respect.
Google had a big media blooper when their algorithm mistakenly recognized a black person as a gorilla [0]. The fundamental problem here is that state-of-the-art machine learning is not intelligent enough. It sees dark-colored pixels with a face and goes "oh, gorilla". Nothing else. The very fact that people were offended by that is a sign that people are truly intelligent. The fact that the algorithm didn't even know it was offending people is a sign that the algorithm is stupid. Emotions, the ability to be offended, and the ability to understand what offends others, are all products of true intelligence.
If you used today's state-of-the-art machine learning, fed it real data from today's world, and asked it to classify them into [good people, criminals, terrorists], you would result in an algorithm that labels all black people as criminals and all people with black hair and beards as terrorists. The algorithm might even be the most mathematically correct model. The very fact that you (I sincerely hope) cringe at the above is a sign that YOU are intelligent and this algorithm is stupid.
*People are overall intelligent, and some people behave more intelligently than others. There are members of society that do unintelligent things, like stereotyping, over-generalization, and prejudice, and others who don't.
[0] https://www.theverge.com/2018/1/12/16882408/google-racist-go...
so, knowledge now is allegedly possession of the future, rather than possession of the past.
This is because the future and past are structurally the same thing in these models. Each could be missing, but re-creatable links.
Also, conflicting correlations can be shown all the time. if almost any correlation can be shown to be real, what's true? How do we deal with conflicting correlations?
For the black man = gorilla problem, an untaught human, a small child for instance, can easily make the same mistake. Especially if he has seen few black people. And well educated adults can also make the mistake initially, even if they hate to admit it.
However, in the last case, a second pattern recognition happen, one that matches the result of the image classifier with social rules. And it turns out that mixing black men and gorillas is a clear anti-pattern and anything that isn't certain is incorrect.
Unlike us, computer image classifiers typically aren't taught social rules, so like a small child, they will tell things without filter. It will probably change in the future for public facing AIs.
Not stereotyping is not a mark of intelligence, it is a mark of a certain type of education. And I don't see why it couldn't be done with the usual machine learning techniques.
I claim it isn't just social rules -- part of that is empathy, which is a manifestation of intelligence that I think is beyond pattern matching.
If a white person were mislabeled as a cat, it would be a cute funny mistake. Labeling people as dogs, not so much. Gorillas, even worse. Despite that gorillas are more intelligent and empathetic than cats. Oh, and bodybuilder white celebrity boxing champion as a gorilla, may actually be okay. The same guy as a dog, no. It makes no sense to a logic-based algorithm. But humans "get it".
A human gets it because they could imagine the mistake happening against them, with absolutely zero prior training data. You don't need to have seen 500 examples of people being called gorillas, cats, dogs, turtles and whatever else.
If you want to say that a hundred pattern recognition algorithms working together in a delicate way might manifest intelligence, I think that is possible. But the point is one task-specific lowly pattern recognition algorithm, which is today's state of the art, is pretty stupid.
That's just one function. That's not the entirety of what the brain (and body) does.
> If you consider pattern matching unintelligent,
What do you think pattern matching IS? Round ball round hole does not require intelligence. It requires physics. The convoluted rube goldberg meat machine what we use to do it, doesn't change what it is. Making the choice of will and approximations, are more signs of intelligence, imo.
> Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs, the people said. With the technology returning results almost at random, Amazon shut down the project, they said.
Granted an article isn't going to get as much attention without an attractive headline but that seems a far more likely reason to have an AI based recruiting recommendation scrapped. The discovery of a negative weight associated with "women's" or graduates of two unnamed women's colleges is notable but if it's tossing out results "almost at random" then...well there seems to be bigger problems?
Men and women are pitted against each other.
Due to the way the media has evolved people consume their own biases and most often just read the headlines.
It's amazon, I can't imagine how many millions went into something like that. We'll almost certainly not get a postmortem but it's definitely intriguing.
Learning history will teach you more about things today than any news source.
Now having known limited capabilities isn't great. But those can and will be worked on. Unknown / unexpected biases wont, making finding them important.
I am a frequent interviewer for engineering roles at Amazon. As part of the interview training and other forums, we often discuss the importance of removing bias, looking out for unconscious bias, and so on. The recruiters I know at Amazon all take reaching out to historically under-represented groups seriously.
I don't know anything about the system described in the article (even that we had such a system), but if it was introducing bias I'm glad it's being shelved. Hopefully this article doesn't discourage people from applying to work at Amazon - I've found it a good place to work.
To say something about the AI/ML aspect of the article: I think as engineers our instinct is "Here's some data that's been classified for me, I can use ML/AI on it!" without thinking through all that follows, including doing quality assurance. I think a lot of focus in ML (at least in what I've read) has been on generating models, and not nearly enough focus has been on generating models that interpretable (i.e., give a reason along with a classification).
Technical talent is both expensive and a rare commodity for tech companies. The non-male engineers I've worked with have always been exceedingly competent, smart, and their differing perspectives invaluable. If there was an untapped market of engineers you'd better believe every tech company would be taking advantage of it.
There is a shortage of male therapists and kindergarten teachers: is that because males aren’t being hired or because there are fewer of them in existence?
There does not exist some magical undiscovered pool of talented female engineers that are being turned away by biased recruiters. It's hard enough to find any sort of talented engineers regardless of other factors. Shit, it's not uncommon to recruit from other countries and cover relocation costs these days.
How exactly do you hire for female candidates without discriminating against the "overwhelming" body of male applicants? I would be really interested to know how this goes on behind the scenes: do you have open positions but cherry pick female candidates while disregarding male candidates from the get-go?
How is this not sexism? Replace female with male in the above quote and tell me it isnt sexist.
So yes, sexism is not symmetric. Same with race.
Please note there is no active discrimination against men, but preference in some cases would be to hire a women. Gender is the only criteria for diversity in here.
Would a preference for hiring men be discrimination against women?
Progamming is also manufactured to be a "male" profession. It used to be that researchers were males, and programming was considered womans work, like doing data entry. When companies like IBM discovered that the best programmers tended to be anti-social males, and that females tended to leave careers earlier to start a family, big tech companies focused all their recruitment on males until progamming became a "male" job. It's similar to how light beer used to be considered a womans drink, but beer companies started running ads showing football players and other masculine figures drinking it, and now it's acceptable for anyone to drink.
Amazon (and a lot of other big non-diverse companies) are therefore hoping what is actually happening is that they are the women's first choice already, but turn them down for some reason, and that they would not have to change a thing about their work process to attract more women, except to start seeing them.
It's obvious why companies like thinking that way, and it's possible to some extent it's true. However, the fact is, if they're playing like this it's a zero-sum game and it is not actually going to improve the diversity numbers.
On my part, I've been wondering. If all these companies want to know where the women who can code are... why not just ask them? Why do you never see a "Female Programmer's Career Survey" with questions such as "On average, over the last ten years, how long have you been looking for a job?" "Would you accept a new job for a $5k raise?" "Have you ever dropped out of a hiring process because of sexism?" Take it out there in the open. Ask the real questions.
Single-generation changes in behavior aren't genetic. They're social.
If tomorrow we say that you have to do 30 chinups to be a waitress, and the job will involve regular fistfights then we count the number of waitresses by gender and say "it must be cultural", we're kind of missing the point. Or if we say "OK now waitresses make 200k and are respected" and watch the numbers shift.
What, pray tell, has changed that made the job more attractive to men, and less attractive to women? You need to be able to answer that question if you're going to make a causal assertion.
I first learned to program 35 years ago. It wasn't fundamentally different then. Hell, we still use programming languages that were in wide use 35 years ago, like C and the unix shell. The kind of thinking required hasn't changed.
So, based on my 35 years of experience, the conditions and the job are basically the same. So again, I challenge you - how is the job different now?
how is the job different? The pay is much higher and so are the entry requirements, that's the biggest difference. You don't get assigned to program punch cards as part of your secretarial role, you have to actively get educated and good to choose it as a career.
I think there's far more places now expecting crazy hours but that's anecdotal I don't have numbers on it. But the languages are different (mostly), the tooling is different, the deployment is different, the scale is different.
Not the parent and I wasn't around either, but I think accessibility of education counters your point, not supports it.
More egalitarianism should in theory be more favorable to women.
Back then I imagine it was much harder to program without access to university computers and education materials.
More women get higher education than men compared to 35 years ago.
More incentives (monetary and otherwise), combined with lower barriers to entry should also be favoring the supposedly disadvantaged.
And yet the drop in F-M ratio since late 1980s has not been overcome last time I checked.
If you assume that there are underlying differences in interests and aptitude, more egalitarianism allows these differences to be expressed more since women are more free to eg. choose a career working with people, like medicine or law. http://www.thejournal.ie/gender-equality-countries-stem-girl... It also raises the bar for inherent aptitude to get into/(the top of) a career, since you're competing against a much wider pool of talent.
The point I was making to the parent was that his point "cross-generation drop in ratio proves it's cultural" doesn't hold up, because there have been many changes across those generations, you're comparing apples to oranges.
That's not true. Firstly, it's more like 40-50 years ago.
Secondly, there are far more women doing software development, but the gender ratio is dramatically different.
Thirdly, that's because male interest exploded with the advent of personal computing the 80s.
Lastly, "programming" as a profession used to be regarded as an offshot of secretarial work, which was dominated by women.
The facts around women in STEM are polluted with a lot of bizarre narratives.
"'programming' as a profession used to be regarded as an offshoot of secretarial work, which was dominated by women". Which begs the question of why women dominated secretarial work (and still do), while as programming became a more respected and better paying profession, it became male-dominated.
It really doesn't. Unless you seriously think punch card programming is the same as modern programming, or that the fact that only women were secretaries that did programming somehow provides data on the relative strengths and inclinations of women and men for programming work at that time.
Look, it's clear that you have no idea of the breadth and depth of data available on this subject, and a trite "sexism/oppression" narrative explains hardly any of it. For instance, the fact that as a nation becomes more egalitarian, the gender disparities in STEM increase, ie. Nordic countries have worse gender disparities than here, despite having less sexism, and oppressive countries like Iran actually have gender parity in STEM fields.
If you want to actually learn about this subject, I suggest reading: https://www.frontiersin.org/articles/10.3389/fpsyg.2015.0018...
The fact is, there's good evidence that women are naturally less interested in STEM-like fields due to a well known psychological attitude on things vs. people. That attitude explains facts like why medicine and law have achieved approximate gender parity overall, but surgery is still dominated by men, pediatrics and family law is dominated by women.
I do find it interesting and noteworthy that gender disparities have grown in STEM while shrinking in other fields. But I believe my explanation accounts for that - that STEM has become more prestigious, which draws men, which forces out women.
The "well known psychological attitude" is begging the question, which seems par for the course on responses here. Is this psychological attitude biological, or social? And if it's biological, how do we explain significant changes in professional proportions that have happened over a mere one or two generations? It seems like a very poor explanation for what you're asserting, contradicting your own stated facts.
If it's social, however, we're back to my explanation - as the prestige of formerly female-dominated careers rises, they become more attractive to men, to the point where men dominate them. It's a much simpler explanation, with no contradictions.
Because you're throwing out wild, unsupported speculation to salvage your narrative, and the original post of yours to which I replied had at least 4 elementary factual errors.
> But I believe my explanation accounts for that - that STEM has become more prestigious, which draws men, which forces out women.
That's not an explanation at all. Why would prestige drive away women? Just because there are men there? Or you think men drawn to prestige don't want women around? Or you think men just flood into any field that has some form of prestige thus drowning out women? So then why aren't the careers they left suddenly dominated by women because all the men left for more prestige? And where are all these men coming from since we have rough equal numbers of men and women? Why are janitors and dangerous jobs dominated by men since those aren't prestigious?
The fact that you think this explains anything or is free of contradictions is frankly bizarre, and just reinforces my point that if you're really interested in this field, you need to more read more and speculate less.
> The "well known psychological attitude" is begging the question, which seems par for the course on responses here. Is this psychological attitude biological, or social?
Likely both, since there's plenty of evidence of things vs. people in toddlers, and this innate preference no doubt gets reinforced and magnified.
In the end, your scoffing at the original poster and "subtly" implying that he's sexist for a remark that is actually well grounded in facts is exactly the problem with debating people on this subject.
Yes, there is sexism in STEM, just like there is in most other fields, but sexism didn't keep women out of medicine or law, they just pushed through and staked their claim. The fact that women haven't done this for STEM which is far less of an old boys' club already suggests something else is at play, and the fact that the same trends are seen across disparate cultures already suggests strongly there's a universal component.
I do think there's a universal component, though, as sexism is seen across virtually all cultures.
Have you noticed how this is inconsistent with your prestige argument?
You're equivocating. You know very well that the type of sexism that kept women from working in virtually all professions, including law and medicine, is not the type of sexism we're discussing now.
Is it competition that is forcing women out, or men ?
> as the prestige of formerly female-dominated careers rises, they become more attractive to men, to the point where men dominate them
What do you mean by dominated, is it the number of people, or is it something else ?
Are you trying to say that once a career path becomes female dominated, men should stay out ?
One very likely explanation is that many women who wanted to be independently financially successful had few choices other than tech back decades ago. Now they have many other choices.
But are they really ?, do you have any data/reference to back it up, I was under the assumption that there are more people, both men and women, working as programmers, than 30 years ago.
Since you're presumably not a woman, and they are, they object to your seeming to be taking it upon yourself to tell them what a woman is.
It's not unlikely that most women don't want to be in software. Most men don't want to be in software! But hundreds of thousands of women are, in fact, working as programmers. Whenever tech news talks about hiring women, presumably, they're talking about hiring these women. Why pretend they don't exist?
That idea is complete poison. In a debate where you're both completely uninformed, the anecdotal evidence of experience is relevant. In any discussion where reason, numbers, research are involved, the gender/race of the person making the argument is irrelevant to the argument. This current idea that only women can talk about female issues, only X about X is pre-enlightenment tribalism.
It presupposes conflict between the tribes too. If in my philosophy I believe that through reason and evidence I can understand your point of view and experiences,there's the possibility of agreement. If you really believe that I can never talk about issues affecting your tribe, or arrive at reason to an underlying truth, there's no point in talking. We might as well just fight to see whose tribe can impose power on whose.
Arguments being inflammatory does not make them invalid. It is fine to have conversations that include inflammatory arguments, if they are made politely, which parent did. But ignoring that the argument is inflammatory, and/or that the group the argument refers to are in fact intelligent people involved in the discussion that may have themselves an opinion, is lacking in empathy. Rhetoric is founded in empathy; that is why it is an art, and not merely a technique.
woah what? no that is not at all what he's saying. If someone tells me "most american men like the NFL" and I don't like the NFL, I would be insane to take that as someone telling me "you're not a real man" and think they're trying to "tell me what a real man is." I can see how someone who is perpetually trying to be a victim might take such a hardline stance, though.
The conversational equivalent using the NFL example would go something like this:
"Why are there no Americans at my favorite chess forum?"
"Americans like the NFL. They're just more into brute force and camaraderie, especially American men. Chess can't really appeal to them. I mean, back in the Neolithic, a modern day chess grandmaster, if he managed to not burn in the sun and see an angry cougar 15 meters away, would probably have died. American men are just closer to their nature. I have this cousin who's American, I have attempted for years to get him to play chess and no dice. Not during the season, anyway."
"I mean sure but there are American chess clubs? Some Americans like chess? Surely the cougar thing is not relevant to my original scope?"
Now, you're a chess loving American. You're not at the first poster's forum because you can't be on all the forums in the world, and it's true love of chess is not exactly common in the US. However, how would you feel reading the description of how apparently literally everyone else in America loves the NFL? Would you feel proud to be American? Would you feel like the poster who made the NFL comment is likely to be an American man? Would you be more, or less likely to read his further arguments on different or related topics, for example, his opinion on American politics?
While I'm sure that there are plenty of women who don't mind those arguments being made, and I think to an extent this forum would select for those women anyway, it remains an argument that is, in essence, 1) prone to being misinterpreted and 2) a little blind to your audience. (the subject is chess - you are talking to a chess fan - for all you know the chess fan also loves the NFL)
And those arguments are less effective than other arguments, for example, the kinds that 1) don't make assumptions of their audience and 2) are closely related to the topic.
(I realise this digression itself is totally off-topic and I'm sorry. I'm not interested in monologuing or haranguing the poster or anything. I hope it, and the little NFL parable, showed you and the people who usually make that argument without thinking about who reads it, a slightly different side, and that you may consider it if in the future you should have longer debates on relevant subjects with people whose opinion and background you don't know a priori. And thank you to anyone who read it.)
Let me play the same game: I know a black guy who is president. Does that mean all black men are political? What about the one who just wants to play sports. Will we now call him an athletic black man, instead of just a black man...
You can generate endless arguments like this depending on your choice of anecdote and label.
Dirlewanger: group X have property Y arandr0x: members of group X without property Y might be offended by "group X have property Y" courir: as a member of group X without property Y I'm not offended by "most members of group X have property Y" arandr0x: <a reiteration of the previous comment> ramblerman: <rhetorically> does saying a member of group X has property Y mean all members of group X have property Y?
I think you (ramblerman) have logically inverted the main claim, which is why it doesn't seem to make sense. Behaviour in line with arandr0x' comment seems perfectly reasonable to me - few people take well to poorly fitting generalisations.
But nobody has ever said that. They say that women are more likely to be.
My father got me into programming, and my colleagues who are women also have a backstory with someone supporting their interest in development.
Just teach kids to code, boy or girl. Not all of them will like it, not all of them will be good at it. But I think a lot more girls would be into it and good at it if they were introduced to it before college.
Tailor it to the kid's interests. My first programs were more socially oriented. When I was 5-6 years old, all my programs were made-up conversations with the computer, where your answers were stored and parroted back by the computer to show that it was "listening". Maybe a boy would have been less into that and more into something mathier, like LOGO instead of BASIC, but it was what was interesting to me at that age. The computer was a form of imaginary friend for an introverted kid like me.
One issue that keeps happening is an over-emphasis on CS-related questions. There are many great engineers I've worked with who didn't do a CS degree, and even though they are brilliant thinkers and talented engineers, too many times the interview question is "solve this problem using <pet CS 101 lesson, like red-black trees>".
And the number of people who are hired who can barely communicate effectively is still shocking. Very few interview questions focus on communication outside the technical realm.
So you can argue there is a bias in recruiting, simply because different people have different criteria for what the best traits/skills to look for is - even though everybody has the same goal, hiring the "best".
I'd also caution about taking Reuters too seriously though. Seems that they've only focused on the gender issue, but this is the money quote:
> With the technology returning results almost at random, Amazon shut down the project, they said.
These interview processes exist sure, and I personally find them idiotic, but is there evidence that they disqualify women more than men?
After we moved to a logic-based test, we were able to hire several more women from interesting disciplines including psychology, math, and biology. The tests involved technical problems written in a general way. For example, thread scheduling was written to instead involving painters, rooms, and drying time. We were able to hire 4 women on a ~50 person team in a very short time, and it worked out pretty well.
No. If you go down that path, then you are implying that women do in fact perform worse at CS-related questions. That's a much bigger can of worms than the bias being implicated here.
Sometimes, they seem more like a secret handshake you need to memorize to get into the boys club than actually useful engineering. Who hasn't had to revise some of these before applying for a job 3 years out of college?
What it does do is effectively exclude applicants who didn't study CS, or who haven't heard of and memorized "cracking the coding interview".
Assuming `fake CS questions == good engineer` is a huge mistake, but one i keep getting downvoted for everytime i mention it. Most rebuttals are usually something like "it's the best system we have", something i find unsatisfying.
While I understand that using IQ tests as hiring predictors is itself a problem, I'm interested in the interplay in predictive ability between the two classes of tests. I think everyone would agree that any primarily intellectual timed test that was _less_ salient to work performance than an IQ test should be binned. What would happen to our interviews then?
This hits home so hard.
And, yes, Bay Area tech hiring is needlessly hostile for men over a certain age as well.
I think the recruiting team at my company would very much be interested in speaking with you.
I guess most men have thick skin or got lucky, so they don't see that, instead they think it's this dream job and everyone should partake in its wonderfulness.
I believe that the lack of women in tech is explained by societal bias against women and the nature of the job.
Does nobody want money in a capitalist society?
This is of course very much dependant on the distribution shapes and I am too lazy to make a thorough analysis - but:
Let's assume that on average females were 10% more efficient programmers - but with the effort to find one female programmer you can find 10 male programmers. How much more effort do you need to find a 10% better programmer - twice as much as for the average one? Even if it was 8 times harder - then still it would make more sense to look for only men than for only women. Of course the optimal way would be to be unbiased and look for any gender.
I think that depends on how your hiring process looks like in general. You might, for example do a benevolent (from a certain POV) discrimination and simply start filtering applicants with something like a naive 20-line Python script that matches applicant names against female names from a dictionary and pushes them to the top of your applicant stack, so to speak.
And there are less tangible or directly measurable, but nonetheless important benefits to hiring women for a business: You can get free publicity and marketing if you run a successful women-only shop, there is a significant demand in the liberal media for female success stories and you can ride that wave.
1. The issue is certainly bigger than hiring. In the many years between birth and looking for a job, there are a lot of societal pressures that will impact what eventual careers people end up in.
2. Hiring managers are people. They are not perfect. They have biases. If someone expects an engineer to look, talk, and act a certain way, that can impact their decision making completely independent of the fact that they want to hire the best people for their company.
Bonus third point: I still see a whole lot of "We want to make sure that the hire fits on the team." This is completely natural, and comes with its own set of built-in biases.
That's exactly the same argument used to justify every regressive policy. If x was true then rational y action would happen.
But that's the poo tof racism and sexism rational y action doesn't happen due to the -ism.
Or maybe women just aren't as smart as men.
Why is the health care field heavily biased in favor or female nurses and doctors? Are women smarter than men when it comes to biology/anatomy?
So, let's think about why we see gender roles in employment. Why are there so few women software engineers? One possible explanation is that women just aren't smart enough. If you don't believe that (and I don't), then you need another explanation. Maybe it's because of sexism. But if you don't want to believe it's sexism (as the OP implied), then what is it? They're not too dumb, and the hiring process isn't sexist, so why? And that's where hands come up empty.
That leads to nonsense like the person on this thread who said women are "wired differently", which presumably makes them less suitable. Which is just a polite way of saying women are too dumb to program, without facing the reality that that's exactly it means.
there seems to be a presupposition here that the the 'natural' proportion of women software engineers is 50%.
So what is the cause, then? Is it biological, or social, or random chance? "Random" doesn't seem likely, especially given how many other professions are male-dominated, and the relative economic and social power of those roles, compared to female-dominated professions.
"Biological", if it doesn't map directly to intelligence, needs another cause - something that can be measured. Do you have a suggestion for this? I don't.
"Social" is the most likely reason, but how is "social" different from "discrimination"? How do you define a social cause for men dominating the industry that can't be readily interpreted as discriminating against women?
I work in personality psychology research, so this whole IQ-centric line of reasoning is very dubious to me. There are many other influential phycological factors involved in people's lives that aren't (as far as we know) a direct result of nurture, and when taken together often make a more significant contribution to people's lives than their score in the single dimension of IQ. Learning disabilities and affective/mood disorders are a big example of this, and personality traits are just as impactful in how a person's life unfolds, regardless of intelligence.
A trait not being the direct result of nurture does not imply it's the result of a traditional long generic process, and this is something that we're only just beginning to scratch the surface of with epigenetics, so it's unlikely that such questions will get definitive answers anytime soon. That being said, the observation that a trait may be determined at birth only suggests that the trait is heritable, but not that it's genetic; those are two separate concepts, and heritability allows for much more variation from generation to generation, such as the case of children of immigrants from poor countries generally being taller than their parents when they're raised in western countries (which is likely due to improved nutrition enabling the full expression of their heritable height).
For example, you could ask the same question about whether the increase in learning disabilities and affective disorders within the past few generations in western societies is also "genetic". The default answer there of course, is that these conditions were only formalized as officially recognized diagnoses recently, and that such traits are only known to be heritable anyway (i.e. there are no definitively known "autism/adhd/etc genes" as of yet), so they're likely caused by the combination of the environment enabling the expression/observation of heritable predispositions. We can then similarly propose a null hypothesis to the male/female divide with the observation that western societies have only recently attempted to become more egalitarian by making various fields more equally attractive than they used to be, along with technological advances creating even more of such equally attractive opportunities, leading to heritable traits expressing themselves more noticeably through choices in the overall job market. In other words, being a professional "gamer" wasn't a viable job option 500yrs ago, but neither was being a professional "camgirl" either (to use two distinct, yet similar and stereotypically gendered "modern" occupations), but being a farmer was, in which case equal male/female distributions among farmers would've been the result of an underlying bottleneck in the pipeline, rather than the lack of one.
To suggest that this issue is either purely "genetic" or purely "social", is severely oversimplifying the matter.
Actually I grew up hearing exactly the opposite - that girls are smarter than boys and that girls "mature" faster than boys.
Maybe anti-male sexism prevalent in the health care and education fields is causing women to prefer those fields.
Fix the sexism in health care/education. Elementary teachers should be 50% men. Nurses should be 50% men. Instead those fields are 90%(!) women! That is a HUGE level of bias and discrimination
Possibility 2: Those fields are female-dominated because they can't get into male-dominated fields.
So what do the pay and prestige look like for female fields, vs male fields? Well, take medical. Nurses (low prestige, low pay) are >90% female. Doctors (high prestige, high pay) are about 70% male.
This suggests to me that there's indeed a huge level of bias and discrimination, but not in the way you think.
Possibility 2: Those fields are male-dominated because they can't get into female-dominated fields.
Men do not work as teachers because the media has painted men as "sex crazed". Most mothers would be uncomfortable with having a male 4th grade teacher for their daughter.
Many women would be uncomfortable having a male gynecologist or a male nurse helping them deliver their baby.
> Doctors (high prestige, high pay) are about 70% male.
Sorry but this breaks your narrative: 60% of new MDs each year are female. However: female MDs are more likely to quit the profession or go part time in order to raise kids. Again, this might show anti-male discrimination because it is not socially acceptable for male doctors to quit work to stay home with the kids.
---
The above suggests to me that there's indeed a huge level of bias and discrimination, but not in the way you think.
-- edit: fwiw, I googled stats. According to the American Association of Medical Colleges, 2017 was the first year ever that female medical school enrollment was greater than male medical school enrollment. I also went to graduation by year as far back as 2002, and it has always been more men than women. So yeah, your statistics are bullshit. Care to offer a source? --
And mind you, being a stay at home parent is considered a low-prestige, low-pay role. To the extent that it's discouraged for men, that's a result of a sexism that puts men in a dominant role and demeans them for doing "women's work".
The idea that men aren't teachers because the media paints them as sex-crazed is absurd. The gender disproportion of teachers existed long before the media mentioned such things at all. And you offer no evidence whatsoever for the assertion.
> Many women would be uncomfortable having a male gynecologist or a male nurse helping them deliver their baby.
And what is your opinion of the above bit of my previous post (since you avoided that in your answer?)
Is it possible men and women weight values differently when selecting occupations?
What is the payoff for pay and prestige and is it the same between genders?
I think the overwhelming evidence suggests it is not the same and that women value different things than men.
Except they're not, they're only empty if you haven't done any reading in this field.
> That leads to nonsense like the person on this thread who said women are "wired differently", which presumably makes them less suitable.
That was your supposition, not the only intepretation of those words. In fact, the weight of the evidence seems to support his statement, but similar to Damore, people like you are just fond of attacking reactionary strawman interpretations of the words actually employed.
> Which is just a polite way of saying women are too dumb to program, without facing the reality that that's exactly it means.
No it's not. "Wired differently" can mean many things, only one of which refers to competence.
There is strong evidence that women are on average more interested in "people" and men more interested in "things". Several references in http://slatestarcodex.com/2017/08/07/contra-grant-on-exagger...
Male nurses now actually can find they have an advantage in hiring because they often have an easier time with the lifting and physical labor being a nurse often requires.
My point is the comparison between nursing and programming is not strong.
Unconscious bias is a thing.
...if and only if...
...there were no other factors at play that cause that market to remain untapped.
For further rational thinking, consider this. If there's a bias, it doesn't mean women won't get hired. It just means they won't get hired for the best positions. Everyone else gets Amazon's cast-offs.
"every tech company would be taking advantage of it" - nope, no one is. I don't know why but my guess is its hard to admit you're doing hiring wrong, hard to hire people who think differently than you, etc.
Of course, in general, you can make a job more attractive (raise salaries, roll out red carpets, install slides...), and you will attract more people. That doesn't prove those people were an untapped talent pool.
Presumably there is a price that would make a high school teacher consider working in tech again. That doesn't imply companies should be willing to pay that price.
I keep seeing all these explanations that are just begging the question.
see: https://stats.oecd.org/Index.aspx?DataSetCode=EAG_PERS_SHARE...
Are there specifics about hiring processes in tech that bias against or scare away female candidates?
1. There's no reason to expect that these women will be unemployed - they just won't be working for Amazon. That's all we know. No point going looking for them.
2. You can't assign intent to hiring decisions made in the training data - there's no reason to believe that men (and why single them out?) "did not want to hire women". Maybe they did. Maybe they have no idea that they're biased - maybe the women making such hiring decisions are just as biased. We have no idea.
3. The evidence that the AI is biased, is that.... the AI is biased. Which means that the training data is biased. Why that is, is a great question - it may reflect unconscious bias in the hiring process, or more obvious old-fashioned biases. It may reflect that the model amplifies some minor bias in the training data and turns it into something much bigger. We don't know.
So yeah, it's biased - the question is why.
Nobody is saying that the biases was caused because it was created by men who didn't want to hire women. That's a fear-mongering straw man.
What people are saying is that there was bias in the training data selected, and so the algorithm exacerbated that bias. Thus, being a cautionary tale about the training data you feed to these things.
" If there was an untapped market of engineers you'd better believe every tech company would be taking advantage of it."
You're assuming rationality where there really is no cause to do so.
If we're being honest, a system only needs to be in a decision-making capacity for discriminatory behavior to be scrutinized, since in many cases human operators will not be able to identify the specific features being used to make decisions about people -- the features could be highly correlated with some subpopulation of protected class. If you take that to be true, the question reduces onto what decision-making roles ML algorithms have that could be discriminatory, and it's hard to argue this is not a massive part of their current and expected roles.
I think this is going to be a long, winding ethical nightmare that is probably just getting started by human-digestible examples such as these. One can imagine things like this one being looked back on as quaint in the naivety to which we assume we can understand these systems. Where do we draw the line, and how much control do we give up to an optimization function? Surely there is a balance -- how do we categorize and made good decisions around this?
As far as I know, a cohesive ethical framework around this is pretty much non-existent -- the current regime is simply "someone speaks up when something absurdly and overtly bad happens."
This is just Simpson's paradox [1] which is notoriously hard to identify because you have to compare the overall with the breakdown. As you say, current-AI probably already has such biases.
This question can be rephrased as "is there a difference between de facto and de jure discrimination?"
My answer is no, causality doesn't matter here: if feature A is a good predictor that some person belongs in group B and not group C, then filtering out feature As is effectively the same as filtering out only group Bs.
If you're hiring therapists, and your candidates take a personality test, and your ML model weights the 'nurturing' feature highly, is that discrimination because it selects against men?
What I don't agree with is the assumption that, in this case, the preferred traits do correlate with fitness, since there's at least one — gender — for which this model is biased even though it has no apparent correlation.
The article notes that Amazon's system rated down grads from two all-women's schools. But it immediately occurs to me to wonder what the algorithm did with candidates from heavily gender-imbalanced schools, which could be much harder to spot.
RPI's Computer Science department is about 85% male, while CMU's is just over 50% male. CMU's CS department is also considered one of the best in the world, and presumably any functional algorithm that cared about alma mater would respond to that. So if the bias ends up being "because of CMU's gender ratio, CMU grads with gender-unclear resumes are advantaged slightly less than otherwise would be", how on earth would someone spot that?
Once you're looking for it, you could potentially retrain with some data set like "RPI resumes, but we adjusted their gendered-words rate" and see if you get a different outcome on your test set. But that's both a labor intensive task, and one that's only approachable once you already know what you're looking for. And even if you do see a change, you'd still have to tease it out from a dozen other hypotheses like "certain schools have more organizations with gendered names, and the algorithm can't tell that those organizations are a proxy for school".
Of course, the counterpoint is that human decisions can't be scrutinized any better, and it's not entirely clear they're less arbitrary or more ethical. At a certain point algorithmic approaches are being scrutinized because they're slightly transparent and testable, so running them on a range of counterfactuals or breaking down their choices is hard rather than impossible. I suspect that's true, but it doesn't really comfort me - humans at least tend to misbehave along certain predictable axes we can try to mitigate, while ML systems can blindside us with all sorts of new and unexpected forms of badness.
Also knowing some people who worked on this, they were VERY cognizant of re-encoding biases from the start of the project, it was one of the main reasons they thought the project might fail.
"Amazon edited the programs to make them neutral to these particular terms. But that was no guarantee that the machines would not devise other ways of sorting candidates that could prove discriminatory." I read that as a very different statement - as written, Amazon corrected two specific instances of keyword gender bias by hand, but couldn't reliably prevent further bias (including gender bias) from arising. That's where tricks like "ask the system to classify gender, and then un-train via that data" come in.
(I don't mean you're wrong, just that if gender bias was accounted for more generally, the article should have said so.)
That said, I think our disagreement might just be a miscommunication on what went wrong in the first place. If you know some people involved, maybe you can help clarify the situation?
The article totally fails to explain why "most engineering resumes are from men" led to an algorithm that downrated female resumes. "Most applicants had brown hair" does not produce a system that downrates blondes if you tell it hair color. So the question is - was the training data biased against female applicants (in which case why wasn't it caught before specific outputs needed modification?), or did something else altogether cause this issue (in which case what?)
If they only trained on who was hired they wouldn't really know if those were good hires.
Facebook is already auto-flagging content this way but it's just a very hard problem (even for humans).
I hate to sound like "that pedantic guy", but I'd argue that the quote above is only partially true. It's the case that some subset of AI techniques "learn from the training data it's given and copies any biases this data exhibits". There are AI techniques that aren't based on supervised learning from a pre-existing training set. That doesn't mean that those techniques can't wind up adopting the biases of their human overlords, but I believe some aspects of AI are less susceptible to this kind of bias, than others.
This is the same line of already-refuted reasoning behind the "I'm just asking questions" in the infamous Google Memo.
To answer my original, rhetorical, question: It's not cynical. It's wrong.
What about citations proving that "software engineering when it was considered a low-class low-skill job" is the same profession as the programming in the past 30years. Or at least that it has the same difficulty / processes.
Btw usage of phrases like "clear history", "so much evidence" (especially when you cite one(!) arguable data point), "already-refuted" does not convince anyone about you being right. It is at best annoying.
From what I've seen this is standard assumption by ordinary people. And it does not target only programming, any office job that involves "sitting at the computer all day" gets that reputation.
And that easily could have not been the case in the 50's (I really don't know). And the profession has clearly evolved (anecdotally: many think it has gotten worse). So your assumptions are really not that obvious. Sorry if that comes accross to you as "arguing in bad faith".
Are you just arguing for the sake of arguing? You're not engaging with my points in any meaningful way. Can we be done with this thread?
You are not engaging in debate in any meaningful way, maybe stop arguing on HN? You are not convincing anyone . . .
[0] https://www.smithsonianmag.com/smart-news/computer-programmi...
You're missing the point anyway. The article made it pretty clear that this AI amplified biases humans already have about women applicants to tech positions. Stating your own biases about women make no sense to the topic, or to the argument you seem to be trying to make.
Janitors are worth talking about as well (women in the same job usually have a different title with less authority and less pay), but high-status, highly-paid, highly influential jobs are where it's most important to avoid bias, and so we talk about those more.
There are enthusiastic people who start with enthusiasm and can keep that enthusiasm. Most begin with practical concerns. And most of those who don't begin with stay with. he
And yes, rich people, upper echelon, people of means are rare in the industry
I see four possibilities here:
1. The algorithm was designed in a completely inept fashion
2. The algorithm design was sound, but ultimately ineffective
3. The algorithm was sound and effective, but results were considered discriminatory.
4. There's something biased about how employees are rated--the data that would feed into the algorithm, which is possibly more of a human element.
Edit: Added fourth possibility
And whatever the cause was, it was not the poor quality of the training data. They tried to stop the model from downranking women based on obvious keywords, only to find it learning to downrank them based on more subtle language cues:
> Amazon edited the programs to make them neutral to these particular terms. But that was no guarantee that the machines would not devise other ways of sorting candidates that could prove discriminatory, the people said.
So the answer is 3 or 4.
If the answer was 4 then they would have probably mentioned the cause of the bias somewhere in that otherwise detailed article. But they didn't, possibly because the cause is controversial - probably option 3 but possibly still option 4.
And then there's the subtle cop-out:
> Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs, the people said. With the technology returning results almost at random, Amazon shut down the project, they said.
If the model was actually useless and returning random noise, then there wouldn't be any bias, and the article wouldn't need to talk about discrimination. This paragraph reads to me like they decided to mention long-tail results (that you'd find in any ML model) as supportive 'evidence' that the model was somehow broken rather than producing valid but controversial results.
Basically, people WANT bias, but they want specific bias. One of the difficulties in training a machine to understand what you find as viable bias vs problematic bias is all the tiny nuances. Yes, you want a great engineer on paper, but you also need to have as diverse a cast as you can in your company (both for optics and creative solutions) AND you need to get people you can afford AND you need someone who's enjoyable to work with etc etc.
Hiring is always going to be part art and part science. There will always be some type of discrimination because of the perceptions of what makes a good qualification for the job. Any hiring group is just going to have their own hierarchy of what they think are the most important skills to have. You can only approach perfection/unbiased hiring, you can never actually achieve it.
Just because some groups have competencies in this area, doesn't mean that others do. I've worked at big tech companies that couldn't get their HR systems to work properly ... IT was abysmal even though we made 'high tech'. Also, it's an internal project, not a product, so the scope of investment etc. might have been very different than otherwise.
> Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs, the people said. With the technology returning results almost at random, Amazon shut down the project, they said.
It looks like the bias wasn't the only flaw.
I’m not sure what the best alternative should be, though. I am a fan of open source work as a sort of code portfolio, but it doesn’t work for every kind of engineering/science (edit: and also would introduce bias against professionals too busy for open source.)
Regarding bias — it seems the only way to truly eliminate it (including unconscious bias) is author-blind reviews, i.e. reviewing code written by a candidate without knowing anything about that candidate’s identity. (And the nice thing about code is it usually doesn’t signal any identity traits of the author via side channels.)
We've had quite a lot of discussions about this internally, and even with humans at the helm with best of intentions about being unbiased, its really easy for a lot of bias to slip in. Even things like the phrasing of questions can introduce bias (i.e. the ol' apocryphal SAT word association problem that had 'regatta:boat').
Should I ever be in a position to hire a colleague, I wouldn't ever do so without having a chat with them.
I spend 8hrs a day in an office with my colleagues (sometimes more than with my wife & kid) and the ones I can't stand is about the only thing wrong with my job.
If we can't even see the person's face over some gender bias hysteria then I wonder how the hell we got here.
People should just get over the fact that men and women are different.
It's great when everyone can coexist perfectly. However, that might not be the best business decision. There's no such thing as "objectively best" just a list of pros and cons to any candidate, and a company's internal preferences.
Yes there’s a lot of “bias hysteria” out there, as you put it, but I would dispute that advocating “author-blind meritocracy” falls into that category.
Quite the contrary: An author-blind review process would actually make any bias impossible — either for or against any particular identity group. It seems to me most people should be able to get behind that, but maybe I’m wrong.
In fact, the main opposition to author-blind meritocracy is the “post-meritocracy” movement which is slowly making its way into open-source projects codes of conduct.
When it gets to the "let's meet" stage, it could be possible that the bias just comes back. Yes this woman made a perfect score on the coding but she's a woman and I'm not going to hire her for reason X that just bubbled up out of my biased brain. I can totally see that happening, unfortunately.
There's plenty of people I would never ever spend time with outside of work, and try to minimize my time with at work.
But that's fine, because 'cthalupa would like to have a beer with you after work' isn't part of the requisites for doing a job on my team. The 'Finding people that fit in with the culture'/'Finding people that I don't mind being around' is how you get monocultures and a lack of diversity in your team.
>If we can't even see the person's face over some gender bias hysteria then I wonder how the hell we got here.
Gender bias is a real thing, a big deal, and certainly not hysteria. There's a lot of ways to reduce it. It doesn't necessarily require never seeing someone's face - though I think automating "skills" related interviews could be a good thing - because you can start with having a structured interview program where you have specific questions to ask and a specific rubric to grade against. Making sure you have solid, unbiased questions, and measure the answers evenly against the same rubric solves the majority of the problem.
>People should just get over the fact that men and women are different.
Well, of course men and women are different. But even for jobs that involve heavy labor, this isn't actually significant - while the average woman has less physical strength than the average man, the type of woman who applies for that sort of job has self-selected into it, and is almost certainly more capable of doing that sort of work than the average woman - meaning it's still not a good indicator even for areas where the differences are largest.
For a white collar job, like the type we're discussing? It's even less relevant. There are extremely few times where you should ever care about gender when it comes to hiring.
You can de-bias by explicitly controlling for gender, but now everyone in your company went to CMU and likes dogs.
The more I see news about what recruiting for ultra-large corporations, the more I think one of two things is true:
* ultra-large corporations are doomed to hire less and less well in a way that is more and more biaised, and we should regulate against such corporations in a way that forces them to redistribute their wealth to SMBs;
* ultra-large corporations need to start exclusively growing through acquisitions, which will have the effect of redistributing their wealth to SMBs, and also of hiring a more diverse base of employees because there is a priori a greater diversity of backgrounds leading to success in the free market than the diversity of backgrounds leading to success in the Amazon interview.
The best thing to be right now is a woman engineer. You can easily get hired within the week.
Unfortunately this doesn't seem to be well known outside of those involved in hiring.
Silicon Valley did a good bit on this https://www.youtube.com/watch?v=Dek5HtNdIHY
It's funny because it's true.
The system failed because they were trying to solve the wrong problem, or maybe more specifically, didn't solve the problem that led to the problems with the AI. Amazon was treating the hiring problem as an efficiency problem alone, and ignoring the bias problem. So they wound up training the AI to do a shitty job much faster than humans ever could be shitty - and, by analyzing the data in a way the human results weren't analyzed, showed the failings of the human hiring process.
Existing process is sexist. Automate to "improve" it, and you wind up with something even more sexist. What this means is that Amazon needs to go back and revamp their whole hiring process to make it fair, before trying to make it faster.
If nothing else, modeling our existing behavior in this way is a great use case for ML. As it allows us to "fast forward" and thus--hopefully!--identify our flaws based on modeled iterations.
Could this be a direct indicator of a powerful subconscious bias in Amazon's existing hiring process?
Yes, but only in the obvious sense that we all already knew: tech companies hire more men than women for technology-focused roles. That's not to say it isn't an issue; it is, but it's nothing new, and almost certainly not unique to Amazon.
Without significant oversight and manual tuning, any training dataset based wholly or in part on current employees is going to demonstrate a bias against women, because there are strictly fewer women. Moreover it's likely that (for a variety of reasons, both intrinsic and extrinsic) fewer women succeed in the interview process as a ratio compared to the number who actually apply.
The difference is that old place transitioned administrative staff to IT roles in the 90s/early 2000s when more things were computerized. Those admins, financial analysts, program analysts were more likely to be female and had degrees in liberal arts, accounting, business/finance, etc.
In the newer place, they filtered based on computer-related degree upon hiring. That automatically excludes many women. Once hired, female candidates advance as well or better.
Anecdotally, I've hired interns in recent years with no tech-specific qualifications as an experiment. If you select for "smart and gets things done" I don't see much of a disadvantage for many roles. You get some duds too, but it wasn't as dramatic a difference as I expected.
I wish companies were more willing to train willing applicants instead of trying to interview for capabilities.
"Demonstrate" may be the wrong word. How about "encode". And in the future, "enforce".
[Edit] On second thought, it would seem like there would be a way to filter out a "raw proportions bias" (like 80% of the resumes in the data set are male) before training.
I can't imagine this not already being a thing, but I haven't really heard of people using this method.
From the article: "Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs... With the technology returning results almost at random"
I am quite curious about details of the model. For example, the single largest contribution to real world interview process variability is interviewer (for resume screening, who screened that resume, etc.). Wouldn't it be possible to code interviewer as categorical variable and separate resume-intrinsic? effect and interviewer effect? They must have tried this, haven't they?
If there are strictly fewer women in the underlying training set, the model can still return something resembling a uniform distribution of candidates while exacerbating the diminished representation of women.
To give a concrete example: you have a bag of blue dice and red dice. There is a supermajority of blue dice in the bag. Your algorithm selects a single die out of the bag on every iteration. The output sequence of dice numbers appears uniform, but there are more blue dice than red dice in the output sequence.
Could this be a direct indicator of a powerful
subconscious bias in Amazon's existing hiring
process?
Maybe - but maybe not.Imagine a company with 2 men in HR, 2 women in HR, 40 men in engineering, and 10 women in engineering. That's with gender-blind hiring, reflecting only the 4:1 ratio of male to female CS graduates.
If you picked a random male hire, there's a 40/42=95% chance they're an engineer whereas if you picked a random female hire, there's a 10/12=83% chance they're an engineer.
Thus if you look over all hires' CVs, due to Bayes’ law the dataset says being male increases the conditional probability you meet engineering hiring requirements - and the ML system picks up on that.
You're right you'd want to look at applicants' CVs - I skipped over that to make the numbers readily comprehensible.
That language is a bit ambiguous, it could just mean that the algorithm failed on a wide variety of jobs beyond engineering. But another reading suggests that the algorithm was not asked "is this person a good fit for this role" but instead "what, if anything, is this person qualified for?"
If that's the case, then the problem starts to make more sense: the algorithm learned a correlation between male-sounding resumes and being hired for engineering roles. That could produce a biased approach even if the decisions in the training data were gender neutral but position-specific. Of course, it would also mean that an Amazon ML team trained an algorithm with inputs that didn't match to its eventual task, and makes me wonder what they used as a test set...
(Anecdotally, Amazon spent quite a while recruiting me for SysEng work I'm wildly unqualified for and uninterested in, even suggesting a switch to applying for that team when I was already in the funnel for something I'm more qualified at. When my resume eventually made it to a syseng engineer, they were rightly baffled that I had landed on in their stack, giving me the sense that something was screwy with how Amazon decides who heads towards which role.)
So the classification function should take into account the resumes of rejected engineers, rather than the pool of resumes of hired employees at Amazon. If someone is seeking a position as an engineer, it is not relevant how much their resume resembles that of HR people, but it is very relevant how much it resembles that of rejected engineering candidates.
If that's the case, then something like having the phrase "women's chess club" in one's resume should not be a meaningful factor for the classifier unless it disproportionately leads to rejection in the current process.
So I doubt it's enough to explain their issue here. I agree that we can't really take any conclusion of their broader hiring patterns from this experiment.
This view that the only thing holding people back is some sort of social or systemic bias seems to be based on nothing except ideology. Incidentally, it's an ideology I also used to hold. Like a good egalitarian I pushed my wife away from sociology and into majoring in CS. She did perfectly well, as did I. More than a decade later she works with people and I work with code. I've no regrets there, but it's not so clear that my persuasion was really the best idea.
Norway is another interesting example here. It is considered by many to be the most gender equal location in the world. Yet you'll still find that nurses are primarily female, doctors are primarily male, and all other 'stereotypical' divisions present in most all developed nations. They tried to change these divisions and with extensive effort were able to effect a roughly constant change in some fields. But again, once that push was relinquished things went just about identically to as they were in very short order. Ultimately we're flexible enough that you can manage to fit a square peg into a round hole at times, but once you stop squeezing that peg goes back to what it wants to be.
Does she make more money than a similar person with a degree in sociology?
At the least, they should interview a certain subset of randomly chosen applicants or else the feedback loop from the interviewing process and the AI is going to grow tighter and tighter.
I don't think this will ever work. There is too much variability in resume wording that correlates to gender and even culture of origin even when you take out names and any other protected class identifying markers. The Dutch tried this and ended up with less diversity.
I'm going to go out on a limb and say you almost want to leave all that identifying data in, but put each candidate into buckets with separate rating algorithms trained against only that "type" of candidate. The top candidates from each culture, and the top candidates from each gender, etc etc, however you want to do it. Feed them into a picking algorithm that builds a composite of what you want your team to look like diversity wise based on the top candidates from each bucket, and go from there.
Don't take my opinion seriously, I'm not an ML guy.
The project was doomed from the start.
Seriously: let us take as given that the AI models are biased. Will you also admit that the existing processes are biased? If so, then what we need to ask is which is MORE biased. It might be complaining that we shouldn't release self-driving cars because on rare occasions they cause accidents.
There is, however, another criterion besides how biased it is: how biased it will be in the future. Human-driven processes have the opportunity to become less biased in the future (also the chance to become more biased, but overall things tend to improve). AI processes that are opaque might lock in bias in a fashion that is unreviewable. I believe that the solution is to build AI models that are more transparent -- that could be BETTER (in terms of avoiding bias) than the human-driven processes we use today.
It's just that generally what seems to distinguish whether something is called "machine learning" rather than "data science and modelling" is that the former is black box and the latter is not.
It has to be better than humans.
That requires the rare genius to fix the problem, or we have to accidentally invent something better than ourselves.
This is a common defense of autopilot systems in cars: it's not perfect, but it's better than the average.
http://blog.alinelerner.com/lessons-from-a-years-worth-of-hi...
That aside, what sucks is that attempts to automate resume scoring rarely look at harder-to-quantify features and focus on low-hanging fruit like keyword occurrences... though in my experience it's such a low-signal document for engineering hiring that the whole thing is a fool's errand.
This should be obvious when testing. Whether the algorithm discriminates should be a top priority for designing these algorithms. That's half the damn math of machine learning. If you can construct an AI, you should know how to test it for flaw in reasoning. It's just another layer of ML to do that. Outliers. It's short sighted to push these things out assuming their output is correct just because it looks 'normal'.
I wouldn't know unless I looked at all the data. But I'm not going to default to the popular opinion because that's literally half or more of the problem.
There's also the idea that lack of women scientist "heroes" can be limiting (lack of role models). Basically the idea that if you stack the cards against a population, you're gonna see population-wide effects.
Given these data points, a biased hiring AI contributes to the problem. Therefore, it should be fixed, along with the above points.
[0]https://www.bbc.co.uk/news/world-40865261
[0]https://www.theguardian.com/technology/2017/aug/13/james-dam...
This one is a bit weird, computer guys were always "nerds" and "geeks" to stay away from.
Depending on the era, we had Einstein, Turing, feinman. Kids my age had Gates (literally the richest man on the planet for my entire formative years), Jobs, Bill Nye. Little further along are the myth busters crew, musk...
We have plenty of heroes to pick from :)
>"that using someone's sex to work out what you think their personality will be like is "like surgically operating with an axe"."
Being phrased by the article as a dismissal of Damore, along with G. Rippon's statements However in the article Schmitt is quoted from, he writes that
>"Culturally universal sex differences in personal values and certain cognitive abilities are a bit larger in size (see here), and sex differences in occupational interests are quite large. It seems likely these culturally universal and biologically-linked sex differences play some role in the gendered hiring patterns of Google employees. For instance, in 2013, 18% of bachelor's degrees in computing were earned by women, and about 20% of Google technological jobs are currently held by women."
He goes on to write that Pyschological sex differences might lead to less than 50% of technology employees being women.
This seems to disagree with Professor Rippon's opinion that
>"but even if you accepted the idea that there are some biological differences, all researchers would assert that they're so tiny that there's no way that they can explain the kind of gender gap that's apparent at Google."
I think there's reason to consider both the societal reasons women might be pressured and excluded from STEM-ey fields, as well as potential inherent differences in interest, and that they can both coexist as considerations, and agree that a biased AI is unhelpful, and many women lack a fair shot of success, however disagree that there is nothing useful in Damore's perspective.
Additionally if such inherent differences are distributed on a bell curve, it would make sense that at cases further along the trail that small differences in populations and their medians are more pronounced.
Helping individuals to overcome biology is much simpler than doing it at population scale.
Where does it state the number of applicants, male vs female?
For now I say it was "handled" in that not only did he fail to demonstrate that female disinterest in engineering, compared to male, is due to inherent psychological differences, and I quoted a couple people far more qualified than me that reached the same conclusion (their statements are in the article. The Wikipedia page is another good summary)
When right wing trolls attacked a female CS lecturer, she wrote a long response here: https://www.vox.com/the-big-idea/2017/8/11/16130452/google-m...
So the bias against women due to decades of societal conditioning leads to less than 50/50 representation because less are applying, which companies are trying to patch by leveling the playing field, making their internal population breakdown identical to the external one.
Seeing shitloads of Indians is a passive effect of that internal/external thing - there are around 1.2 billion Indians...
I must say, I am frustrated by this being brought up in every discussion of women & STEM. Want to discuss the leaky pipeline from physics PhDs to full professor in physics? "Maybe those women did a PhD in physics despite not being interested, and they just didn't notice before!"
* One says "executed concentration camp prisoners in Kosovo". (Yeah, ok, I'm kidding. How about "Executed a plan to reduce production costs by 30%"?)
* The other says, "Won the Women's World Chess Championship 3 years in a row".
The first has five stars (thanks to "executed") and the second has three (courtesy of "women's"). Which are you more interested in?
There was a study published a while back that looked at PISA data and found that girls and boys were pretty evenly represented among the kids who were at the top in STEM [1].
But it also found that for the boys in that group quite often STEM was the only thing they were outstanding at. In other areas they were average to good.
For the girls, on the other hand, they were often excellent at something else in addition to STEM, with them often even being better at that something else than they were in STEM.
People have a tendency to pursue a career in one of the areas they are very good at.
This suggests that boys who are very good in math, etc., are more likely than similarly good girls to pursue it as a career because that is their only choice if they want to go into something they are very good at. The girls are more likely to have math, etc., as one of two or more possible careers in areas they are very good at.
In pop culture terms, STEM boys are more like Martin Prince, and STEM girls are more like Lisa Simpson.
[1] I didn't save the link and have failed to find it with Google. Anyone have it?
I'm hoping industries that hire young are seeing different numbers than I did, because that should signal a shift in older ones that hire senior discipline engineers after a decade or so.
Edit: that said, companies should continue to do what they can to remediate this, but I am furious that the government has done almost nothing about the issue. The underrepresented remain exactly that.
The painful stuff is when it's obvious and provable, because it highlights all the times it can be questionable as to whether it occurs.
In my experience in the industry, this is a laughable statement to make. It's a shame that unless one is a victim of unconscious, systemic bias, one is so much less likely to acknowledge it as a problem, that actually exists, and hurts people all the time.
If that's your point, I guess my counter argument is myself, not a victim, very aware of the problem, and acknowledging it as a problem as my post.
There's also the victims of conscious bias that would probably be able to acknowledge the problem...
Am I misreading your post?
It doesn't matter if it's intentional or not to a victim. It's still the same system, same cause and effect, same yield of powerlessness.
Whichever way you want to see it.
+ If it's intentional it's not unconscious.
+ If it's unconscious it's part of a culture that tolerates the behavior to the point that it doesn't get questioned.
+ If it does get questioned, eventually people are just playing dumb or it becomes intentional - if it's provable that it continues to occur.
It was still just a trickle, for the same reason you stated - very very few woman apply to tech positions.
Time and time again people (mostly men of course) keep asking "but why? why aren't there more women in the field?" Time and time again they keep saying "but I don't see any sexism in the workplace, it's nothing like it used to be, it's practically a meritocracy these days!" Yes, indeed, it truly is a giant mystery.
And yet, at the same time there is a constant deluge of stories about rampant sexism in the industry. Of all sorts, at all levels, at almost every company, and often of shockingly regressive character even up through the present time. There are countless stories in the industry of how women in tech are persistently denigrated, how men talk over them in meetings, how their ideas are ignored until they come out of the mouth of a man, how sexual harassment is ubiquitous, how they are routinely excluded from workplace culture through extremely male-centric activities that include things as ridiculous as morale events or even meetings held at strip clubs.
All of this takes a toll, and that toll is ultimately to stunt the careers of women in tech and to push women out of the industry entirely. Working in tech as a woman is climbing a hill with a much steeper slope than it is for guys. Women routinely get passed over for promotions, are routinely underpaid, routinely do not receive credit for their ideas, and routinely experience more hostile working conditions (through bias as well as sexual harassment). So they leave. They find something better to do with their time because they just can't take the stress and harassment anymore or because it just does not provide the same return on investment as it does for guys.
And we know this. We know this from studies and exposes and a torrent of anecdotes from individual women who have been in the field for years or decades. Some people (guys) have a tendency to write off each and every one of these stories and studies as somehow individual aberrations or outliers which don't have any bearing on the fundamental overall character of the industry, but this is a mistake, they are absolutely representative. The problem of over-representation of white men in tech cannot be solved by "fixing the pipeline" in the educational system nor can it be solved by making hiring processes perfectly unbiased (or even biased towards women) because the real problem is much bigger, it's systemic, widespread misogyny throughout the entire industry. That will take a tremendous amount of work to fix, but once the industry stops treating women as second class citizens (or exotic outsiders) and stops pushing them out of the industry through its toxicity then the problem will mostly fix itself.
People do let them. You can't force what people are interested in and you can't let in that which does not exist. In fact many places in tech give preference to women applicants, because they don't apply often and the companies want more women. They're just rare to see. :(
There's no grand conspiracy. The truth is much less exciting: Women and men have different preferences, generally speaking.
Women are more interested in people (i.e. healthcare). Men are more interested in things (i.e. engineering).
Most nurses are women. That's not because women are activity trying to keep men out. It's because fewer men are interested or apply!
We also don’t see women in the most dangerous jobs. No one seems to have a problem with that, just as no one has a problem with most healthcare jobs being dominated by women. As they shouldn't, because people should be allowed to pursue and apply to what they want to.
Achieving that level of a standard is a balance.
There shouldn't be excuses being made. All that can do is contribute to the perpetuation of the conditions that presently exist, because the core issue isn't being identified.
Furthermore, if the core issue is the excuse itself, then again, this is covered under the umbrella of 'if people would just let them'. The secondary issue would then be that the core issue isn't being questioned.
I want to consider your, er, argument, in best possible faith, but you've given me almost nothing to work with here.
Sweden tried what?
Failure means what?
It's surprising to me that Amazon didn't (apparently) try different models for different populations. Sure, it might open you up to criticism, but there are some good data-driven reasons to do so. Women's colleges won't show up with regularity on men's resumes, for instance. Similarly, there are fraternities and sororities around engineering and STEM that may provide different signals, but won't appear equally distributed on men's & women's resumes. Language use on resumes does differ by gender, and using "Captured value of $100 million by..." rather than "Created value of $100 million by..." may describe the same project. (I gotta say, using verbs at all seems silly, since it really is about how well you market, rather than what you did.)
So, curious about the model. Different models for different subsets of the training data can lead to big wins.
I also wrote a similar AI for finding surrogate pregnancy candidates and it also showed bias against men.
Goes to show how AI can fail and be incredibly sexist.
The actual newsworthy part, which is getting slightly stale, is that it was influenced by the data bias.
I am not even a fan of Amazon but I think this is unfair to them. They did the right thing here.
I wouldn't go so far to say that running your history through an AI is necessarily proof of anything though (esp. in court). Imagine using another company's historical data to suggest that they're discriminatory.
It is clear whatever executives in charge of this project haven't the first clue of whatever it is they are doing. It lacks not only any technical deliberateness but also fails the common sense test. I really wouldn't want to be working for such people.
I don't think we're anywhere close to general AI so any AI system out there is built on what we feed it. If your music suggestion algorithm has a biased AI that's one thing, but when you're using AI to make critical decisions in society like who gets hired/recurited, medical diagnoses and other things you need to be extremely careful.
When designing AI models the broadest set of viewpoints should be considered. Any piece of AI is simply a reflection of it's creators, we need to make sure that some sort of equitable consensus is reached before deploying AI to critical human scale issues.
Neither matters much, neither makes sense. What they should have been doing was training it on the resumes submitted by employees who then went on to be very successful within the company. Those are the successes. Those are what you want more of. And, probably far less likely to be weirdly biased by gender or race or whatnot.
Race: Human
Gender: Yes
Age of legal contractual consent: Yes
If you've ever trained a NN, you'll know that they are exceedingly clever in finding patterns that fit what you're training for. You can remove the word "women's" and other obvious things from being considered, but I promise you, if there's another non-obvious patterns that are more likely to apply to the women candidates, the AI will find them and use them.
Your submission was made 11 hours ago. If they merely applied an upvote to that existing story after 10 hours, the point time value rot would be such that it would have greatly reduced impact at ranking the story to the front page (ie it would be nearly useless for discovery purposes).
Doesn't this mean the headline is incorrect, incomplete and/or (intentionally) misleading?
This is the most interesting part to me. It’s suggesting that men and women think about things differently. Which, if true, suggests a lot of other possibilities that have relevance when hiring for certain positions.
But, it also means that bias is now measurable...
...and there were more than a few papers at NIPS that were directly dealing with "fairness" in a NN, aimed at addressing and using these issues and effects.
My only worry is that they are aiming for 50/50 which doesn't reflect the underlying gender ratio of the developer pool.
If this is what they want then they have to feed their neutral algorithm pre-biased data to get their expected results.
So there may still be a 8:1 ratio of men:women but according to the article, it would seem than even that 1 women would have been negatively impacted.
I am not arguing the validity of their conclusions, just saying that I don’t think that it’s about equal numbers: given a male and female applicant, according to the article, women would have had an unequal chance of passing the screening. (Now if that inequality was for valid reasons, that to me, is an open question, but the article indicates that there was unjustifiable bias.)
Fact is, there are more men than women working in tech, especially seriously hard-core stuff like these big companies need. This is most likely out of their own personal volition, and no matter how much outreach we do, this is likely to stay the case for generations.
Of course, that's also an egregiously wrongthink position to take. Double plus ungood.
I can easily see a bias towards a particular gender arising, even when a team had intentions to do, simply because of how the data is selected.
________
Case 1 : They used similarity statistics between candidates that were hired and applicants.
This is the easy one. No need to label datasets. The approach is semi-supervised. It will also 100% cause a bias towards employees with similar profiles as those already working at Amazon. (ie. men)
________
Case 2 : They manually labelled/ranked a dataset of resumes and assigned them scores. (more likely)
Here the implicit bias of the mechanical turks / rubric would be visible. If higher scores were assigned to traditionally masculine activities, then male resumes would stand out. I doubt this was the case though, as people at Amazon are generally competent enough to not make such a trivial mistake. Also, gender only gets mentioned in non-technical skills (mean's team, women's club, etc), which in general are not the most relevant part of the profile anyways.
________
Speculation:
1.
> Instead, the technology favored candidates who described themselves using verbs more commonly found on male engineers’ resumes, such as “executed” and “captured,” one person said.
I wonder if this has anything to do with gender at all. All good resumes that I've read use action words irrespective of gender. Maybe type A personalities use words like "captured" more often than type B ones, and the % of men in type A categories are greater. Would discrimination against women in such a case be unfair ?....Maybe.....Maybe not.
2.
Extracurricular activities are more prominent on weaker resumes than stronger ones. So, the model may be weighing down resumes with too much extracurricular fluff vs technical skills. Men's activities are rarely prefaced with the word "men" in it. (they would just say Football team, Chess team). Women's activities on the other hand, always have the word "women" attached to it. If the extracurricular activities were penalized, then the words inside of them, including "women" would also be penalized. Thus, the model learns a latent gender bias without any bad intentions.
_________
In ML, one of my favorite statements is : "The model is only as good as the data it is trained on." If the data is not sufficient, rich enough or prepared in the correct manner, then unintended consequences are nearly guaranteed.
There's only so much you can sweep under a rug.