Could this be a direct indicator of a powerful subconscious bias in Amazon's existing hiring process?
Could this be a direct indicator of a powerful subconscious bias in Amazon's existing hiring process?
Could this be a direct indicator of a powerful
subconscious bias in Amazon's existing hiring
process?
Maybe - but maybe not.Imagine a company with 2 men in HR, 2 women in HR, 40 men in engineering, and 10 women in engineering. That's with gender-blind hiring, reflecting only the 4:1 ratio of male to female CS graduates.
If you picked a random male hire, there's a 40/42=95% chance they're an engineer whereas if you picked a random female hire, there's a 10/12=83% chance they're an engineer.
Thus if you look over all hires' CVs, due to Bayes’ law the dataset says being male increases the conditional probability you meet engineering hiring requirements - and the ML system picks up on that.
You're right you'd want to look at applicants' CVs - I skipped over that to make the numbers readily comprehensible.
That language is a bit ambiguous, it could just mean that the algorithm failed on a wide variety of jobs beyond engineering. But another reading suggests that the algorithm was not asked "is this person a good fit for this role" but instead "what, if anything, is this person qualified for?"
If that's the case, then the problem starts to make more sense: the algorithm learned a correlation between male-sounding resumes and being hired for engineering roles. That could produce a biased approach even if the decisions in the training data were gender neutral but position-specific. Of course, it would also mean that an Amazon ML team trained an algorithm with inputs that didn't match to its eventual task, and makes me wonder what they used as a test set...
(Anecdotally, Amazon spent quite a while recruiting me for SysEng work I'm wildly unqualified for and uninterested in, even suggesting a switch to applying for that team when I was already in the funnel for something I'm more qualified at. When my resume eventually made it to a syseng engineer, they were rightly baffled that I had landed on in their stack, giving me the sense that something was screwy with how Amazon decides who heads towards which role.)
So the classification function should take into account the resumes of rejected engineers, rather than the pool of resumes of hired employees at Amazon. If someone is seeking a position as an engineer, it is not relevant how much their resume resembles that of HR people, but it is very relevant how much it resembles that of rejected engineering candidates.
If that's the case, then something like having the phrase "women's chess club" in one's resume should not be a meaningful factor for the classifier unless it disproportionately leads to rejection in the current process.
So I doubt it's enough to explain their issue here. I agree that we can't really take any conclusion of their broader hiring patterns from this experiment.
Yes, but only in the obvious sense that we all already knew: tech companies hire more men than women for technology-focused roles. That's not to say it isn't an issue; it is, but it's nothing new, and almost certainly not unique to Amazon.
Without significant oversight and manual tuning, any training dataset based wholly or in part on current employees is going to demonstrate a bias against women, because there are strictly fewer women. Moreover it's likely that (for a variety of reasons, both intrinsic and extrinsic) fewer women succeed in the interview process as a ratio compared to the number who actually apply.
The difference is that old place transitioned administrative staff to IT roles in the 90s/early 2000s when more things were computerized. Those admins, financial analysts, program analysts were more likely to be female and had degrees in liberal arts, accounting, business/finance, etc.
In the newer place, they filtered based on computer-related degree upon hiring. That automatically excludes many women. Once hired, female candidates advance as well or better.
Anecdotally, I've hired interns in recent years with no tech-specific qualifications as an experiment. If you select for "smart and gets things done" I don't see much of a disadvantage for many roles. You get some duds too, but it wasn't as dramatic a difference as I expected.
I wish companies were more willing to train willing applicants instead of trying to interview for capabilities.
"Demonstrate" may be the wrong word. How about "encode". And in the future, "enforce".
[Edit] On second thought, it would seem like there would be a way to filter out a "raw proportions bias" (like 80% of the resumes in the data set are male) before training.
From the article: "Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs... With the technology returning results almost at random"
I am quite curious about details of the model. For example, the single largest contribution to real world interview process variability is interviewer (for resume screening, who screened that resume, etc.). Wouldn't it be possible to code interviewer as categorical variable and separate resume-intrinsic? effect and interviewer effect? They must have tried this, haven't they?
If there are strictly fewer women in the underlying training set, the model can still return something resembling a uniform distribution of candidates while exacerbating the diminished representation of women.
To give a concrete example: you have a bag of blue dice and red dice. There is a supermajority of blue dice in the bag. Your algorithm selects a single die out of the bag on every iteration. The output sequence of dice numbers appears uniform, but there are more blue dice than red dice in the output sequence.
This view that the only thing holding people back is some sort of social or systemic bias seems to be based on nothing except ideology. Incidentally, it's an ideology I also used to hold. Like a good egalitarian I pushed my wife away from sociology and into majoring in CS. She did perfectly well, as did I. More than a decade later she works with people and I work with code. I've no regrets there, but it's not so clear that my persuasion was really the best idea.
Norway is another interesting example here. It is considered by many to be the most gender equal location in the world. Yet you'll still find that nurses are primarily female, doctors are primarily male, and all other 'stereotypical' divisions present in most all developed nations. They tried to change these divisions and with extensive effort were able to effect a roughly constant change in some fields. But again, once that push was relinquished things went just about identically to as they were in very short order. Ultimately we're flexible enough that you can manage to fit a square peg into a round hole at times, but once you stop squeezing that peg goes back to what it wants to be.
Does she make more money than a similar person with a degree in sociology?
At the least, they should interview a certain subset of randomly chosen applicants or else the feedback loop from the interviewing process and the AI is going to grow tighter and tighter.
I can't imagine this not already being a thing, but I haven't really heard of people using this method.