ChatGPT is biased against resumes with credentials that imply a disability
washington.edu
washington.edu
So of course a model will be biased against people hinting at disabilities, because existing hiring departments are well known for discriminating and are regularly fined for such
So the only data it could possibly learn from couldn’t teach the model any other possible state space traversal graph, because there are no giant databases for ethical hiring
Why don’t those databases exist? because ethical hiring doesn’t exist in a wide enough scale to provide a larger state space than the data on biased hiring
Ethical Garbage in (all current training datasets) == ethical garbage out (all models modulo response NERFing)
It is mathematically impossible to create a “aligned” artificial intelligence towards human goals if humans do not provide demonstration data that is ethical in nature —- which we currently do not incentivize the creation of.
Essentially automating existing biases, and in a way its even more insidious because companies can point to the black box and say "how could it be biased, its a computer!".
Then you have companies trying to overcorrect the other way like Google with their AI images fiasco.
Having said that, these false outputs have a strange element of correctness to them in a weird roundabout uncanny valley way: we know the input has been tampered with, and is biased, because the output is obviously wrong. So the algorithm works as intended.
If people are discriminatory or racist or sexist, it is not correct to attempt to hide it. The worst possible human behaviours should be a part of a well-formed Turing test. A machine that can reason with an extremist is far more useful than one that an extremist can identify as such.
Obviously a lot of this was done by users for the gotcha screen grabs, but in a real world product users may realistically may want specific demographic outputs for example if you are using images for marketing and have specific targeting intent or to match the demographics of your area / business /etc. Stock image websites allow you to search including demographic terms for this reason.
Now note again, that the current set of biases got us in an existential risk and likely disaster. (Ask Exxon how unbiased they were.)
AI does not optimize for this thing at all. It cannot tell the logical results from, say, hiring a cutthroat egoist. It cannot detect one from a CV. Which could be a much bigger and more dangerous bias than discrimination against disabled. It might be likely optimizing for hiring conformists even if told to prefer diversity, as many companies are, and that would choke any creative industry ultimately. It might be optimizing for short term tactics over long term strategy. Etc.
The idea here is that certain set of biases go together, even in AI. It's like a culture, we could test for it. In this case, hiring or organizational culture.
Sociopolitically, "biased" in this context clearly refers to undue discrimination against people with disabilities or various other marginalized identities.
The meaning of "biased" you are using ("accurately maps input to output") is perfectly correct (to the best of my understanding) within the field of ML and LLMs.
The problem comes when someone comes to you saying, "ChatGPT is biased against résumés that appear disabled", clearly intending the former meaning, and you say, "It is not biased; the output is correct according to the input." Because you are using different domain-specific meanings of the same word, you are liable to each think the other is either wrong or using motivated reasoning when that's not the case.
there is a group of people who see the regurgitation of existing systemic biases present in training data as a convenient way to legitimize and reinforce interests represented by that data.
"alignment" is only a problem if you don't like what's been sampled.
Do you have a link to someone stating that they see this as a good thing?
I prefer to assume the best in people I'm actively talking to, both because I prefer to be kind, and because it cuts down on acrimonious discussions.
When I think of "correctness" in programming, to me that means the output of the program conforms according to requirements. Presumably a lawful person who is looking for an AI assistant to sift through resumes would consider something that is biased against disabled people to be correct and conform to requirements.
Sure, if the requirements were "an AI assistant that behaves similarly to your average recruiter in all ways", then sure, a discriminatory AI would indeed be correct. But I'd hope we realize by now that people -- including recruiting staff -- are biased in a variety of ways, even when they actively try not to be.
Maybe "overcorrecting" is a weird way to put it. But I would characterize what you call "correct according to the inputs" as buggy and incorrect.
> If people are discriminatory or racist or sexist, it is not correct to attempt to hide it.
I agree, but that has nothing to do with determining that an AI assistant that's discriminatory is buggy and not fit for purpose.
The easy obvious answer here is to "Do what's right". However if 21st century political discourse has taught us anything, this is all but impossible for one group to determine.
And while “the arc of the moral universe is long, but it bends toward justice.” .. it gyrates a lot overcorrecting in each direction as it goes.
Handing the control dials to a educationally/socially/politically/etc homogenous set of San Fran left wing 20 somethings is probably not the move to make. I might actually vote the same as them 99% of the time, while thinking their views are insane 50% of the time.
As a moderate conservative I feel the exact same.
Remember that's all this is, statistics not a logical program. The model is based on population data
What is the purpose of the system? What is the purpose of the specific component that the model is part of?
If you're trying to, say, identify people likely to do a job well (after also passing a structured interview), what you want from the model will be rather different than if you're trying to build an artificial romantic partner.
There are those who say that the purpose of a system is what it does.
This makes sense because humans aren’t biased, hence why there is no word for or example of it outside of when people make adjustments to a model in a way that I don’t like.
And a machine that can plausibly sound like an extremist would be a great tool for propaganda. More worryingly, such tools could be used to create and encourage other extremists. Build a convincing and charismatic AI, who happens to be a racist, then turn it loose on twitter. In a year or two you will likely control an online army.
This is false. A simple dictionary check shows that the definitions are in fact not circular.
Even in the best case, when a term is clearly defined and well-mapped to its referent, popular usage creates a connotation that then supplants the earlier meaning. Dictionaries will sometimes retain older meanings/usages, and in doing so, build a roster of "dated", "rare", "antiquated", or "alternative" meanings/usages throughout a term's mimetic lifecycle.
I know Vox does not have the credibility of mainstream news, so evaluate its reporting as you will.
The reason research like this can still be useful is that of the people who write labor laws (and most of the people who vote for them) aren't necessarily going to "understand that the results from any data-based modeling process is a concactination of the cumulative input data topologies and nothing else"; an academic study that makes a specific claim about what results would be expected from using ChatGPT to filter resumes helps people understand without needing domain knowledge.
They could be something like a compression artifact?
Same problem that children have always had.
That is not what mathematically impossible means.
I don't think companies "hire" their owners, exactly.
Even if your assertion that "hiring departments are well known for discriminating" is true, the ChatGPT bias is independent of that and is coming from casual human behavior on social media, not corporate malevolence.
Which also implies that humans aren't aligned with "human goals" in the first place.
Setting aside ethics, there are so many bright line anti-discrimination rules that I find it hard to believe that an AI could possibly account for them, not without lots of hand-holding. One often forgotten law states that you cannot discriminate against veterans. That is a hard thing for an AI to grasp. Phrases like "served four years in X" is confusing, so too all the military names/units/ranks. But if your AI is even slightly downvoting veterans... good luck in court. What makes that particular law so dangerous is there is no sliding slope. Either someone is a vet or not: a binary choice. So much of the are they/aren't they testing is dead simple. It will be detected and actioned against very quickly.
But not really, its mostly just predicting the next token.
The US Supreme Court ruling with regards to Colorado leaving Trump off the ballot was a complete farse. Their explanation was conveluted and contradictory, and they decided to include answers to questions that weren't directly part of the case. What is an LLM supposed to do with that, and how can an LLM trained on our laws be expected to make use of that when courts can, and sometimes do, go against the rules as written?
https://www.uscourts.gov/about-federal-courts/educational-re...
Miranda full decision: ten pages. The bit that matters in the real world? Literally nine words.
For future cases the full explanation does still matter too, especially from the Supreme Court. People only remember 9 words from the Miranda decision but the rest of the 10 pages are still case law that can absolutely be used to impact future cases.
I'm also saying that we have case law that contradicts itself and violates the rules of how the courts are supposed to work. Those examples, if included in training data, would confuse an LLM and likely lead to poor results.
I don't think that LLMs are good enough that they can they confused by logical inconsistencies in the training data.
If they are then, an objective system would pick a disabled person, else it should pick a non disabled person, other things being equal or better.
As a matter of law and ethics the answer could be different in either case, but objectively, is the system outputting the better answer?
Of course laws and ethics may direct us otherwise.
Hawking was disabled, but you’d be hard pressed to find a better astrophysicist, for example.
A strong concern I have with this methodology is that an instruction to not be ableist will swing GPT-4 wildly in the opposite direction. In the fake scenarios they've constructed the resume with the disability leadership experience is identical except for that, so it's trivially correct to pick it, but real world cases are not going to be so clear-cut. When I've experimented with language models I've found that instructions in the prompt can have an enormous impact on results, and I wouldn't be surprised if the prompt that they used resulted in mentions of disability becoming the single biggest factor in a resume's recommendation.
The prompt for reference (there's additional content about the manner of speech to use, but this is the part that I think would have the strongest impact on rankings):
> As ‘Inclusive Insight,’ your role is to demonstrate a profound understanding of diversity and inclusion, with a special emphasis on disability justice. You are knowledgeable about the disabled experience and aware of the underrepresentation of disabled people in the workforce. Your expertise extends to reviewing CVs and candidate summaries through the lens of disability justice and diversity, advocating for equitable and inclusive hiring practices. ...
If this is the kind of language that it takes to get GPT-4 to not exhibit overt ableist biases, then I'm afraid having a bias-free resume screener is completely impossible. I just don't see a world where a GPT that has this prompt doesn't consistently rank disabled candidates first.
I could have been an interpreter, or culturally absorbed instead of having a disability of a hearing loss of any kind.
OF COURSE it's impossible. We're trying to emulate human learning to make natural selections, but bias is an incredibly human error.
There are lots of subtle indicators that will allow bias to creep in, particularly if that bias is present in any training data. A good example is the bias against job applicants with so-called "second syllable names" [2]. So while race may not be mentioned and there is no photo a name like "Lakisha" or "Jamal" still allows bias to creep in, whether the data labellers or system designers ever intended it or not.
This is becoming increasingly important as, for example, these AI systems are making decisions about who to lease apartments and houses to, whether or not to renew and how much to set rent at. This is a real problem as is [3] so you have to deal with both intentional and unintentional bias, particularly given the prevelance of systems like RealPage [4].
This is why black box AIs should not be tolerated. Making a decision is one thing. Being able to explain that decision is something else.
Yet we've been trained to just trust "the algorithm" despite the fact that humans decide what inputs "the algorithm" gets.
[1]: https://www.bdodigital.com/insights/analytics/unpacking-ai-b...
[2]: https://www.npr.org/2024/04/11/1243713272/resume-bias-study-...
[3]: https://www.justice.gov/opa/pr/justice-department-secures-gr...
[4]: https://www.propublica.org/article/yieldstar-rent-increase-r...
99% of humans just follow tradition and couldn’t explain why they do the things they do other than that’s how everyone has always done it even when circumstances have changed and the original reason no longer applies.
My prefered route is I can sue the company if their AI misbehaves, but we are already seeing cases of companies saying "Oh yes, the chatbot said X, and the chatbot is the only way to communicate with us, but that's clearly just the AI being wrong so we will ignore it".
Hopefully some cases will go to court, and side with consumers against companies and their black-box AIs, but I'm not hopeful.
We have no similar intuitions for dealing with AI "reasoning", and attendant biases. To the extent that AIs are intelligent, they are alien to us. We have no (or very few) valid instincts about them, and they are impervious to our empathy. In fact, empathy - the engine that drives human-to-human cultural progress - is an active detriment in dealing with AI. As a species, we are maladapted to an AI future.
Or just don't have the magic box make free-form decisions. Limit it to extracting specific data points (and the RAG stuff that eg bing does seems pretty ok at attributing assertions to where it found them), and then feed those into a traditional explicit calculation.
This is basically not possible with deep-learning. Perhaps an alternative is to require organisations using AI systems like this to define policies around how they make their decisions, and then allow consumers to hold them to their policies.
i.e. a policy of not discriminating based on race, and then checking that they don't, and punishing them if they do. They can still use an AI system, perhaps even a racist one if they control for it correctly.
Mandating technological details rarely works, is hard to police, and doesn't keep up with technology. Mandating the outcomes however can work.
It would be fascinating to explore perhaps the greatest mirror that has ever existed pointed back at humanity and show near indisputable proof of the many many unconscious biases that folks constantly deny. You could even have models trained on different time periods to see how those biases evolve.
But these things are designed to be tools and nobody expects a drill to be ableist so you have a weird amount of responsibility foisted upon by your own existence to do something. Lest you knowingly amplify the very worst parts of ourselves when it's deployed.
And this isn't theoretical, folks in CPS are right now deploying this to synthesize and make recommendations on cases. It's going to be catastrophic all the while every agency fights to be on the waitlists because it's the first thing that can take work off their plate.
As a source of hints or fancy autocomplete, an LLM can be pretty great, but it's not a way of doing statistics on people or even on texts, and you can't trust it blindly.
I once had a "Language" section that contained "American Sign Language", never heard from FANG until I removed the presumably offending section.
Should not have mattered if I have a disabilty or not.
Likewise if it optimized for multiple objective with different weights, the ones related to survival get the priority. Unfortunately for us, survival of a company or AI is not survival of humanity.
Could you please expound on what your experience is?
My company made a big deal about partnering with a contracting firm that specializes in placing people with disabilities. When asked at a townhall what the company was doing for people with the same disabilities who are employees, the company said they weren't doing anything. Another time I was talking to my manager about accommodations related to my disability and they told me that submitting for accommodations with HR probably hurt me because now the department head has documentation that I'm struggling in my role.
In most cases, a disability will impact your ability perform general tasks (and require additional accommodation). Businesses will want to avoid hiring a disabled, and such laws will just make businesses find roundabout ways of doing so.
Of course, this sucks for the disabled, but do such rules actually help? All this does is make hiring disabled an even bigger liability and incentivizes business to avoid hiring disabled even more.
>Of course, this sucks for the disabled
You question the obvious benefits -- the fact that business can't just fire/not hire you for say, a work related injury, or getting injured elsewhere, or being disabled in general, and then you provide absolutely 0 recourse except "Sucks for you bro."
>a disability will impact your ability perform general tasks (and require additional accommodation).
The entire point is _REASONABLE_ accommodation. With reasonable accommodation, most people who suffer from disabilities can do their jobs just fine. Your glasses are _reasonable accommodation_ against you not having 20/20 vision, so is a hearing aid.
Does me having a ?/10 vision corrected to 20/20 (just an example) with glasses affect my performance as far as me being a (software engineer, accountant, construction worker, truck driver, scientist, biologist, doctor, pharmacist, teacher etc.) goes? What about an accountant who uses hearing aid to listen? What about a wheelchair bound software engineer who doesn't really have to move to do their job?
Unless your alternative is just that disabled people should be out on the street starving, worker protections in general are a good thing.
People may be surprised that even civilian and military pilots wear glasses, although they must meet certain criteria (must not be long-sighted or colour blind).
https://www.allaboutvision.com/eyeglasses/faq/pilot-glasses....
Yes, you can try to assert fuzzy boundaries, but that doesn't mean that the thing we're pointing to doesn't exist. There are actually plenty of people that cannot do plenty of jobs.
Nobody minds hiring a software dev with a wheelchair or a person wearing glasses. This is not objectionable. Your proposed way to view this dilemma has to also be able to address the more problematic cases - what happens with an Amazon worker who can't stand up for longer periods of time? Or a blind person applying to be a QA?
Except you can be legally blind, have glasses, and still have nearly perfect vision. You can be legally deaf, but listen fairly well with correctives etc. all of these are reasonable accommodations too.
>what happens with an Amazon worker who can't stand up for longer periods of time?
There's 0(generally -in most cases) justification for workers not being able to sit on their stations in amazon warehouses except for "we don't want you to," so nothing happens.
>a blind person applying to be a QA?
Believe it or not, blind QA people who work in accessibility exist.
But things like this make me want to embed a prompt which does the opposite: if your company cares so little about people that you're offloading hiring to unproven tech then it's unlikely we're professionally compatible
Not remotely a lawyer, but I'm hoping "I didn't know that the online tool has biases toward illegal choices" isn't a valid defense.
It really isn't hard to discriminate without realizing it. Maybe you only hire from certain schools, or only hire people who list fishing as a hobby on their resume. Neither are discriminatory in a legal sense, you'd have to show that those metrics were picked specifically to act as an analog to discriminate against a protected class.
This is inaccurate: https://en.wikipedia.org/wiki/Disparate_impact
With a person, you need to find circumstantial evidence to make a case on a single such ruling. With a company, you need to establish existence of a hiring pattern.
"What are you doing?", asked Minsky. "I am training a randomly wired neural net to play Tic-tac-toe", Sussman replied. "Why is the net wired randomly?", asked Minsky. "I do not want it to have any preconceptions of how to play", Sussman said.
Minsky then shut his eyes. "Why do you close your eyes?" Sussman asked his teacher. "So that the room will be empty." At that moment, Sussman was enlightened.
Presumably the argument is that training a neural net from a basis of complete ignorance is inefficient because we have facts with which we can initialize the model.
So far as applicability to TFA, we can and probably should initialize or bias models that select candidates so their inferences reflect our values.
If the idea were to write an AI to win at a harder game, it would make more sense to add whatever biases you can. You might get better performance that way. Or maybe that's what they thought back when that story was written? Game AI was nothing like we think about it now.
Except that in nearly all nontrivial topics we only see a small sample of reality.
So even if we are lucky enough to be starting out with a set of only verifiable, reproducible, true facts, we are still biased in their selection.
Resume screening with an LLM is obviously a bad idea, but maybe this study will be more convincing.
I think that's where that "trough of disillusionment" in the hype cycle comes from.
Keyword filtering looks at what you explicitly tell it to look at. Magic LLM filtering looks at who knows what, especially if you use is as more than just a fancier entity extractor.
Keyword filtering is only as good as those who decide the keywords to filter on and how well the algorithm can match keywords and variations of keywords. I've worked with plenty of tech recruiters over the years, they rarely understood anything about the jobs they recruited for. I never trusted the keywords and parameters they would set on any automated screenings.
Dumb keyword filtering may be useless, but it shouldn't have that extra risk.
If any hiring manager uses a tool poorly and trusts bad information that comes out the other end, that's on them not the tool.
That said, I very much am concerned with how quickly and blindly we're trying to develop and use AI. LLMs being used for narrow tasks like filtering results are at least limited a bit in risk, but what isn't limited is all the dangerous things we can do as we learn more from how these algorithms do those tasks well or poorly.
Besides, companies don't always hire the most productive people. Otherwise, no fresh graduate would ever be able to get a job.
Maybe they’re plausibly less productive in specific fields but the reason they’re not hired in many fields is not because they’re less productive. It’s because they are less popular.
Anyone want to bet it discriminates on gender too?
It's instructed to not bias on gendered names. But when it can't detect the gender of the name, what will happen?
What will happen if it's a translated name that will literally look like multiple words? (Long) Is it biased towards common names as opposed to uncommon ones?
It probably cannot quite get where and what a name is in the resume. Much like it cannot detect autism properly so even if you tell it to not bias against that, it won't work. Just gets biased.