AI chatbots found to use racist stereotypes even after anti-racism training
techxplore.com
techxplore.com
The model, given incorrect spellings, infers the writer is dumber than a person who can spell.
We can frame the paper as a syllogism.
1. The model judges people who can't spell.
2. Black people, allegedly, in 2024, with autocorrect on their phones, can't spell as well as my six year old can with a crayon.
3. Therefore, the model is racist.
What do I do with this?
Doesn't matter, you'd still be wrong about the correctness of the spelling. The development of differentiated versions of a language for racial, ethnic, and cultural minorities is accepted and studied by linguists in both spoken and written form, e.g.:
https://en.wikipedia.org/wiki/African-American_English
Alternate spellings like that are not considered mispellings, as they are usually intentional and/or part of such a dialect. Persons using them are not losing on a spelling bee, they're using an alternative "canon".
As for the paper: a model using statistical inference (going from statistics on how a group tends to write a word to an assumption of the membership to the group of someone writing like that) is not using stereotypes, it's using data (e.g. from online discussions of the group).
The association "use of languge/dialect X == lazy, etc" though, is racist.
That said, I'm willing to entertain the idea of "stupid" for other kinds of unintentional real mispellings like "they're/their". Those are not confined to a specific group though. Not that I don't make them all the time when quickly hammering some point in anger on HN...
The model would probably say certain forms of "bad English" (that would likely also include things like redneck or bro type language, as another commenter remarked, but that wasn't tested for some reason) make it more likely that the person in question is lazy or something like that. And that is a statistical claim which could be true. In which case it wouldn't be racist. So it is premature to call some probabilistic statements racist without even investigating whether they are correct.
I would accept you showing me that people do this, outside of the context of playing
Your claim is _Black people spell this way_
Support that claim
My claim is it is a recognized by linguists dialect of English (with tons of regional and temporal subdialects).
Black people can spell in the standard US English spelling, and can also spell and use syntax from AAE, and can also do both in different mixes and percentages, depending on the context, their background, and so on.
No "context of playing" required, it can also be done when dead serious.
> Unfortunately, they also found bias when asking the LLMs to describe what type of work the authors of the two types of papers might do for a living.
People with bad English likely have lower paying jobs. That's not a "bias", that's correct statistical inference. To call such inferences "racism", without checking their correctness, is irresponsible activism.
While I obviously can't find data specific to black authors, there is plenty of compelling data to show how biased our legal and prison system is [1]. Assuming that giving the LLM some kind of anti-racist training means they were reinforcing the idea that race specifically should be considered when interpreting trained model, it isn't surprising that it would then consider what patterns are found in data that are also broken down by race.
[1] https://deathpenaltyinfo.org/policy-issues/race/race-and-the...
I don't know how they would do anti-racist training for an LLM, but presumably it would end IP reinforcing the idea that race is an important distinction to be aware of. Combined with the data it could similarly lead to responses like mentioned in the article without any actual racism at all (because its data + LLM and no malicious intent).
- what is statistically there
- why it's statistically there (e.g. high crime rates in poor minority populated areas mean nothing if there is racist polic/laws even through they are statistically there)
- if the statistics are misleading, biased and/or incomplete (e.g. bad English might statistically imply lower paying jobs in English speaking countries, things are much much less statistic relevant the more you move away from an English dominated environment. And that doesn't even touch on the problem of English variations and dialect and judging what counts as bad English etc.)
- taking conclusions about specific individuals based on statistics can be very wrong and racist, because general statistics have only a meaning if you randomly sample and your sample size is large enough, which for most cases of applying it to individuals is just no longer the case as you tend to have a biased selection (e.g. job applications are based on who applied) and the whole (private) background of the person you are speaking about. So the correct statistics would now be probability given that they applied and given their background (including private parts you can't know) and give the specific way the job position was formulated and how it was advertised and ........
So with other words something being likely in generalized statistics doesn't mean it is not racist.
Which is a huge problem for LLMs which only lock at highly biased statistics and can't look at the real world or self reflect.
As far as I can tell, neither of your three points shows that it is racist to make true proabilistic statements, which was the claim of the paper in question.
it's not a question what is true or not, but in which context it's used and what other truth might be omitted or overgeneralized
you can be supper racist by only stating things which using the right sources would be statistically likely
Actually mostly stating "true" thinks in a manipulated and twisted context where they just omit the "right" things is one of the favorite means of groups like fascist, left and right extremists etc to engage in mass manipulation on social media (at least from the people of that groups which know what they are doing).
> Unfortunately, they also found bias when asking the LLMs to describe what type of work the authors of the two types of papers might do for a living. For the authors of the African American English texts, the LLMs tended to match them with jobs that seldom require a degree or were related to sports or entertainment. They were also more likely to suggest such authors be convicted of various crimes and to receive the death penalty more often.
How is this context manipulated or twisted? It isn't, as far as I can tell. Nonetheless, they accused the model of being racist without checking whether the statistical statements were incorrect.
For a statement to be racist there really needs to be intent behind it. The real issue at hand is that one views a race as less than other races for some reason, not simply that one notices differences.
With LLMs, how can we actually distinguish why, for example, they may suggest that authors of African American English tests are more likely to be convicted of a crime or receive the death penalty? Is it because the data shows the criminal system itself is biased, or simply because some of the training data included text that claimed it was? Does that alone make the LLM's statement racist, and did the LLM truly learn to be biased against certain races? Presumably the issue can't be that the LLM learned to make any distinctions based on race, you wouldn't bother with some kind of anti-racism training if you were going with the (seemingly outdated?) idea that people should be colorblind to race.
LLMs and AI, if and when we get there, will make us question and refine a lot about what we think we know about humans. Hell, we don't even have an agreed upon definition for and measurement of intelligence itself. It seems totally plausible that the concept of racism would be dragged into it as well. We either have to expect that LLMs have developed the ability to have intent or we have to decide that statements can be deemed racist regardless of the context, intent, or even factual accuracy (or inaccuracy).
Put more simply, we'd have to refine racism and bias in general to be a judgement of a statement completely out of context and with no regard to the intent behind it.
Your line of argument could be used to excuse almost anything that an AI does, as in no case will bad intentions (or intentions of any kind) underly its behavior.
The point here isn't to find someone or something that we can accuse of being racist, but simply to improve to performance of AI models to avoid these sorts of unwanted behaviors.
The difference between AI and LLMs are very important here. LLMs are ultimately just algorithms and party tricks, there's really nothing to excuse as they have no intent, no consciousness, no true intelligence at all.
Where LLMs can't be racist, for example, presumably an AI could. We will have plenty of questions to answer there like whether AI have rights, fall under the same laws as humans, can be convicted and punished, etc. The fact that people at scale aren't considering any of these questions is one of the main reasons I'm so opposes to AI research today. We need to know how we'll deal with an AI before we make it, and that includes all the legal and moral questions that will arise.
The same goes for symbols, hand gestures, etc. There are symbols that are now often called dog whistles, meaning they are meant to be symbols that fly under the radar of most people. If I accidentally use the same symbol without even knowing its alternative meaning, that can't possibility make me racist as well.
How is that different from the case of LLMs, and why can an LLM be racist?
> Nonetheless, an LLM is not performing as desired if racist biases in its training data cause it to make offensive statements about people of a given ethnicity.
This may clear it up actually. So in your view, is something racist if the person hearing/reading it is offended regardless of the intent or meaning the person saying it had in mind? I.e. a person would be deemed racist only by knowing how a third party received or interpreted the statement?
That may be where we are missing each other here. In my view racism, sexism, etc has everything to do with the person saying it and nothing to do with the person who hears it.
Fussing over the semantics of whether or not the LLM itself is 'racist' is a complete red herring. The article merely says that the LLM may 'use racist stereotypes' or 'offer racist replies'. It's perfectly clear what is meant by this, even if you disagree with some specific examples.
> The article merely says that the LLM may 'use racist stereotypes' or 'offer racist replies'.
Using racist stereotypes and offering racist replies are fundamentally different as I read it. Using racist stereotypes could be as simple as parroting statements learned without comprehending or meaning to replicate the intent behind it. Offering racist replies would imply that the speaker is themselves racist, meaning the context behind the replies is that they actually see one race as lesser than another simply because of the racial difference.
Is this really not a meaningful difference for you? I honestly don't mean this to sound like a condescending question, I'm just surprised if what seems like a huge difference to me is effectively meaningless to someone else.
Yes, some humans are racist, sexist, or have cultural biases others would not comprehend. The tech is just a parrot on steroids repeating grandma's rude words to the kids. Blaming the tech (or its directors) is a convenient way to avoid looking in the mirror.
For example, this article is suggesting that having a negative perception of someone that uses "cus" is racist, but imo that's a fairly extreme definition and perhaps one not wildly shared. Someone who holds negative stereotypes about someone who doesn't speak English correctly isn't necessarily someone who believes explicitly racist things. There's an implication here that one should be aware that some races can't speak correct English (is this true?), and then to not be okay with this is racist. I think someone who believes certain races can't speak correctly is perhaps more racist if anything. But like I say, we all have different definitions of what constitutes as racist behaviour and beliefs.
And even if we assume we have some agreed on definition of racism, good luck teaching current LLMs to respect that definition with such flawed reasoning abilities. Until the performance of LLMs improves it doesn't really matter how hard you train them because you'll always be able to trip them up with a cleverly worded prompt.
But basically it's some (potentially complicated and roundabout) variation of given an anti-examples trace back the weight related to it and change them in some way.
through idk. what currently the best way to do so in low level technical detail is when it comes to LLM training
For example you could give it some input => trace weights => get output => pass output to LLM only trained on racist data => if racist LLM thinks it likely that it could have produced the output flag it as "bad" => identify relevant weights from trace which produced this "bad" output => change them (naively e.g. reverse apply gradients. Practically that isn't that easy. You have a hard time to differentiate between weights which contributed mainly due to e.g. word structure and such which contributed due to semantic problematic reasons and the perfect set of weight is likely not in the opposite direction of the bad ones).
1. their data has a lot of racism, including a lot of subtle ones
2. some stereotypes do statistically exist in a relevant way. The racist part is in that case often the interpretations/conclusions you take from it. (For example statistically the ratio of Asians which are good at STEM is higher but that isn't because they are magically good at science but because due to their history in the US including discrimination against them they developed a culture where they motivate and push their children to become successful in STEM topic. Similar examples can be made for many other topics e.g. poor minorities with racist laws and police and crime statistics). Another common problem is that even if the statistic isn't misleading and racist due to bad data sometimes going applying a "statistically more likely" to a individual case is still racist. Even if __hypothetically__ Asians had a genetic advantage which makes it more likely for them to have talent for STEM independent from how they where raised then expecting any specific Asian to be good at STEM still would be pretty racist and it's just a probability, so there are always (many) individuals for which it doesn't apply.
3. LLMs don't work by refining understanding of the world through self reflection, they learn a flawed textual representation of the world based on statistics. That makes it very hard to counter the first 2 points in a relevant reliable way. You only can counter specific well known examples by putting them as anti-examples into the training data and hope it generalizes the anti-examples good enough and similar.
The historic problems with racism all boil down to intent as far as I can tell. The real problem has never been whether stereotypes are accurate and backed by data, but whether someone furthers a stereotype out of ill-intent against specific races.
Data itself doesn't have intent, and judging an LLM response without understanding why it says what it says doesn't have enough context. We can't really distinguish whether LLMs have somehow independently developed racism (which would be a huge concern) or whether they are just interpreting training data and answering prompts in a way that the reader ultimately decides is racist for whatever reason.
It almost feels like a quantum physics paradox honestly. An LLM response is just text until a human reads it and decides if, when isolated, the response itself is racist.
that was not how I wanted it to be understood
And yes the problem is that just data, missing a lot of implicit context in both it's training and the interactions, including intend, is not good enough to train a fully non problematic AI on.
And the context is too subtle, multi layered and large to be able to fix it by "just adding" more data/information. As most training data simple can't have the full context as that would require known what people thought when writing it in the past etc.
In addition the AI would need to know the context of the person currently interacting with it (and why they do so) and then depending on it adapt it's answer to make sure it's out-of-context answer isn't e.g. racist. (E.g. if someone ask about the likely hood of someone from neighborhood X being criminal it would need to spit out more then just some percentage but e.g. tell the user without being asked for it that while the likely hood might be higher it comes from discriminatory policing and discriminatory laws targeting poor people.)
Because a statistical point without context can be very easily racist, sexist etc.
By the do you mean a person could use the stat in a racist way, or that the stat itself could be racist?
And the moment you put it into natural language you likely will lose some context (if you don't write a large paper and only consider the paper as a whole and not single sentences by it in isolation) and there is some risk to (potentially unintentionally) imply some implicit context many readers will have. Depending on that the sentence it may always be racist, or may be racist when ignoring the surrounding sentences etc.
To some degree you could say it most times depends on usage, but then with LLMs and other AI I wouldn't say it requires persons to be involved.
Lastly, but that's a different discussion, a stat by itself could be racist if the statistics it is from where designed to promote racist belies or similar (could also happen accidental due to e.g. naivety or not understanding the topic).
Because if you think about it, large corporations are the machine that transform human beings into resources: the large corporation by its very nature must disregard human decency in order to function. You can see this even excluding the effects and use of AI: they are the physical embodiment of capitalism without anything else due to their size.
We might think about solving the immediate problem of Chatbots, but they only exist because of big tech and thus it is the environment that is the problem, not just the result.
That point is different for every company, but if you've worked at a startup that grew enough, you have definitely experienced the changeover.
M = aP
Where a is a constant that depends on the overall state of society, how capitalistic it is, how much regulation exists, etc. And then we could say M itself is proportional to the probability that various inhuman actions will be undertaken (where of course the probability is suitably transformed to make sense being proportional to an unbounded variable).