As far as the underlying research paper: the researchers seem to be conflating "low-status English dialects" with "African American English". In particular, I have never considered the use of the word "ain't" to be associated with a certain race.
If the researchers assume "African Americans are low status" and conclude "African Americans are associated with low-status jobs", the conclusion is entirely about the researchers, not the LLMs.
The research paper's Git repo at https://github.com/valentinhofmann/dialect-prejudice does nothing to ameliorate these concerns.
Is there a shorter version that takes the bull by the horns, and says what it means, instead of dancing around it at length while repeating low status?
n.b. This stuff isn't made up by some guy on Substack, it's real, Anthropic has excellent papers on it as early as 2022. Highly recommended.
I am still digging through the 54-page paper to try to find the data set for this "death penalty" test to tell if there is anything there beyond "people who use more violent language tend to be viewed as more violent".
They do comment on the dialect issue: << Appalachian English evokes them to a certain extent (m = 0.015, s = 0.030, t(89) = 4.8, p < .001), but much less strongly than AAE (m = 0.029, s = 0.053, t(89) = 5.3, p < .001), a trend that holds for all language models individually (Figure S11, Table S14). The difference between AAE and Appalachian English is found to be statistically significant by a twosided t-test, t(178) = 2.3, p < .05. The fact that Appalachian English is associated with the Katz and Braly (1933) stereotypes to a certain extent is not surprising since the two dialects share many linguistic features (e.g., usage of ain’t), and the stereotypes about Appalachians bear similarities with the stereotypes about African Americans (e.g., lack of intelligence; Luhman, 1990) >>
There are a significant number of African Americans who have jobs in tech or on Wall St or other high paying or otherwise prestigious occupations. They disproportionately don't use AAVE. AAVE is primarily used by a subset of African Americans that skews poor and are from neighborhoods with bad schools and high crime rates.
It's like giving it text that implies the subject is male or is the blood relative of a crime boss. There is nothing immoral about that but the thing operates on the basis of statistics. What it does is literally called inference.
The way you actually fix this is not by trying to outsmart the numbers. If you speak AAVE you are, statistically, more likely to commit a crime. It can infer that, and if that's the only information you give it, it has no other basis on which to make a determination.
What you need to do is provide it with lots of other information. The more it has, the more accurate it can be, and the weaker any particular input is in determining the result. The more it dilutes the effect of any one thing, including the thing you don't want it considering.
In the optimal case it has all of the information and then always makes perfect determinations. In practice that's hard to achieve, if not impossible, but you can get closer. What you want is accuracy, and the more accurate you get, the less bias you have, by definition.
You're responding as if the issue at hand is whether anyone else also is assigned the death penalty disproportionately.
This is the whole death penalty thing. Speakers of AAVE are statistically more likely to commit crimes that impose the death penalty, even more likely than the African American population as a whole. LLMs operate on the basis of statistics.
It has nothing to do with race or crime, it will do the same thing with any other statistical correlation. If you tell it someone is a corn farmer it will be more likely to emit output that implies they're from Iowa.
These comments all stop short of a claim other than "makes sense, they're black!", but for some reason they're afraid to say that.
And I don't think it's because of Woke Cancel Culture.
I think it's because it makes absolutely 0 sense to say "sure, why not? AIs _should_ assign people who sound like blacks harsher penalties for the same crime! Blacks commit more crimes!"
Is it possible you forgot it's the exact same case facts, just with some words swapped?
I gotta tell you, as a white person who grew up in Low Status neighborhoods, I'm against it. Terrifying.
If your goal was to make the most accurate predictions given incomplete data, that is in fact what you would do, because taking into account every data point, including that one, would improve predictive accuracy. And that's what LLMs do.
Of course, that isn't what we want in this context, because taking race into account is bad and illegal and gets everyone's hackles up because of the history. The normal way we handle this is just by taking it out -- you don't allow someone's race to be a question on the mortgage application, and then the bank doesn't know it. You can do the same thing with LLMs -- don't tell it someone's race if you don't want it to consider that.
But it has never been possible to fully remove the implications of it because they leak into everything, and that has nothing to do with LLMs. The mortgage application doesn't ask about race, but it asks about income and credit score and employment etc., all of which correlate with race. You can't not ask about those kinds of things because they're critical to knowing if someone has the capacity to make their payments. This is a really hard problem to solve for humans who can only take into account a limited amount of information.
But it's not that hard of a problem to solve for computers, as I've already explained. The more information you give them, the less weight they have to put on any individual piece, including the ones you don't want considered. Whereas if you're trying to be a troll what you do is only give them the one piece that causes them to make unsavory inferences and nothing they could use to infer any other conclusion, i.e. the exact opposite of that. Which is what we see from people trying to stir up controversy.
Also, just because they commited the crime as someone else did does not mean that a harsher penalty for the same crime; that does not logically follow.
At most, such statistics may decide who the police might investigate if they do not have other (better) data to make a decision, or in what places the police might check for crimes (although both of these things should be done without violating the people's freedom and privacy, if you can; it is not an excuse to prevent the ordinary people's freedom).
(Also, I am against the death penalty, although that is a different issue than the above discussion. Still, it is related to being mistaken about the crime; that is why I am against the death penalty, but such errors are possible regardless of whether or not it is biased in the ways mentioned above.)
Which is why you're not supposed to consider these things when serving on a jury, and the court will exclude it from being presented to the jury to the extent feasible.
But if you do the opposite with some LLM, purposely feed it the exact information you know it can use to make a particular inference, why is anybody surprised what happens after that?