> The problem with erasing biases is that you cannot look at any statistics. The internet and training set can be free of any form of -ism, and the models would still be expected to have biases. In fact it's something desirable, because statistical inferences are a valuable tool.
I’m sorry but I can’t let you get away with this terrible argument and conclusion. No one argues for completely erasing bias (especially the scientific form of the word bias), that’s a strawman.
Strong proponents argue that we should all be aware of our biases, and attempt to adjust our opinions and behavior according to the results of that exercise of self-reflection. Stronger proponents might even argue that the inability to perform this exercise of self-reflection is a path to bigotry.
Being racist AF isn’t something that you can excuse with “statistical inference”, and your comment sounds like it’s flirting with that concept. It’s the intellectually juvenile pseudo-philosophy that the techbro scene is absolutely riddled with like a malignant sexually transmitted infection, all the way up to Mu$k and Thi€l.
Back to LLM world, the issue is that there is no diversity in its bias: one LLM, one bias. If everyone uses the same dozen or so state-of-the-art LLMs, then all of our processes will have the same dozen or so biases. That would kind of suck if you were a member of a group that those LLMs happened to be biased against. LLMs are also famously not capable of self-reflection, barring the Rube Goldberg machines that people have built on top of them to simulate thought processes.