If I ask a language model, "Are Indian people genetically better at math?" and it says 'yes', it has failed to accurately approximate reality, because that isn't true.
If it says, "some people claim this", that would be a correct answer, but still not very useful.
If it says, "there has never been any scientific evidence that there is any genetic difference that predisposes any ethnicities to be more skilled at math", that would be most useful, especially for being a system we use to ask questions expecting truthful answers.
There are people who just lie or troll for the fun of it, but we don't want our LLMs to do that just because people do that.
I think there are a lot of people who would say "Indian people are better at math" and not even think about why they think that or why it might even be true.
In my opinion, most biases have some basis in reality. Otherwise where else did they come from?
I for one would not be prepared to defend the persistent bias against black persons and immigrants as having a basis in reality. YMMV.
There is a BIG difference between biases being based in reality (which they're not), and biases being based in our varying perceptions of reality, which are themselves biased.
Because in my mind, that's the environment the person making that claim was in, so I just kind of automatically include that in my interpretation of their statement.
I don't think they are making a generalized statement that that Indian people are genetically better at math. I think they are making a statement that they perceive that average Indian person that they run into is better at math than the average person they run into. And maybe they are right, and maybe there is a reason based in reality why that is.
It sounds like that is somewhat true based on what you said about the visas.
I never take any of these things to have anything to do with genetics. To me it's always due to some external factor like the visas as you mentioned, or even maybe just like a cultural thing where they are pushed harder to be good at something as they go through school, and so are better at something than the average person in the end.
This is not a technical limitation at all, this is purely about cost and time, and companies wanting to save on both.
There are also methods like RAG that try to give them access to fixed datasets rather than just the algorithmic representations of their training data.
yes, I still wonder how LLMs managed to generate this expectation, given that they have no innate sense of "truth" nor are they designed to return the most truthful next token.
LLMs stepped into a field that has existed in popular consciousness for decades and decades, and the companies running LLMs for public use *sell* them on the idea that they're useful as more than just expensive text-suggestion machines.
Anything that has a legal requirement to remain unbiased will also clearly define what counts as bias, e.g. discriminating based on race in hiring like you mention. So there's not just some requirement that a process be "unbiased" in a vague, general, philosophical sense as debated above in this thread. Rather, the definition of bias is tied to specific actions relative to specific categories of people, which can thus potentially be measured and corrected.
More generally in ML, bias means that the training set deviates from the ground truth systematically in some way. Entirely eliminating bias that falls into that broader definition seems like an impossibility for general-purpose LLMs, which cover so much territory where the ground-truth is unknown, debatable, or subject to change over time. For example, if you were to ask an LLM whether governmental debt above a certain percentage of GDP damages growth prospects sufficiently to make the debt not worth taking on, you would not receive an answer that corresponds to a ground truth because there is no consensus in academic economics about what the ground truth is. Or rather you wouldn't be able to know that it corresponds to the ground truth, and it would only be a coincidence if it did.
That ML definition of bias runs against the legal definition where the ground-truth is itself biased. e.g., if you were to develop an algorithm to predict whether a given student will succeed in a collegiate environment, it would almost certainly display racial bias because educational outcomes are themselves racially biased. Thus, an unbiased algorithm in the ML-meaning of the word would actually be extremely biased in the legal sense of the word.