If (for example) 66% of Doctors are male and 34% female then it's not reproducing "existing structures of oppression" it's inferring something about reality.
If (for example) 66% of Doctors are male and 34% female then it's not reproducing "existing structures of oppression" it's inferring something about reality.
And if you think that people won't use the idea that the outputs are unbiased because the computer isn't programmed with the same prejudices that produce the inputs, I have some algorithmically-generated investment advice involving a bridge to sell you
That's fine but it isn't the goal of these algorithms. It isn't the reality that is useful for them to learn. It's a different problem to try to build some kind of "unbiased" ontology rather than just to learn about words. Feel free to research or create solutions to this other problem, it sounds interesting.
Suppose, for example, that I gave this same statistic to someone and then asked them to select from a pool of 100 applicants for 50 available places in medical school. Let's assume that there's an equal # of male and female applicants and that their exam results are all similar. Do you think that knowing about this 66-34 split might influence the gender balance of the final selection?
The whole point of training and using machines is to make more accurate, more useful decisions in a complex world.
That can't happen if we give them data that isn't borne out by reality, or tell them to ignore data that is.