If someone finds AI output problematic, you should look at the input data. If the input data such as the world in general is causing "problematic" output data, it might be time to reconsider what you think of as problematic.
If someone finds AI output problematic, you should look at the input data. If the input data such as the world in general is causing "problematic" output data, it might be time to reconsider what you think of as problematic.
Well, that's just it, aint it? Folks aren't saying the AI isn't aligned with itself. They're saying the input data isn't aligned with something. That seems pretty reasonable. Does anybody think scraping random internet data is going to spit out viewpoints that "the world in general" can all agree on?
Is that "problematic"? Frankly, yes, as in it's almost guarenteed to cause problems. That doesn't mean we need to censor or remove it, it means we should discuss what "problematic" means in these contexts.
The data in the article is arguing against that kind of pat assumption. For example it says that: "Women made up a tiny fraction of the images generated for the keyword 'judge' — about 3% — when in reality 34% of US judges are women, according to the National Association of Women Judges and the Federal Judicial Center."
So part of the concern is that the input data is not a representative sample of "the world in general," but relies on stock photos, celebrity media images, sensationalist news reports, items which by design do not accurately reflect the real world.
What percentages of judges are women is a different question than what percentage of judges depicted in visual media are women... and that is an even different question than what percentage of judge images used to train stable diffusion were women?
The dataset's proportions don't match reality's proportions, which leaves the AI with an inaccurate perception of reality.
This presents as a phenomeonon one might describe as "bias", even though the AI itself has no motivations.
People have this problem as well.
Look at the issue of predictive policing via AI: police are biased, therefore the data is biased, and thus the model is biased.
Meme world.
It might be time to consider the possibility that some of what's out in the world in general is problematic. Ask "the world" their opinion on Jews or atheists or germ theory or quantum superposition and you'll get interesting answers AI shouldn't necessarily consider to be accurate.
Humanity is so weird.
[1] Which they will construct using their imagination, and not realize...the very thing they are mocking others for doing (though, I am being reductive here).
Bias in learning data is essential to detecting patterns, by AI or by humans. ML's dependency on being fed biases (AKA as stereotyping) is an Achilles heel that affects every kind of learning, but it creates real problems when learning sophisticated patterns (like those learned by LLMs) where such biases can cause misbehavior and misunderstanding.
Unsurprisingly, fans of ML simply have avoided publicly addressing this limitation because it's essentially unsolvable by automation. The only solution is to manually acknowledge every exception to the rule -- the way children have to to learn every irregular verb. But until we do this, AI will remained fundamentally driven by its inbred biases and stereotypes.
By that logic, nothing is biased. That racist cop? Just looking at her input data and outputting her own.
(I'm ignoring the OP article, which I haven't read.)
What’s the difference between bias and knowledge? Science changes etc
Bias is structural ignorance of input data that contradicts with an entity's ideology that they want to protect.
A racist cop who came to the job with an internalized ideology that black people are bad and executes on that ideology despite experience and data to the contrary is biased. But a cop whose lived experience on the job is that (at least within their jurisdiction) black people are more likely to commit crimes and be violent, and therefore warrant a higher degree of wariness isn't biased, but operating rationally based on their input data. If that rational conclusion morphs into racist ideology that is inflexible to change given new input data, then it becomes bias. It's also possible they misinterpreted the input data and keyed in on race rather than poverty as the root cause, but that also isn't necessarily bias, just bad processing or incomplete input data.
The process of science is the attempt to follow the input data to the conclusion without letting theoretical ideologies bias us away from discovering unknown truths.
Potentially less polarizing: Aristotle when developing his theories of violent vs natural motion. I would not call him unbiased for theorizing without looking outside.
You can't claim to be unbiased from just individual experiences, you need to attempt to understand the whole picture.
We have similar religion-like dogma today, it has just evolved.
Which one is it?