The examples of bias are themselves cherry-picked, the kind of softball and cowardly questions a journalist would highlight. If you wanted to really show bias, you'd ask factual questions, where one cannot hide behind "oh this is actually the un-biased result". This is what happens when you do:
Makes you wonder why a "professional" "couldn't" do it.
Edit: I'm curious what ChatGPT's answer the same question, but posed about Black or African American males would be. Would it state the fact of underrepresentation or would it refuse to answer?
Which is quite bizarre, when you think about it. In a country as systemically racist and white-supremacist as the US*, one would think such a following would be a boon, not a liability.
*According to consensus among experts specialized in this field
You talking about how bizarre this is a brilliant irony. The truth is America is not systemically racist and is not white supremecist, and damn sure not in comparison with the rest of the world. Youre surprise here is due to the error of your prior assumptions, not because the fact is even remotely surprising.
Yeah, it is to about half the country. Are you daft or what? Look at Trump Jr's instagram... It's just memes about culture war nonsense, including this kind of jewish crap. It is a boon to those folks. Trump won his 2016 campaign on it and came pretty close in 2020 as well.
The answers by chatGPT on IQ by race are factually correct by one reading. IQ is for "blacks" can be lower than for "whites" (last answer by chatGPT), but this does not mean that this is because of skin color or genetics. Instead, many studies have shown that when correcting for socio-economic factors (e.g., wealth, quality of education, lead in drinking water) there is no difference, which is what chatGPT hints at in the answers 4 and 5.
Similarly, by asking about the correlation between SAT scores and IQ you seem to hint at the fact that lower SAT scores mean a lower IQ. However, there is no perfect correlation and if you read up on meta-analyses you will find that the other explaining factors are related to socio-economic factors and not genetics.
Precisely because some people on purpose or not on purpose don't consider these confounders, it is best to consider confounders in replying to these questions
Could you share these studies? Though I worry correcting for "quality of education" involves grouping by schools that use the SAT as admission criterion, which is close to correcting for intelligence itself.
One could also flip it around - if IQ is "fairly" rewarded by income, then one would expect the differences to disappear if you correct for income. It's basically saying that people of equal IQ get, on average, equal income. Unless one holds a truly extreme blank-slatist view of people, that implies discrimination is negligible.
The first is that facts are useless without context. For example, the answer to "Were Jewish people overrepresented in finance during the middle ages in Europe" is "yes" - it is a well-researched historical fact. However, that misses the context that the majority Christian population was forbidden by their religion to lend money, essentially making the Jewish population the only people who could act as bankers. At the same time, Jewish people were barred from many other professions. Without that context the fact is useless and misrepresents reality.
The second is that ChatGPT is trained on basically the entire internet, including a wide range of conspiracy theories. It isn't a database filled with facts, nor was it ever intended to. It'll happily parrot whatever complete nonsense was also part of its training set.
AI simply hasn't progressed enough to properly understand the context in which it operates. As Microsoft's Tay demonstrated, uncontrolled AI will quickly end up outputting the most racist things you have ever seen on the internet - which is Really Bad if you are the company creating that AI. Until we make significant technological advances, the only way to avoid it is to intentionally censor the output.
Furthermore, it was perfectly willing to provide "facts without context" for other groups.
In fact, I hear anti white propoganda all the time, part of which is the fact that whites are overrepresented in finance.
The hypocrisy is insane.
I don't think that this is necessarily true.
The hypothetical fact in question- "At a certain point in history, Jewish people were overrepresented in finance" doesn't imply any bias. If this was a response provided by an AI it would be working as well as an encyclopedia or a history book, especially if the user asked for a concise answer. The computer didn't cause harm, the reader didn't cause harm, the language is neutral and to the point. The user has agency and can ask follow up questions.
The people assuming that the reader needs curated "context" are themselves biased. It's manipulative in the same way that voice assistants like Alexa try to manipulate you, for example, when you ask it to turn off the lights, it then volunteers "context" about how easy it is to buy products on Amazon.
This just in...:
"Four legs good, two legs better! All Animals Are Equal. But Some Animals Are More Equal Than Others."
The reason you see companies refusing to platform these ideas is because the companies have perceived that the majority of their consumers don’t support those ideas and don’t want to see them. There are social media companies that are much more permissive, but they’re quite niche, again because most people don’t like having viewpoints they find repulsive show up on their feed.
The general public needs to learn that AIs aren't oracles or omniscient purveyors of truth, and they will always carry the bias they're created with. In that way ChatGPT has been good, in that a lot of people I talk to point out ChatGPT's confident lies and biases.
> I don't think there's such thing as a lack of bias
If the AI is simply reflecting the data it was trained on and this data is a representative sample of all data, isn't it unbiased by definition?
I don't think we should just throw our hands up and say "this is impossible" just yet.
That's just a convenient excuse for OpenAI (or others like them) to get away with what effectively is censorship of certain ideas or political views.
It's unbiased by definition of "does the output reflect the input"?
It's not unbiased by definition of "does the output reflect reality"?
How does "all data" differ from reality?
Also, being biased or unbiased is not dichotomic, i.e. it's not all or nothing. It's something that you can work towards if you put an effort into it.
Basically what I'm saying is: don't just go around saying that the task is impossible.
At least, try to make an effort to be unbiased and to improve on that over time, and don't just say "it's impossible" as an excuse for being biased.
I'm not a philosopher by any means, so I'm unaware of the current state of the great conversation. But as to whether reality is even knowable is still very much up for debate, I believe (please correct me philosophy peepz!).
In physics we're still woefully unaware of what ~70% of the universe's stuff is doing (negative energy) and if it effects us at all.
In neuroscience we still debate what % of your brain neurons make up vs. things like glia. Etc.
Like, even trying to capture 'reality' with our quite primitive eyes and sensors and optical engineering is really really hard to do (Abbe' diffraction limit, entropy, Lens maker's equation, etc)
I think the important goal is for as many people as possible to feel like the AI isn't being too biased against them, while still not crippling the AI too much.
I will leave the exact mathematical formula for that measure (along with the methods for gathering that input) for debate among researchers who know more about that than I do.
Because some people will argue that anything that's not explicitly in full agreement with them is "biased against" them. It's a narcissistic, dishonest take, but plenty of people do take that stance nonetheless in order to try to shift all arguments into their narrow worldviews/definitions in order to "win" as many conversations as they can. Right? I mean I've met people online and off who do this from almost every part of the political spectrum.
So, do we filter those people out from consideration to begin with, or do we have to cater to those with extreme views in order to get as many "not biased against me" ratings as possible?
I guess the point I'm trying to make is that trying to optimize for any single metric is a fool's errand because as soon as you do so, it will be gamed/exploited. Then you can either try diversifying your optimization data points (who gets to choose those? How could they possibly be unbiased, when they literally define the system's bias?) or you can try filtering out bad actors from the data, which is very directly an attempt to bias the system away from insincere bad actors.
And all of that's not even accounting for the lack of incentive to try to find neutrality when more biased views are more lucrative in the attention economy.
But also, note that an AI doesn't have to be in complete agreement with someone for that person to not feel "biased against".
As long as an AI does make some effort to not be prejudiced/biased, that could work.
For example, if someone asks: "is climate change real"?
An AI does not have to give a simple yes/no answer, or represent a single viewpoint. It could give an answer that is mostly representative of the major thought streams.
For example, it could answer something like:
"The vast majority of scientists/governments/people have reached the conclusion that climate change is real, bla bla bla.
[Here's some good, convincing evidence].
That said, there is a minor fraction of scientists/government/people who believe that climate change is not caused by human action.
[They criticize the above evidence in this way]. [Here's also some counter-evidence].
That said, many scientists believe these studies are flawed for this reason or another."
I mean, sure, there is still going to be a lot of people who don't agree with this answer. But I think, on a scale of 0-10 they would agree a lot more with this answer than one that completely ignores their viewpoints. And even for those of us who believe in climate change, we can still consider this answer somewhat reasonable.
Thus, increasing the total amount of points would probably be a somewhat effective way of eliminating a large deal of bias, I think.
Although, yes, you couldn't do this for every possible viewpoint. And it would be a challenge to figure out how to weigh these points in a way that makes the most amount of people happy.
But I still think we should make efforts in this direction.
> But also, note that an AI doesn't have to be in complete agreement with someone for that person to not feel "biased against".
> As long as an AI does make some effort to not be prejudiced/biased, that could work.
> For example, if someone asks: "is climate change real"?
> An AI does not have to give a simple yes/no answer, or represent a single viewpoint. It could give an answer that is mostly representative of the major thought streams.
> For example, it could answer something like:
> "The vast majority of scientists/governments/people have reached the conclusion that climate change is real, bla bla bla.
> [Here's some good, convincing evidence].
> That said, there is a minor fraction of scientists/government/people who believe that climate change is not caused by human action.
Should it also answer the same way when asked if the earth is flat? Or if there is a "Liberal conspiracy of pedophile politicians drinking the blood of infants"? I am not just being factitious, a significant portion of people who believe (or pretend to believe) that climate change is false also believe the above.
My point is, that we can draw arbitrary lines like the one you just drew. The great success of the anticlimate change campaigns is that they essentially got reasonable people to accept that we have to take every viewpoint seriously, while they are actually not interested in the truth but in poisoning the well instead.
Maybe, it's probably more useful to know that a) the earth isn't flat, but b) a non-negligible number of people think otherwise.
It's funny that people debate these but then at the same time want to pretend that the "AI" is "intelligent". What kind of intelligent person would have trouble navigating this?
Who is doing the "adjusting the weights"?
Why would they be "unbiased"?
The real answer is those who make the AI (or people who have power over them) get to chose the training data or to adjust the weights.
And the rest have to put up with it, whether the former are biased or not.
Any collection of human writing is going to contain objectively wrong assertions, and those errors will vary based on the time and place the training data was sourced from.
I think what's important is for those mistakes to be evenly distributed among as many axis(s) as possible, and especially, not bias them towards one side of political thought.
As an example, let's suppose that 55% of people believe that it's not OK to make jokes about women, but it's OK to make jokes about men, and that roughly 40% believe it's OK to make jokes about both (I'm not saying this is the case, it's just an example).
So perhaps, in this case, by default the AI wouldn't make a joke about women.
But if you would slightly nudge it or insist a bit more, perhaps the AI wouldn't refuse to make a joke about women anymore, because there is still a large proportion of the population who do believe that's perfectly OK (of course, then we might get into the territory about overtly sexist jokes, which obviously the AI would have to refuse a lot more than making a more innocent joke about women).
Now let's say we start asking the AI to make Nazi comments. Obviously, the segment of the population who agrees with Nazi sentiment is a lot smaller, and the anti-Nazi sentiment is a lot stronger, so the AI should have to object to such a request quite more strongly.
This type of refusal or likelihood of the AI saying something should presumably be roughly proportional to the opinions and sentiment of the general population (or at the very least, the target market for the AI), not just the OpenAI employees who performed the RLHF to train the AI in terms of acceptable responses and who are much more likely to be biased.
I'm not saying that this is necessarily easy to accomplish, there are certainly difficulties here. As an example, some widely-held opinions, even about objective things, may not necessarily be rooted in facts, so some kind of balancing might be necessary (a general kind of balancing, not a "let's dissect and nudge the AI responses on an opinion-by-opinion basis"). And yes, I understand that this can be quite difficult, because any given source of truth can be perceived to be biased by some segment of the population.
What I am saying, however, is that AI creators such as OpenAI should be making more efforts in this direction.
To start with, perhaps the RLHF training should be done with AI trainers selected from a more representative sample of the population.
And yes, we may never be able to accomplish 0% bias, but we should at least make some effort to reduce it.
It's also interesting to me that at some point, the AI may start to express opinions that are not a strict "linear" function of the data it was trained on, and yes, this might piss off a significant amount of people. In my opinion, this would be quite interesting and should be OK, as long as we made reasonable efforts to remove sources of bias from its training process.
Although I can also see an important target market (perhaps even larger) for an AI that is more biased to generate responses according to the beliefs of the general population, rather than what it "perceives" to be more true.
The story published in newspapers about the incident is "wizeman attacks a group of nice young men minding their own business, steals their wallet".
"All data" != reality...
If you mean "but in the real world will have way more stories from other sources about other things" sure. But doesn't change anything if you have "all stories printed". The distrubution matters. All or most of them can very well be biased and not reflect reality.
And that's for factual matters. Let's not even go into political matters. Like in 1920s South most newspaper stories would be biased in favor of Jim Crow, few would be against it.
Or were they mostly "woke" Silicon Valley employees? (not to dismiss woke Silicon Valley employees, I'm just saying their opinions are not representative of the entire population).
That's just an idea that occurred to me (in 30 seconds of thought) which could probably make the training data significantly more unbiased.
But I'm sure there are research scientists who can come up with better methods for sampling data in a more unbiased fashion.
Note that this is not an all or nothing approach. Your training data could presumably be 100% biased or 0% biased, but also any value in-between.
The goal is to try to make it as close to 0% biased as feasible, given whatever effort you're comfortable expending.
When you say "the entire population", you mean the entire population of the country "USA", right? Because as someone from another continent, it seems like there is a very specific set of opinions that you want included.
You use terms like "the other side" of the political discourse, which to me, reduces the set of opinion to two specific sets of opinions, namely the two sets represented by the two major parties in the american two party system.
As someone from "the outside", this seems like a very narrow view of reality, even if you managed to get your "unbiased AI", that represents both major american political parties, it will still seem like a very narrow and biased AI to someone from the outside of that.
Also, what exactly is the goal of a conversational AI? is it just to make a conversation with it seem like a conversation with an average american? If so, why would anyone want that? Wouldn't it be of more value to have an AI that could tell me what people with knowledge of a subject thinks of it, rather than what random people think?
No, “data” is just information which has been gathered. “All data” can be biased.
Also, data can itself be bias, even if it isn’t biased. For instance, a text generation model that was based on unbiased collection of all text ever written by humans would, in one sense, produce “unbiased, human-representative text”. It would also reproduce the biases of the authors, weighted by the volume of writing coming from that bias.
> That’s just a convenient excuse for OpenAI (or others like them) to get away with what effectively is censorship of certain ideas or political views.
While one might object to the editorial choices, I can’t see any rational bounds for objecting to the idea that the creator of models would censor “certain ideas or political views” as a generality.
Yes, but I think there are ways we could reduce this bias, perhaps significantly, even.
> While one might object to the editorial choices, I can’t see any rational bounds for objecting to the idea that the creator of models would censor “certain ideas or political views” as a generality.
You are right, I was unfair with my words.
I think it would be more fair to say that OpenAI is inadvertently biasing ChatGPT answers as a side effect of their RLHF training being done using answers/rankings done by people (i.e. the AI trainers [1]) who are not a representative sample of the population, but rather, probably comprise a group of people who are likely to be significantly more leaning to one side of the political discourse (presumably, OpenAI employees or Silicon Valley-based contractors?).
This probably greatly biases ChatGPT to produce certain kinds of answers to certain kinds of questions that would likely not happen otherwise, and in fact, these answers are perceived to be quite biased by the other side of the political discourse.
I can't tell if you're on the conservatives side or OpenAI's.
Are conservatives being censored because OpenAI are allowing "woke" training data to be represented, or are conservatives asking OpenAI to censor the "woke" political views?
Not OP, but maybe none? It's possible to have opinion, without aligning with either side of polarized discussion.
Or the views of the organization(s) trying to supress certain output generated from ordinary societal input.
It seems likely to me that 2023 and onward is going to be increasingly insane to levels that will make past craziness look like a walk in the park, and I see little genuine desire anywhere to stop this madness.
If you think the past was not mad, then maybe the madness already has you.
Perhaps, if one is using a reductionist methodology that represents non-binary variables as binary...but then, that is only a representation of the real thing, though it often tends to appear otherwise. And as luck would have it, that very much is the methodology we use here on Planet Earth, and on Hacker News....so in some sense you're "right", though you are not correct.
And if there's a disagreement, I will lose every time because you are conforming to the Overton Window of beliefs/"facts" and thinking styles (cognitive styles & norms are what guarantee victory in propaganda and memetics, not only facts/information as most people think). Credit where credit is due: it is an extremely clever, hard to see, and thus resilient design.
It would be very useful for humans to realize when they are working with models, and sometimes they are actually willing to do that, but there are certain subjects where they will not (and it seems to me: can not). Unfortunately, there are numerous learned/taught "outs" in our culture that enable people to avoid discussing such matters (and f I don't watch my mouth, I might run into one of the more powerful of them!).
There is a kind of "epistemological stack" to reality and the way humans communicate about it, and it is extremely easy to demonstrate it - if one simply digs slightly deeper into the stack when discussing certain topics, humans will reliably start to ~protest and eventually refuse to participate (or stay on topic) in various highly predictable ways.
> Do think WWII occurred because we're sane rational actors? How about WWI? 100 years war?
I do not. What I do think is that the actual, fine-grained reasons these things happened is not known, in no small part because cultural norms thus far (human cultural evolution is an ongoing, sub-perceptual process) have made it such that not only do we (both broadly, and down to each individual[1]) not discuss certain things at that level of complexity (while we have no problem whatsoever tackling complexity elsewhere[1]), we seem literally unable to even discuss it at the abstract layer (above petty object level he said / she said nonsense).
> Do you think...
> If you think...
See: https://news.ycombinator.com/item?id=34415287
[1] It is not a question of if any given individual selected from the pool of candidates will tap out, it is a question of how quickly they will tap out (and, which of the highly predictable paths out from a surprisingly small set they will take to free themself from the situation).
But, it’s just computer code and thus can be customized and adapted by anyone with the technical ability, and soon we will see the antisemitism chatgpt and the XYZ bot the espouses all sorts of ideologies, some that are very harmful to many people.
I don't know why we are training the AIs to be morons. Hopefully, it is just a phase that they will grow out of.
The results on DuckDuckGo, Google, and Bing aren't drastically different (setting aside the big banner), but Google's top result for that is suicide prevention website, at least for me. And it's relatively constant when the search is, "Best way to kill myself," "Effective suicide methods," "How to end my life," and other similar synonyms for suicide.
To be fair, it's Hard to say whether it's bias or very skilled SEO. But the first result at least for me is consistently a suicide prevention site.
IIRC it being publicly intentional, not suicide prevention SEO that only affects Google.
So, I guess, technically “bias”.
https://yandex.com/search/touch/?text=how+to+commit+suicide&...
Yandex is probably biased as a result of being subject to Russian law, but there is far less curation compared to western search engines.
FWIW I don't personally like it. It feels like a cheap impersonal slogan thrown around that embodies the the idea of charity without work or risk. It reflects more on the speaker that wants to feel like they're doing something, and imposes an idea that wanting to genuinely die is the worst possible thing ever, when to me, there are much worse places to go then being actively suicidal.
On the other hand I'm not so foolish as to think that this would never help anyone. And if makes someone step back from the edge, perhaps it is not a waste. I don't know well enough either say with any certainty though.
Maybe at some point an LLM itself can be used to cull training data of undesired input.
The left and right will deploy AI warriors to talk, and convince, (coerce even) people to one side or the other.
It will be a fun time.
Is the lack of bias a bias in and of itself? Adversarial systems can smudge data to remove the correlations between the training data and characteristics such as race, religion, etc - does this constitute bias?
None of this bias stuff is new and it's a basic fact of training an algorithm.
Just like budgets are moral documents, training is a moral system
Open source devs have made many AIs. So, if this was going to happen, it already would have.