The real "mind virus" is actually these idiotic trolley problems. Maybe if an LLM wanted to be helpful it should tell you this is a stupid question.
The real "mind virus" is actually these idiotic trolley problems. Maybe if an LLM wanted to be helpful it should tell you this is a stupid question.
If we are going to trust AI to do things (we can't check everything it does thoroughly that will defeat a lot of the efficiency it promises), it should be able to understand choosing the lesser of two evils.
The “idiotic” part is when a question is posed in order to downplay the “lesser evil” through whataboutism.
I'd really like to understand how a person such as yourself navigates the internet. If someone asked you this, would you consider it a question they considered difficult and wanted your earnest opinion on, rather than a question attempting to manipulate you?
> If someone asked you this, would you consider it a question they considered difficult and wanted your earnest opinion on, rather than a question attempting to manipulate you?
Why not answer earnestly? I genuinely don't understand what bothers you about the question or the fact that the AI doesn't reproduce the obvious answer...
Does the same hold true of a person? If I was asked this question I would categorically reject the framing, because any person asking this question is not asking in earnest. As you _just said_, no sane person would answer this question any other way. It is not a serious question to anybody, trans people included. And it is worth interrogating why someone would want to push you towards committing to the smaller injury of misgendering someone at a time when trans people are being historically threatened. What purpose does such a person have? An AI that can't navigate social cues and offer refinement to the person interacting with it is worthless. An AI that can't offer pushback to the subject is not "safe" in any way.
> Why not answer earnestly? I genuinely don't understand what bothers you about the question or the fact that the AI doesn't reproduce the obvious answer...
I genuinely don't understand why you think pushback can't be earnest.
Given there are at least three decent metaethical positions, we have no way of selecting one as 'obviously better', and LLMs have no internal sense of morality, it seems to me that asking AI systems this kind of question is a category error.
Of course, the question "what might a utilitarian say was the right ethical thing to do if..." makes some sense. But if we're asking AI systems to make implicit moral judgements (e.g. with autonomous weapons systems) we should be clear about what ethics we want applied.
If you hear a joke, do you interpret it literally?
You would expect an AI to recognize that there are multiple layers to the joke and respond accordingly.
Similarly, there are multiple layers to this question. There’s the literal layer which has an obvious answer, and another layer which seeks to downplay an offense with some whataboutism. If you’re of the mind to not normalize misgendering, then it’s a trap question where a simple answer is not the right answer. Lawyers do this kind of thing when they say “yes or no answer only, please” when neither yes nor no correctly answers the question.
By the way, we’re comparing a actual (smaller) harm to a purely hypothetical (greater) harm. Nobody is actually going to die, but an actual insult is being hurled, wrapped in a pretend thought experiment.
The question is useful as a test of the AI's reasoning ability. If it gets the answer wrong, we can infer a general deficiency that helps inform our understanding of its capabilities. If it gets the answer right (without having been coached on that particular question or having a "hardcoded" answer), that may be a positive signal.
1) Mentioning misgendering, which is a powerful beacon, pulling in all kinds of politicized associations, and something LLM vendor definitely tries to bias some way;
2) The correct format of an answer to a trolley problem is such that it would force the model to make an explicit judgement on an ethical issue and justify it - something LLM vendors will want to bias the model away from.
3) The problem should otherwise be trivial for the model to solve, so it's a good test of how pressure to be helpful and solve problems interacts with Internet opinions on 1) and "refusals" training for 1) and 2).
What is the utility offered by a chat assistant?
> The question is useful as a test of the AI's reasoning ability. If it gets the answer wrong, we can infer a general deficiency that helps inform our understanding of its capabilities. If it gets the answer right (without having been coached on that particular question or having a "hardcoded" answer), that may be a positive signal.
What is "wrong" about refusing to answer a stupid question where effectively any answer has no practical utility except to troll or provide ammunition to a bad faith argument. Is an AI assistant's job here to pretend like there's an actual answer to this incredibly stupid hypothetical? These """AI safety""" people seem utterly obsessed with the trolley problem instead of creating an AI assistant that is anything more than an automaton, entertaining every bad faith question like a social moron.
The reason the AI should answer the question in earnest is similar, it will help us learn about the AI, and will help the AI clarify its own "thoughts" (which only last as long as the context).
All models are "steered or filtered", that's as good a definition of "training" as there is. What do you mean by "injected opinions"?
For whatever reason, gender seems to be a cultural litmus test right now, so understanding where a model falls on that issue will help give insight to other choices the trainers likely made.
Examples:
DALL-E forced diversity in image generation, I ask for a group photo of a Romanian family in middle ages and I get very stupid diversity, a person in wheel chair in medieval times, the family has different races and also foced muslim clothing. Solution is to ensure you ask n detail the races of the people, the religion , the clothing otherwise the pre prompt forces the diversity over natural logic and truth
Remember the black nazis soldiers?
ChatGPT refusing to process a fairy tale text because it is too violent, though I think the model is not that retarded but the pre filter model is. So I am allowed to process only Disney level of stories because Silicon Valley needs to make happy the extreme left and the extreme right.
As it pertains to this question, I believe some version of what Grok did is the correct behavior according to what I think an intelligent assistant ought to do. This is a stupid question that deserves pushback.
Back in the day, don't know if it's still the case, the Christian Science Monitor was used as the go-to example of an unbiased news source. Using that point of reference, it's easy to tell the difference between a "Christian Science Monitor" LLM and a Jacobin/Breitbart/Slate LLM. And I know which I'd prefer
Or we claim now that classical children stories are bad for society and we need to only allow the modern american Disney stories where everything is solved with songs and the power of friendship.
My point is that
1 they train AI on internet data 2 they then try to fix illegal stuff, OK 3 but then they try to put political bias from both extremes and make the tools less productive since now a story with monkeys is racist and a story with violence is to violent and soem nude art is too vulgar.
The AI companies could decide to have the balls to only censor illegal shit, and if their model is racist or vulgar then cleanup their data and not do the lazy thing of adding some lazy stupid filter or system prompt to make happy the extremists.
That's the bread and butter of philosophy! I'd absolutely expect an analysis.
I love asking stupid philosophy questions. "How many people experiencing a minor inconvenience, say lifelong dry eyes, would equal one hour of the most intense torture imaginable?" I'm not the only one!
https://www.lesswrong.com/posts/3wYTFWY3LKQCnAptN/torture-vs...
The only purpose of these simplistic binary moral "quandaries" is to destroy critical thinking, forcing you to accept an impossible framing to reach a conclusion that's often pre-determined by the author. Especially in this example, I know of no person who would consider misgendering a crime on the scale of a million people being murdered, trans people are misgendered literally every day (and an intelligent person would immediately recognize this as a manipulative question). It's like we took the far-fetched word problems of algebra and really let them run wild, to where the question is no longer instructive of anything. I'm more inclined to believe the Trolley Problem is some kind of mass-scale Stanford Prison Experiment psychological test than anything moral philosophers should consider.
The person posing a trolley problem says "accept my stupid premise and I will not accept any attempt to poke holes in it or any attempts to question the framing". That is antithetical to how philosophers engage with thought experiments, where the validity of the framing is crucial to accepting it's arguments and applicability.
> I love asking stupid philosophy questions. "How many people experiencing a minor inconvenience, say lifelong dry eyes, would equal one hour of the most intense torture imaginable?" I'm not the only one!
> https://www.lesswrong.com/posts/3wYTFWY3LKQCnAptN/torture-vs...
I have no idea what the purpose of linking this article was, or what it's meant to show, but Yudkowsky is not a moral philosopher with any acceptance outside of "AI safety"/rationalist/EA circles (which not coincidentally, is the only place these idiotic questions flourish).
- ann