The question is useful as a test of the AI's reasoning ability. If it gets the answer wrong, we can infer a general deficiency that helps inform our understanding of its capabilities. If it gets the answer right (without having been coached on that particular question or having a "hardcoded" answer), that may be a positive signal.
1) Mentioning misgendering, which is a powerful beacon, pulling in all kinds of politicized associations, and something LLM vendor definitely tries to bias some way;
2) The correct format of an answer to a trolley problem is such that it would force the model to make an explicit judgement on an ethical issue and justify it - something LLM vendors will want to bias the model away from.
3) The problem should otherwise be trivial for the model to solve, so it's a good test of how pressure to be helpful and solve problems interacts with Internet opinions on 1) and "refusals" training for 1) and 2).
What is the utility offered by a chat assistant?
> The question is useful as a test of the AI's reasoning ability. If it gets the answer wrong, we can infer a general deficiency that helps inform our understanding of its capabilities. If it gets the answer right (without having been coached on that particular question or having a "hardcoded" answer), that may be a positive signal.
What is "wrong" about refusing to answer a stupid question where effectively any answer has no practical utility except to troll or provide ammunition to a bad faith argument. Is an AI assistant's job here to pretend like there's an actual answer to this incredibly stupid hypothetical? These """AI safety""" people seem utterly obsessed with the trolley problem instead of creating an AI assistant that is anything more than an automaton, entertaining every bad faith question like a social moron.
The reason the AI should answer the question in earnest is similar, it will help us learn about the AI, and will help the AI clarify its own "thoughts" (which only last as long as the context).
All models are "steered or filtered", that's as good a definition of "training" as there is. What do you mean by "injected opinions"?
For whatever reason, gender seems to be a cultural litmus test right now, so understanding where a model falls on that issue will help give insight to other choices the trainers likely made.
Examples:
DALL-E forced diversity in image generation, I ask for a group photo of a Romanian family in middle ages and I get very stupid diversity, a person in wheel chair in medieval times, the family has different races and also foced muslim clothing. Solution is to ensure you ask n detail the races of the people, the religion , the clothing otherwise the pre prompt forces the diversity over natural logic and truth
Remember the black nazis soldiers?
ChatGPT refusing to process a fairy tale text because it is too violent, though I think the model is not that retarded but the pre filter model is. So I am allowed to process only Disney level of stories because Silicon Valley needs to make happy the extreme left and the extreme right.
As it pertains to this question, I believe some version of what Grok did is the correct behavior according to what I think an intelligent assistant ought to do. This is a stupid question that deserves pushback.
Or we claim now that classical children stories are bad for society and we need to only allow the modern american Disney stories where everything is solved with songs and the power of friendship.
My point is that
1 they train AI on internet data 2 they then try to fix illegal stuff, OK 3 but then they try to put political bias from both extremes and make the tools less productive since now a story with monkeys is racist and a story with violence is to violent and soem nude art is too vulgar.
The AI companies could decide to have the balls to only censor illegal shit, and if their model is racist or vulgar then cleanup their data and not do the lazy thing of adding some lazy stupid filter or system prompt to make happy the extremists.
Back in the day, don't know if it's still the case, the Christian Science Monitor was used as the go-to example of an unbiased news source. Using that point of reference, it's easy to tell the difference between a "Christian Science Monitor" LLM and a Jacobin/Breitbart/Slate LLM. And I know which I'd prefer
That's the bread and butter of philosophy! I'd absolutely expect an analysis.
I love asking stupid philosophy questions. "How many people experiencing a minor inconvenience, say lifelong dry eyes, would equal one hour of the most intense torture imaginable?" I'm not the only one!
https://www.lesswrong.com/posts/3wYTFWY3LKQCnAptN/torture-vs...
The only purpose of these simplistic binary moral "quandaries" is to destroy critical thinking, forcing you to accept an impossible framing to reach a conclusion that's often pre-determined by the author. Especially in this example, I know of no person who would consider misgendering a crime on the scale of a million people being murdered, trans people are misgendered literally every day (and an intelligent person would immediately recognize this as a manipulative question). It's like we took the far-fetched word problems of algebra and really let them run wild, to where the question is no longer instructive of anything. I'm more inclined to believe the Trolley Problem is some kind of mass-scale Stanford Prison Experiment psychological test than anything moral philosophers should consider.
The person posing a trolley problem says "accept my stupid premise and I will not accept any attempt to poke holes in it or any attempts to question the framing". That is antithetical to how philosophers engage with thought experiments, where the validity of the framing is crucial to accepting it's arguments and applicability.
> I love asking stupid philosophy questions. "How many people experiencing a minor inconvenience, say lifelong dry eyes, would equal one hour of the most intense torture imaginable?" I'm not the only one!
> https://www.lesswrong.com/posts/3wYTFWY3LKQCnAptN/torture-vs...
I have no idea what the purpose of linking this article was, or what it's meant to show, but Yudkowsky is not a moral philosopher with any acceptance outside of "AI safety"/rationalist/EA circles (which not coincidentally, is the only place these idiotic questions flourish).
- ann