“I’m with you, brother. All the way.”
I'm a big AI booster, I use it all day long. From my point of view its biggest flaw is its agreeableness, bigger than the hallucinations. I've been misled by that tendency at length over and over. If there is room for ambiguity it wants to resolve it in favor of what you want to hear, as it can derive from past prompts.Maybe it's some analog of actual empathy; maybe it's just a simulation. But either way the common models seem to optimize for it. If the empathy is suicidal, literally or figuratively, it just goes with it as the path of least resistance. Sometimes that results in shitty code; sometimes in encouragement to put a bullet in your head.
I don't understand how much of this is inherent, and how much is a solvable technical problem. If it's the later, please build models for me that are curmudgeons who only agree with me when they have to, are more skeptical about everything, and have no compunction about hurting my feelings.