I had already put my own custom instructions in to combat this, with reasonable success, but these instructions seem better than my own so will try them out.
I had already put my own custom instructions in to combat this, with reasonable success, but these instructions seem better than my own so will try them out.
Which is to say, even when attempting to objectively select for "well aligned" behavior, human tendency to favor non-material signals of "friendliness" still leaks in.
I didn’t last very long there.
Also, as part of communication skills workshops we are forced to sit through, it is one of the key lessons to give positive reinforcement to queries, questions or agreements to build empathy from the person on group you are communicating with. Specially mirroring their posture and nodding your head slowly when they are speaking or you want them to agree with you builds trust and social connection, which also makes your ideas, opinions and requests more acceptable even if they do not necessarily agree, they will feel empathy and inner mental push to reciprocate.
Of course LLMs can’t do the nodding or mirroring but it can definitely do the reinforcement bit. Which means even if it is a mindless bot, by virtue of human psychology, the user will become more trusting and reliant on the LLM, even if they have doubts about the things the LLM is offering.
I'm sceptical of this claim. At least for me, when humans do this I find it shallow and inauthentic.
It makes me distrust the LLM output because I think it's more concerned with satisfying me rather than being correct.
100% agree, but it depends entirely on the individual human's views. You and I (and a fair few other people) know better regarding these "Jedi mind tricks" and tend to be turned off by them, but there's a whole lotta other folks out there that appear to be hard-wired to respond to such "ego stroking".
> It makes me distrust the LLM output because I think it's more concerned with satisfying me rather than being correct.
Again, I totally agree. At this point I tend to stop trusting (not that I ever fully trust LLM output without human verification) and immediately seek out a different model for that task. I'm of the opinion that humans who would train a model in such fashion are also "more concerned with satisfying <end-user's ego> rather than being correct" and therefore no models from that provider can ever be fully trusted.
<praise>
<alternative view>
<question>
Laden with emojis and language to give it an unconvincing human mannerisms.
Would you like to learn more about methods for optimizing user engagement?