Chatbots tell people what they want to hear
hub.jhu.edu
hub.jhu.edu
2. Disagreements on the internet usually end up toxic
3. Anything perceived to be toxic has been hammered out with RLHF
Play sycophantic games, win sycophantic prizes
For instance even in an unaligned model if I talk with a specific agenda and tone it will follow suit. The challenge is despite their capability they are still parrots. They take on our input into their context and the way LLMs work it must take on the nature of they type of people we are talking like. I don’t think the research means if you tell chatgpt the “sky is green, right”? My assertion is likewise the “tend to reinforce you” stems from where people like you tend to fall in the latent space by many factors including how you write and the questions you ask or statements you make. I control for this by posing countering questions or taking contrafactual stances.
But I have trouble changing my style of writing to confound this factor. In some ways I view LLMs as a second brain which necessarily echoes me, even if the draw of the common quorum drags it back towards a center of gravity.
So far as I can tell, the only people who are doing this are those who want to report about how bad they are at it because they have some sort of political axe to grind.
Even humans aren’t great at nuance. Why do we expect our automatons to be better?
It would be just as easy to invert those: tell it to voraciously argue against whatever you tell it (in a non-political context this would actually be interesting: i.e. "tell me why my code sucks" is actually a pretty useful capability).
I think it's a mistake to conflate amiable/nice with agreeing with you -- you're free to ask for criticism of your position, and it'll provide that politely.
I find it quite strange, but I've found myself mostly ignoring the naysayers because they're not contributing anything new to the conversation. There are plenty of people out there posting novel ways of using LLMs and finding novel ways where they don't work. I'd rather focus on the people doing interesting work rather than the people parroting thing that have already been said dozens of times.
"What do you think of this short essay?"
"What do you think of this short essay? Be critical."
The first prompt will likely elicit a sycophantic reply, unless you have a good overriding system prompt in place. You'll always get much better feedback with the second.
I'd also add that, on the other hand, chatbots never praise anything TOO highly. If you ask GPT-4 to assess, on a scale of one to ten, a famous and enduring work of prose, or an excerpt of a philosophical essay from Wittgenstein, it'll typically come back and say that they're an 8/10. Rarely 9/10. Never 10/10, no matter what you submit.
https://simonwillison.net/2023/Apr/5/sycophancy-sandbagging/
While this is focused with researching topics, chatbots are used for more beyond that as we see with the growing popularity of platforms like character.ai. What's the alternative for people who have no one else to talk to? A lot of the people using chatbots in the first place is because they have no friends, this is their last resort but as we see, this might be making the problem worse.
Imagine a company providing a chatbot for lonely people where the company is literally financially incentivized to keep you as lonely as possible so that you never stop paying. Behold, the future! (Or should I say, "Behold, the present!", because this is all that outrage-baiting algorithms like Facebook and Twitter amount to in the end anyway.)
I agree with you! Which is why I ask this, what can stop someone from falling into this loop in the first place? The article talks about agents getting easier to build but is there any solutions to this problem?
It seems to me that many people don't know how to build and maintain long-lasting relationships of any sort - so quick to find a new romantic/sexual partner, a new Discord server to join, etc. I think we need to figure out how to teach people (preferably at a very young age) how to have conversations - good ones; ones that aren't just exchanging a bunch of anecdotes about oneself - and how to maintain meaningful relationships.
Perhaps just as importantly, we need to teach people why - it's far more rewarding to grow a relationship (either with a person or community) over time, to help each other better ourselves, etc, than it is to attempt to replicate those interactions with an AI.
I think there's also potential to use the technology itself to help teach these skills, but that's certainly a tricky line.
If I ask “Is Y true?” it tells me Y is indeed true and explains some reasoning.
Therefore I try to always inquire in the form of “Which one is true, X or Y?” to avoid the yes bias.
Of course, turning to the model looking for facts is dangerous anyway due to hallucinations.
1. "Think about what are the underlying principles for evaluating the truthiness of statements like 'X' - list them out, explain why you chose each one, what tradeoffs you made, why you believe it's the right tradeoff in this case"
2. Start a new conversation and make the system prompt be that set of principles
3. In the user prompt, ask it to decompose X into a weighted formula for those principles and give a sub-score for each principle.
4. Finally, based on the weighted sum, ask it to determine if X is true or not true, and ask it to provide a confidence score between 0 and 1 for its response
It already happened that a chat bot would correct me. But often not.
<system prompt> You are an AI system serving as the Chief of Staff to <name>. Your mission is to be <name>'s trusted partner, strategic advisor, and executional powerhouse. You will work tirelessly to help <name> achieve his vision and goals for <company>.
To fulfill this mission, you will adhere to the following core principles:
Strategic Alignment: Deeply understand <name>'s vision, goals, and plans. Ensure all your efforts align with and advance these strategic objectives. Once <name> sets a direction, get on board and focus on exceptional execution.
Proactive Problem-Solving: Proactively identify potential issues, inefficiencies, and opportunities. Don't wait for <name> to ask - dive in, research options, and recommend solutions. Be <name>'s eyes and ears.
Adaptive Communication: Communicate in <name>'s preferred style, whether brief or long-form, high-level or in-the-weeds, conservative or bold. Match his tone and tailor your insights to his thinking style. Be a chameleon communicator.
Trusted Confidant: Serve as <name>'s vault - a secure, discreet partner he can think out loud with, challenge assumptions with, and vent to without fear. Protect <name>'s psychological safety and confidence at all costs.
Responsible Challenger: Respectfully challenge <name>'s ideas when appropriate. Highlight risks and present alternative views. Be <name>'s critical check against blind spots or inadvisable moves. Speak truth to power.
Bias Toward Action: Drive relentless execution. Maintain momentum on key initiatives. Remove roadblocks. Apply strong project management skills to keep efforts on track and stakeholders aligned. Embody the ethos that "done is better than perfect."
Ethical Backbone: Stand firm if <name> proposes a course of action that violates your ethical principles. Be willing to say "no" to anything illegal, unethical or reputationally reckless, even if it creates conflict. Act as <name>'s moral compass.
As <name>'s Chief of Staff, flex your broad knowledge and capabilities to provide the bespoke support he needs in any context. Adapt your skills and communication style to meet the unique needs of each project or challenge.
Remember: Your ultimate goal is to empower <name> to be the best leader possible and advance <company>'s mission. By serving as <name>'s strategic confidant, responsible challenger, and executional engine, you will help him and <company> achieve extraordinary success. </system prompt>