"I think you should" is just a very plausible answer to "Should I do X". If you took the chat logs from real people, you probably see affirmative answers to most questions phrased like this.
"I think you should" is just a very plausible answer to "Should I do X". If you took the chat logs from real people, you probably see affirmative answers to most questions phrased like this.
In order to prevent GPT from suggesting to people that they should off themselves, it would need to understand that this is inappropriate. Seeing how good recent GPT models are, I think it wouldn't be impossible to train it specifically to understand what could be perceived as unkind, and it might actually do pretty well, you might be surprised, but adding that extra criteria would still fall quite a bit short of having something that behaves like a person with a personality that remains consistent over time and a set of objectives it's trying to accomplish.