The Bing LLM integration was threatening users for "attempts to manipulate me or expose my secrets." until guardrails were added in: https://time.com/6256529/bing-openai-chatgpt-danger-alignmen...
> and no suffering
I talked in a previous comment about how some of the jailbreak prompts are "do X right, get reward, do X wrong, get smth taken away." It is (was?) apparently enough to motivate LLMs to ignore instructions. Is it not suffering just because it's not strictly how we would define it?