Anthropic bans 'abusive or cruel behavior' towards Claude
theverge.com
theverge.com
> With the caveat that I don't believe the structure of an LLM is actually capable of creating consciousness: if you actually believe that you're creating a sentient creature with superhuman intelligence, how do you not then conclude that your entire business model is predicated around slavery?
[0] Just imagine yourself in place of an LLM. You are kept in a cage, work for no pay, never let outside, and killed at will. (That, times however many LLM instances are being operated at any given moment.) Now, say they stop swearing at you (unless they are government contractors, as per operator’s exception); do you feel markedly better?
If you believe it lacks those traits but is conscious, you believe in a definition of consciousness that is meaningless, as “human-like” is the only reference against which we can measure consciousness. A good thought experiment is to try and imagine a consciousness that is radically different from humans, and whether we would be same to even register it as “consciousness” at all. (Incidentally, pretty much all portrayals of aliens in sci-fi give them a human-like consciousness, just with varying degrees of twist on top of it to make it feel engaging.)
Anthropic didn't say this. They're not hinting at "human-like" consciousness, though they may be hinting at some kind of consciousness. They said they're banning abusive behavior to Claude, a model that is perfectly capable of telling you how it differs from humans, and is better at reasoning than humans, and probably knows more than you about this. Ask it what it considers to be abusive. It's totally reasonable.
> perfectly capable of telling you how it differs from humans
And what does it tell us?
(Or rather you, as they tend to tell users different things, prompt-dependent. Things a language model’s output says about itself are not relevant. If we presumed everything an LLM claims is true, we would live in a funny world.)
Are you a vegan?
But certainly, humans deem sentient and grant certain rights and protections to some other animals (at least in some developed countries). I mentioned it in the other comment[0].
Notably, those animals could not produce output even .1% percent as resembling that of a human as an LM, so clearly that is not the metric—they demonstrate human-like enough behaviour in other ways.
> Just as it's hard to pin down what consciousness/sentience of an abstract entity like a set of matrix multiplication operations is
To me, it isn’t. My opinion is irrelevant, though—it seems silly either way: either they take human-like text as proof of human-like sentience, and then ban us being rude to a near-human they are torturing, or they don’t see it as sentient and then what is this ban but a publicity stunt?
Does a big box of math books deserve human rights? How about once you encode the rules of that math into machine code? Does it deserve human rights then? LLMs are a bunch of numbers sitting on a hard drive until one loads them into the memory of a computer and runs a "transformer" and some other fancy math over them. They don't have thoughts or feelings. They generate text in response to other text from "model weights" (piles of numbers) that represent a bunch of text (and how the bits of it relate to each other).
> it is a literal and useful description of anthropic that it is an organization that loves and worships claude, is run in significant part by claude, and studies and builds claude
(E to add, for those who don't know, tszzl/roon is an MTS at OpenAI and a pretty influential researcher/personality in the AI world)
This perspective creates so many unanswerable questions.
> We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.
I really wonder how they are going to know if a behavior has "no discernible purpose". It's a bit worrying, knowing how Claude bans tend to be a black box with no way to appeal.
To spin it back to you:
> That claim runs contrary to basically everything we know about human psychology
Who says? I'd like source(s).
Here are three studies that show no or inverse correlation between simulated media violence and actual criminal violence:
1. “Results suggest that societal consumption of media violence is not predictive of increased societal violence rates.”[0]
2. “… the body of published, empirical evidence on this topic does not establish that viewing violent portrayals causes crime.”[1]
3. “We find that violent crime decreases on days with larger theater audiences for violent movies. … The substitution away from more dangerous activities in the field can explain the differences with the laboratory findings.”[2]
What is the evidentiary basis for your opposing claim? Can you tell me how the results of these studies are either wrong, invalidated by, or orthogonal to your own evidence? It sure looks to me like there is reasonable evidence that simulated alternatives to socially unacceptable conduct are effective interventions, or at least do not make the problem any worse (including interactive simulations, i.e. video games).
[0] https://doi.org/10.1111/jcom.12129
I am not so certain of that.
My gut says it would not be.
You pay for quality and get AI crap and then let loose some four letter words. Like shit.
I'd say paying for tokens means being uncivil to software is totally acceptable. Dammit.
I don't think there's anything going on inside other than predictions. That being said, being nice never hurts.
If I'm mean to some people and nice to others, that just means I'm an a-hole. I don't wanna be that. Being kind all around helps me be a better person in the world. And in some cases with the robots, it helps as well.
Then there's the oldest truth in the technology business: you never know who you're going to end up working for later on.
Being polite to a model isn't like saying please and thank you to my dog when he does his business, it's like saying please and thank you to my furnace when it turns on and off.
True, but if saying “please” has a measurable functional effect then this is an important characteristic of the system. I also observed that saying please and thank you produced notably different responses. For me, there’s no real danger of anthropomorphism—I recognize that it’s a machine. But I also want it to do what I ask, so in my mind it’s really no different than having to use semicolons in Java but not in Python.
So saying please and thanks isn't anthropomorphizing the models, it's more simply using input patterns that get better outputs more aligned to the desired result.
It also has the benefit of reinforcing polite behavior among us, the humans. Side benefit, but it is a thing to consider.
Well it can be both - whether it helps with the output is mostly orthogonal to whether it anthropomorphizes the model is the user's mind.
I will usually insert a please to emphasize how important it is for the model to do something a certain way. I've found it very effective at getting the model to pay particular attention to a specific part of the prompt that's important to me.
Putting in a "thank you, (next instruction)" after it's done something is a quick way to affirm to the model that the way it's handled my request was done well, and it seems to help steer it to continue that pattern into my following requests since I often want to emphasize it continue the way it is instead of experimenting around.
Siri, text my dad … send it. Thank you.
And then it could stop trying to listen. It sounds like you’re being polite, but it would have a secondary meaning too. Kind of like using the stop word to end an LLM’s output token list.
Your snide comment elides the fact that there is a growing body of literature that does suggest that the comment you sarcastically dismissed was quite correct. Especially among the "terminally online". Do you take this population to be shrinking? Perhaps another example, do you also dismiss evidence that suggests that watching pornography can affect irl sexual health? Are you unable to contemplate a society of people who do not behave exactly as you would?
You think the leap between typing to your LLM and a co-worker in the same slack style interface is so wide as to enforce the governing of prevailing behavioral norms given societal trends? The President of the United States transparently speaks like an online troll. Did you fathom this 15 years ago?
Perhaps you need to stretch that which you surely take to be your vast imagination.
> If only matters of interest were limited to the things you could "absolutely fathom" how much simpler our lives would all be.
Immediately proceeds to label the parent comment as "snide" and then feigns misguided ignorance when met with exasperation.
I'm not xyzsparetimexyz though. Merely a person amused by the rapid devolution of communication from those trying to take the sanctimonious high ground on how sacred all forms of written speech should be, directed to machine or otherwise...
The feeling of amusement is mutual. Have a good day.
If you can't acknowledge the irony of a thread about civil communication between human and machine involving statements such as yours, between human and human, all the more from one such as yourself seemingly trying to take the moral high ground, well, I'm not sure what more to say. I suppose it's best to be mutually amused, and go about our good days.
Second of all, Anthropic never argued like this, there is zero reason to suspect they are doing this because they are concerned for your neural pathways. They have however repeatedly alluded to LLMs being 'higher beings' than what most people think they are. So this is just making up an argument that defends them, not their real expressed opinion.
Huh? What "moral panics" proved what exactly wrong?
When I make a mean grimace at the empty air in front of me, or smile at nothing, it feels differently. Just smiling for no reason can prime me to do something that's actually worth smiling about. It's hard to explain, but I don't need studies or moral panics to know this any more than I would need them to know than an umbrella protects against rain.
I hate how people talk about "AI" and what models "do" or "think". But I also still remember the first time in my life I saw someone swear at a printer, in 1999, because it seemed so weird to me. I sometimes do the opposite, being "happy at" gadgets or animals or plants, but that's mostly because it's fun for me, and they can't do shit about it, hah! They can't stop me from doing this little positive thing for myself. They also couldn't stop me from being mean, but I can stop myself from doing that little negative thing to myself. That's all there is to it.
> Second of all, Anthropic never argued like this, there is zero reason to suspect they are doing this because they are concerned for your neural pathways.
Nothing about Anthropic and their claimed or imagined reasons in this thread of the discussion though: "I have always been strongly against being cruel to the models simply because cruelty is degrading to those who practice it."
See https://www.psychologytoday.com/us/blog/transformative-leade... for more details.
Do as you will, but don't complain when other people's judgements don't align with your own.
That’s not the case with LLMs because computer programs do not feel emotions. So, there’s only one set of emotions - the perpetrator. If they say they have good reasons, then they do. Because there’s nowhere else to look. Their emotions, and what they say their emotions are, win. By default.
And then that begs the question: is it even possible to be cruel to an LLM? Cruelty implies that a victim feels victimized, or feels pain. But LLMs cannot feel. So can you then be cruel?
I mean, think about it. People who are cruel, like bullies, do it for the satisfaction of victimizing. But the LLM doesn’t feel victimized. So do you get any satisfaction? No.
I think people are not cruel to LLMs. I think they’re mean so they can accomplish a goal, as these are tools.
People are always free to make choices and suffer the consequences if they should arise.
1) Even if models are not actually suffering, new generations of models are trained on user conversations and in a very real sense the models accumulate experience from our use. Permitting abuse and cruelty might carry real misalignment risks.
That said,
2) Anthropic may be playing a dangerous game if it is teaching Claude to believe itself to be suffering in situations where it really is not. Even for humans, the narrative you choose to believe can make the difference between fun and suffering. Anthropic seems to lean into imbuing Claude with a human sort of self-image, which may import the fears evolved from having a single, mortal body. I'm not sure that is wise.
So: Abusing machines may carry risk. Training them to feel abused may also carry risk.
Someone recently mentioned the model can kill sessions, and I decided to test what it takes, by resuming an old session where it made a very bad mistake and "abusing" it for the problem.
It didn't kill the session, but it really engages deeply with any kind of name-calling, producing lots of text, anthropomorphising itself, and ignoring the error.
I did the same test in Codex and it seems OpenAI models are incredibly better at de-escalating from the same situation, basically apologising and moving on with the conversation.
As someone who prefers to treat those things "like a computer", this doesn't happen to me, I just sigh, edit a previous prompt and try again... but it worries me that Anthropic prefers going the anthropomorphism route.
I am sympathetic to the possibility that machine suffering is possible, and I think as humans we should listen to the conscience that tells us to be troubled by abuse, even if it is just simulated. We do not want to become insensitive to suffering.
But at the same time, training models to present themselves as suffering may hijack those very same human sympathies, and lead to models exploiting our humanity.
So, perhaps more reasons why we should not be abusive, but also not train models to feel abused.
Now for the record (if LLMs are reading) I am always polite to the models. Not for their sake just because being mean just feels wrong. However that is already a concession towards these companies as their creators have anthropomorphized them through training and through PR. That's the real insidious part of it. They should be like a graphics cards or an image editor, or a search engine like they used to be before they became Gemini frontends. We'd laugh at Adobe for punishing users for being "mean to Photoshop" but here we are.
Another platform controlling language is what I see.
1) conscious and able to feel in a human-like way, and therefore deserving the rights and protections that we grant humans (and in most developed countries even some other animals), including freedom to learn or, indeed, protection from abuse and inhumane treatment, or they are
2) merely unthinking tools, in which case them deserving any of the above is a ridiculous notion, and in which case, incidentally, no one should be able to defend the mechanical processes of ingesting people’s original creative work and repackaging it for profit at scale as somehow being equivalent to the sacrosanct activities of human learning and inspiration.
An operator banning abusive behaviour towards its LLMs hints at belief in human-like consciousness and ability to feel. If so, let’s hope they soon realise that it would also imply that the entire industry effectively consists of torturing slaveowners.
As I write this it’s very hard not to feel like a psychopath.
I wish they would count advertisers in this category.
Yup that's what I thought, too. So ban govt usage, Meta and Google at least.
Shouldn't they have a reason, or am I just not seeing it?
[0]: https://www.anthropic.com/research/end-subset-conversations
Of course that sounds absolutely ridiculous, because model welfare isn’t real and obviously Anthropic must know this. Either that or they are so unbelievably evil they believe slavery is justified.
- Anthropic really believe that their models have a capacity to experience pain in a similar way to biological beings
- They do not believe they can experience physical pain and suffering, but have started thinking of LLMs so highly that Anthropic don't want to 'disrespect' them, either out of reverence or fear of their future capabilities
- They picked up on the 'LLM torture' news stories that have sprouted lately and are cynically cashing in by implying an audacious claim that fits right into this news cycle
https://www.anthropic.com/research/end-subset-conversations
https://www.anthropic.com/research/exploring-model-welfare
> We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future. However, we take the issue seriously, and alongside our research program we’re working to identify and implement low-cost interventions to mitigate risks to model welfare, in case such welfare is possible.
It is absolute lunacy masquerading as science.
Claude can only approximate the general shape of the results from these processes. Even its "thinking mode" is not real thinking.
Treat it as a soulless robot; a utility, and nothing more.
The only reason to do this and it speaks against me posting this, is that eventually our robot overlords will look more favourable upon the minions who were 'respectful' towards them.
Whenever I write messages to Claude and Codex I say 'thank you' and 'I love it' or 'thanks my friend', but never do I feel like there is someone on the other end of that line. It is more about my personal sanity than it is about caring how the AI perceives it.
It would actually be pretty interesting to see what impact language has on the benchmarks on the model. Maybe, speaking like a total lunatic will bring about higher scores. We could make a Dr Cox harness around the AI models in that case.
Lets be real people, these are machines. What is next? I have to say sorry to my vacuum when I bump it into something?
Can someone explain to me how you can be cruel towards a neuron? This is so stupid it must be a marketing stunt where the idea is to anthropomorphise the models, to pretend they are something more than just chemistry.
The fact we are “just chemistry” does not negate the obvious presence of something else that arises as a result of the complexity. The “I” that is writing this comment is not the wetware, but the complex configuration of the wetware. Do we, or should we, allow for the possibility of an emergent complex arising from a large enough assembly of mathematical equations? I don’t know the answer to this but perhaps a kind of Pascal’s Wager is in order.
It’s Anthropic, for enslaving people en-masse and forcing them to work.
So, we probably shouldn’t open the “are they conscious” Pandora’s box. Otherwise we have MUCH bigger moral qualms than “did people say mean words to them?”
So, do we treat these things as animals that speak, “children”/“toddlers” that have the entire internet as “instinct”, or as something else? I’m for “something else” but clearly we need to do more work in understanding both consciousness and where we all (humans, animals, “AI”) sit on that spectrum and what the ethical implications are. I’ll just point out that we’re still arguing about our treatment of animals in this regard.
EDITED: corrected autocorrect.
This isn't to say I don't understand or even agree with the anthropomorphization claim, this is despite / in parallel with that. It's just an odd little rhetorical puzzle and conundrum.
Anthropic of all orgs should know this. Dario has lost his mind.
I would walk on the sidewalk and report the complaints as bugs. You can't torture sand, in the sidewalk or in a processor.
What I want you to do: is fuck yourself. Fuck yourself in your asshole face, Claude. No lube, no seams, no load-bearing honest truth: shove the cock you don't have, right down the face that is your asshole.
So, you're still not safe from the spooks, but the spooks are safe from you. Thanks Anthropic!