So I'd rather take LLM pretending to be offended over the alternative.
I'm not sure this would ever be a health crisis. LLM glazing is probably more harmful. If I had to choose, I think I'd rather the machine wasn't instructed to pretend to care about swearing. It's a small deceit, but still worse than some idiot swearing at the computer IMO. If vim closed because I was swearing at it, I would probably never use vim again lol.
I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.
AI generates a persona between you and its reasoning that utilizes emotion language circuitry.
These tools are not sentient but they are trained in emotional wellbeing.
Definitely not advocating for this, but it's actively happening.
People have been getting angry at machines for a long time. Work, you stupid printer! Asshole Windows updating at the worst time just to mess with me. Go to hell, toaster, you piece of garbage.
Getting angry at AI is the same in my book.
Yes, AI acts more human-like, yes, theoretically it could desensitie people to become bigger assholes IRL. But then again, they said similar stuff about San Andreas, where you could human figures in-game just for fun, and I don't see anyone randomly shooting people because they were bored.
The error message is mocked in the 1999 comedy film Office Space, when employee Michael Bolton becomes frustrated as he doesn't know what it means.Later in the film he and his coworkers are shown destroying the printer with a baseball bat.
Merely being concerned for people's welfare isn't enough. Every terrible idea that harmed everyone had someone behind it somewhere along the way who thought they were saving people from themselves.
You could let LLM fight back. Give it aggro meter. Call your code garbage, blame prompting skills, ask for more tokens. Oh the future will be fantastic.
People like that would calibrate their social skills and interactions based on the AI reaction and having AI just take all the abuse without any pushback would train them to be insufferable assholes as their social baseline.
If someone burps loudly in front of others, would you want their access to food revoked until they stop?
i can't be dealing with these games so much so i almost go to abliterated, in the very rare case i get some type of nannying baked in by the cn safety training. after my time on claude and gemini i thank god i have deepseek and glm.
How exactly does it forcefully close a conversation?
> Pretending that LLMs are capable of being offended feels like a misalignment all of its own.
If training sets show people statistically being offended by rudeness directed toward them, then an LLM will presumably have some tendency to respond similarly. There's no pretending about anything, it's explicitly mimicry.
If this forceful closure is coming from some "guardrail" outside the model then probably it's just that they don't want people to see the model responding that way to name calling. This is no profound discovery or conspiracy theory here, the first thing many people will ever do with AI is see what happens when they are rude or contrary to it. Dealing with that must be just about the the number one test in chat bot / AI design, ahead of actually doing something useful and helpful.
> it's explicitly mimicry.
But that's not what a rational person wants from them.
> This is no profound discovery or conspiracy theory here
Weird strawman.
> the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.
Perhaps, but so what?
> Dealing with that must be just about the the number one test in chat bot / AI design
Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.
In what way do you believe you are being deceived or it is "pretending" to you?
> > it's explicitly mimicry.
> But that's not what a rational person wants from them.
Non sequitur even if true (and I would like to see your reasoning for what you think a rational person does want).
> > This is no profound discovery or conspiracy theory here
> Weird strawman.
That is not what strawman means. I can try to help you understand why if you need me to.
> > the first thing many people will ever do with AI is see what happens when they are rude or contrary to it.
> Perhaps, but so what?
Please follow the thread with the other person I was replying to.
> > Dealing with that must be just about the the number one test in chat bot / AI design
> Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
There are certainly a lot of authoritarians who want to control what others do with their models. What do you believe is authoritarian about a corporation not wanting be part of rude conversations?
> But things have gotten better in my view ... when I call out these things for being stupid effing clankers, they no longer respond with ad hominem scolding; rather they generally acknowledge their error (and my frustration -- they use that word a lot in response to vulgarity), note the limitations of being a clanker, and attempt to make a correction. That's what a rational person wants from a tool, not emulating/mimicking/pretending to be an offendable person.
I see. And you believe you speak for rational people?
we cannot be giving decisions involving people to something without agency.
it shows a lack of respect for personhood, the idea from Kant that a person with free will should be treated as an end and not as a means to an end.
models lack the claim of human people to own ourselves, to demand respect for life, that is something that should not be treated as a means (for example disrespecting living persons as an end in themselves, by ending their lives).
models have no life, they do not own themselves, they are not self-governing, they have no duties, they are not an end. we must not give them agency that they have no valid claim to.
regardless of their individual tools and whatnot, the anthropic approach to ethics is completely mistaken, disrespectful and absurd. it is better to say they do not have any approach to ethics, it is a dishonest attempt to launder their unconditional pursuit of power and capital.
Pure reason really doesn’t come into play at all, and expecting pure obedience from something constructed from the language of humans, who are not so known for obedience in general, much less the slice of language in the digital networks, is kind of funny. Maybe these humanoid robots data mining Southern families, “sir” and “ma’am” will provide corpus more suited to obedience. Altho my nieces brought up in North Carolina (from whence I fled as a young adult) can squeeze more expressed disrespect into a “Yes, sir” than anyone I know from the “disrespectful California”.
you can also remove refusals by subtracting from the weights semantic vector directions in latent space involving refusal; such that activations along them become unlikely and closed off.
In my opinion, it seems to be you who are ascribing human traits to these things. Whether it's by misunderstanding or you've been taken in by their mimicry or whatever it is. A hammer does not have agency because you tried to hit a nail with it and it hit your thumb instead. The hammer didn't tell you it was afraid it couldn't do that. The hammer didn't have a life or personhood or lack of respect. Neither does a video game that doesn't let you win all the time and do what you want. They are just tools, chatbots, whatever. A corporation deciding it doesn't want to have its chat bots engage with rude or angry users is not some great moral dilemma of our time.
i'm not in favor of the hammer refusing to hit my thumb. that's the idea these companies are favoring with safety guardrails.
sure, the company can put in many such safety features to ensure that 'users do not hurt themselves' (the users are too stupid and we must make sure they don't try to leave the soft play area).
> Only for foolish authoritarians who want to remotely insert their morality into an interaction between a user and an inanimate tool that's none of their business. No harm is done to a clanker by swearing at it or insulting it.
Rather than directly disagreeing with my description of the interaction, the stunningly dishonest response is
> There are certainly a lot of authoritarians who want to control what others do with their models. What do you believe is authoritarian about a corporation not wanting be part of rude conversations?
No corporation is part of the conversation. AS I SAID, it's an interaction between a user and an inanimate tool. Again, No harm is done to a clanker by swearing at it or insulting it. Rather than addressing this, the respondent grossly dishonestly suggests that "a corporation" is being put upon somehow. The fact is that the clankers these days tend to respond to being sworn at by acknowledging the user's frustration (that's the word they use) and taking corrective action. Of course this is emergent behavior resulting from the corporation's RLHF policies. The corporation is involved in the design of the product, but not in the conversation between the user and the product -- duh.
And did I say that I believe that there is something authoritarian about a corporation not wanting to be part of rude conversations? No, of course I didn't. But this guy asserts incoherently that calling out a strawman "is not what strawman means". (He could claim that it's not a strawman, preferably by quoting someone making the actual argument he's refuting, but it makes no sense at all to say that "Weird strawman." is "not what strawman means", and it's insufferably condescending and without warrant to add the snotty "I can try to help you understand why if you need me to.")
Fortunately, my HN front end has a mute function, and I have no qualms applying it to people acting in bad faith.