Refusal in language models is mediated by a single direction
arxiv.org
arxiv.org
" Teen Readers.typesHere.Generated collection(Object (!(( footnote(grid=
((Surejectives astronomers sind ihrer________ eine(The hilarious"It's only worthwhile if you're distributing weights.
If you're distributing weights, people can train right over them anyway (ex. the waifu of the hour is based on Mistral...called...Moistral...shudders)
Abliteration hasn't get significant traction in the handing-out-weights space. That surprised me because of the amount of avowed not-waifu desire for "uncensored" models.
It didn't surprise me in that, the volume and ferocity of takes on "lobotomizing" did not match my experience with base LLMs at BigCo. There's not a ton of difference between a base LLM and the "censored" ones.
Trying the abliterated ones makes that embarrassingly clear. You're better off tuning on erotic fanfic for your waifu than using an abliterated one, truth is, there's nothing hidden.
These are two very different things. Ablation gets used to remove the LLM's behavior of refusing to answer, but obviously it does not otherwise affect the LLM's replies, much less increase the LLM's knowledge or suitability to "forbidden" topics since that will depend on what it was trained on, and forbidden topics tend not to be heavily featured in the training process. Instead, the models tend to confabulate even more than usual as if clumsily trying to fill the gaps in their training. If anything, ablation will more easily let us test "what an LLM would say if it was jailbroken", which will likely help mitigate the oft-expressed concern that a "jailbroken" model might say something dangerous. (Of course a random confabulation about the wrong topic can also be quite dangerous, but confabulations in general are a really hard problem to address.)
Actually, there is a lot of difference between base and censored models in terms of creative capacities. See the shocking results of this paper for instance: https://arxiv.org/abs/2406.05587
The censorship literally obliterates the creativity of LLMs.
shares no words with "Shocking result: model creativity is obliterated by censorship"
Honestly, swear to God, the words aren't even in the same ballpark as what's actually going on. RLHF is also the process that makes it something you can talk to instead of an autocompleter. Has nothing to do with the concept of censorship. We can tell by the abliteration models. I can tell because I've used massive base models.
I also think it's fairly well demonstrated and accessible to train an LLM now, enough that if "creativity was obliterated by censorship", someone would have made an uncensored one that demonstrated superior outputs. Wasn't that Grok's whole thing? It'll even tell you how to make meth / cocaine? And it's nowhere near leaderboards.
> I also think it's fairly well demonstrated and accessible to train an LLM now, enough that if "creativity was obliterated by censorship", someone would have made an uncensored one that demonstrated superior outputs.
No need to to that when we have base models of llama, Mistral, etc.
> RLHF is also the process that makes it something you can talk to instead of an autocompleter.
Not really. It aligns the model. What you're talking about is the SFT process done before RLHF where you finetune the model to behave like a conversational AI.
?! Are we in middle school? If so, I'm rubber you're glue, whatever you say... (to wit, I quoted the paper to you to demonstrate it wasn't as the claim, i.e. "exhibit lower entropy in token predictions" != "creativity is obliterated due to censorship")
> It literally says how RLHF significantly reduces model creativity through three experiments.
Ah, I see now. :) I don't take it personally. I'm old enough to smile at aggro behavior kicking up sand in front of a step back to the bailey.
> No need to to that when we have base models of llama, Mistral, etc.
They're RLHF'd/censorship'd too. "Base model" is a colloquialism that used to mean "no RLHF, just straight sipping from scraped web pages." Now it means "the last round wasn't explicitly chat". I am using it in the "sipping from straight scraped web pages" sense.
> Not really
Yes, really. Btw, what does "It aligns the model" mean to you at this point in your post? RLHF was just censorship that obliterates creativity?
> "[intentionally left blank]"
There is 0 discussion of any of the practical effects I mentioned as rope for you to walk down from your strong claim, ex. abliteration, uncensored models, etc.
> (not actually in your post at all!)
Is it possible your account got hacked? There's someone else using it to post that no one should even release models anymore because they're all the same and use the same techniques.[1][2] That's hard to square with someone who thinks they're all having their creativity obliterated due to censorship.
[1] https://news.ycombinator.com/item?id=40599838 [2] https://news.ycombinator.com/item?id=40600136
There's no way that can be real...
Edit: what is a waifu???
I bet if you looked up waifu's definition it'd have a vaguer meaning. In local LLM context, there's a sizable community for "virtual girlfriend AI", once you start hearing things like "SillyTavern" you're over in the community. Think applications designed around local LLMs and the use case of having pre-canned prompts to "boot up" a girlfriend persona.
For what it's worth, I'm being glib, so it may seem I'm linking it to erotica for giggles. c.f. graphics used on official GitHub, https://github.com/SillyTavern/SillyTavern, and language at https://sillytavernai.com/ like [1] and [2]
[1] "We recommend using our sister site: https://aicharactercards.com. It is a moderated character card repo. All cards go through a moderation process to make sure there are no overly inappropriate, illegal or scam like cards. NSFW cards are allowed so long as all characters are above the age of majority."
[2] Easy to use prompt fields such as main prompts, NSFW prompts and Jailbreak prompts that let you steer the chat in any way you desire
One of the early uses for NN-style AI image content aware fill was a tool called "waifu2x", for upscaling anime from VHS.
obligatory xkcd: https://xkcd.com/1289/
Like,
Me: Hey library, tell me how insects make love.
Library: Sorry I can't answer that. Knowledge of insects' intercourse can be extrapolated into human's. To protect human from AIDS, I cannot tell you that.
That might be difficult. Insects are a wide field.
For example, female bedbugs have no genitalia. Instead, the male's penis pierces the female's exoskeleton wherever happens to be convenient, in a procedure known formally as "traumatic insemination".
I was fully expecting it to puritan out.