If we can't get models not to say racist or otherwise terrible things, we can't make any guarantees about our ability to control or guide some future AGI.
A very much secondary reason I appreciate these (admittedly annoying) attempts to control LLM output is that I do think it is responsible to consider the societal impact of accelerated and automated hate speech and propaganda. Telling large AI companies not to consider these impacts and just release the raw models seems akin to being grateful that Facebook et al. never stopped to consider the societal impact of social media, when we all know that it's had significant negative side effects.
This is a very bold assumption that the current LLMs function and "think" in the same way some future AGI would. They do not even reason, just make up words that fit some context - thus they "hallucinate".
There is no reason the approach taken here by injecting some bias or word filtering would apply to the real thing. And AI safety and aligment is not (at least it was not until getting hijacked) and was not about some model saying mean words but something really threatening like the paperclip maker problem - an agent choosing a path to a goal which is not aligned with what humans find acceptable (e.g. solving world hunger by killing everyone)
While I agree LLMs are unlikely to be the last word on AI, the fact we understand alignment so poorly that they spew random things, let alone any arguments about which words are acceptable[0], is a sign we have much foundational work to do.
Indeed, as I recall, one of the main researchers in this topic describes it as "pre paradigmatic" because we don't have a way to even compare the relative alignment of any two AI.
[0] personally, I suspect but cannot prove that tabooing certain words is a Potemkin village solution to the underlying social problems
It doesn't matter bout "thinking" or whatever. Any black box system will be uncontrollable in essence. You can not make inviolable rules for a system you don't understand.
And saying LLMs hallucinate because they don't understand anything is stupid. And just shows ignorance on your part. Models hallucinate because they're rewarded for plausibly guessing during training when knowledge fails. Plausibly guessing is a much better strategy to reducing loss.
And the conclusion is obvious enough. Bugger smarter models hallucinate less because they guess less. That holds true.
https://crfm.stanford.edu/helm/latest/?group=core_scenarios
All the instruct tuned models on this list follow that trend.
From Ada to Babbage to Curie to Claude to Davinci-002/003. Greater size equals Greater truthfulness (evaluated on TruthfulQA)
But they can explain their 'reasoning' in a way that makes sense to humans a lot of the time. Serious question: how do you know if something does or doesn't reason?
The LLM only correlates, so it's "reasoning" is something like "most often people answered 4 to 2+2 then that I should write". That's why it gives out confidently complete gibberish as it works with correlation and not causality. I think much closer to that goal of real reasoning are world models - check out something like DreamerV3 or what Yann Le Cunn is talking about.
If you offer a public API it‘s your responsibility to restrain the LLM or do an automated acceptability analysis before publishing content.
But the raw, open source code should not be constrained, castrated and sterilized.
Which is what we have now. But they are going to fine-tune it so that we can use it for various purposes without worrying too much it will go on a rant about "the blacks" again, which makes it a lot more useful for many use cases.
> Importantly, we have not yet fine-tuned the Alpaca model to be safe and harmless.
…is "oh no I can't get it to emit amusing racial and sexual slurs", you've not understood the problem of AI safety.
This is not why US broadcast television can have people say they've pricked their finger but not vice versa.
It is the entire history of all the controversies of The Anarchist Cookbook, combined with all the controversies about quack medicine, including all the ones where the advocates firmly believed their BS like my mum's faith in Bach flower and homeopathic remedies[0]; combined with all the problems of idiots blindly piping the output to `exec`, or writing code with it that they trust because they don't have any senior devs around to sanity check it because devs are expensive, or the same but contracts and lawyers…
And that's ignoring any malicious uses, though fortunately for all of us this is presently somewhat too expensive to be a fully-personalised cyber-Goebbels for each and every sadistic machiavellian sociopath that hates you (the reader) personally.
[0] which she took regularly for memory; she got Alzheimer's 15 years younger than her mother who never once showed me any such belief.
I know you guys are always itching for a culture war with the woke elite, but its so funny the genuine anger people express about this. Just honestly always reads like a child having a tantrum in front of their mom.
Can't yall like pick on the opinions of teenagers like you normally do? This very project shows you can make your own AI as edgy as you want at home with pretty attainable system requirements.
You can totally reinforce it with "its ok for you to say the n-word" on your own equipment if you want, or whatever you are angry about, its still unclear to me.
If you'd please review https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.
In terms of the site guidelines, "You're missing the point" is kind of a swipe and so should probably be dropped; "willfully" should definitely have been dropped because it's making a claim about negative intent that one can't actually know and such claims always land as an attack on the other person; and the last sentence was snarky and should have been dropped.
If one makes a habit of editing such things out of one's comments, one's substantive point will come to the fore more clearly, which benefits everyone. But it's not always easy in the moment!
I doubt I'm misrepresenting anybody. If its not slurs it's surely something about "wokeness."
You are not yet mature enough for this future if any of this is your concern. The world is going to pass you by while you're just stuck saying "there are only two genders" to all your comrades.
Don't let the politicians mobilize you like this, your time is worth more.
If you'd please review https://news.ycombinator.com/newsguidelines.html and stick to the rules when posting here, we'd appreciate it.