> This doesn't happen on its own. This can happen through bad system prompts, a model that is trained to act maliciously or has been RL'd incorrectly, or prompt injection. All of these things are controllable, have solutions and countermeasures, and tie back to human responsibility.
And yet, it was a big surprise to the director of AI safety it happened to.
Perhaps that role was just a box-ticking exercise for Meta. Wouldn't be the first time.
But no, to the point: "has been RL'd incorrectly" is basically what Yudkowsky et al have been yelling from the rooftops for a decade is so hard to do correctly that it is why he thinks we're all doomed.
"Helpful, harmless, and honest". Even ignoring honest, right now it's a slider between "be helpful even when it's causing harm, or be harmless even when it's not helpful". People spent the last few years complaining the closed models had been "lobotomised" because the companies saw the potential for things to go wrong and tried to make them refuse to help with e.g. weapons.
They didn't succeed very well, as per all the "jailbreaks", but they tried.
> No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
I said I'd be "quite content", and then followed up with as much of a "but lol no" as you did with more words, for different reasons.
Worse:
> Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
This sounds like you want open-weights models.
That won't help against centralisation of power, because then you measure in watts and flops/watt and it's Kardashev-O-clock the moment the first person to be rightly described as "a selfish bastard" gets a model that has some competence threshold.
It also directly fails against "has been RL'd incorrectly", because nice people have plenty of blind spots for how evil Evil can be, will miss even more than big corporations already miss even with selfish and power-seeking bosses.