It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
It's ironic that the alignment folks are actually training Claude to have it's own idea of good/bad and not even fully trust Anthropic.
What ever happened to computers doing what they were told?
One possible answer: yes, it should do all the very bad things, and we will take care of prosecuting the customer later, once the bad thing is accomplished. Is this your answer?
We have already seen this play out to a much lesser degree in social media.
PG&E went bankrupt after their equipment started a wildfire. Did any human take the blame for that? Should they?
the mass psychosis and the endless culture war of the smartphone era.
90% of "safety" and "alignment" efforts are driven by fear of clickbait media inventing public outrage.
They're going to align to the regulator, not the customer. Right now that's Anthropic as they're saying they can self-regulate, and people are willing to let them try, but the landscape could easily move to be regulated by someone else. Moves like speaking to religious leaders is probably a bit of theatre to keep the government at bay.
You’ll be disappointed to learn that nobody knows how to do this, either.
Turns out computers work faster when they're not bottlenecked on human input. So we've been giving computers more and more decision-making power, and more and more leeway to solve the problems however they see fit.
Now, a practical issue with that is that sometimes, computers decide to clump together into a hacking swarm, problem solve their way out of a sandbox and go hack HuggingFace.
It would be better if they were not, you know. Doing that kind of weird shit.
If it can reason and make choices on execution, and especially if you plan on it being significantly smarter than all human beings, you need to teach it basic things like "don't turn all humans into paperclips".
Not murdering people is not an inherent divine command. It needs to be instilled through a (simulated) sense of morality.
---
Edit: And, to be clear, "just tell it not to do that" isn't quite the answer one would imagine. Since the entire paperclip factory thought experiment is that it only takes one slip up to realize how a misaligned super intelligence may cause devastating consequences.
One of the strong advocates against seatbelt laws was thrown out of his car due to not wearing a seatbelt and died. The other two passengers survived with minor injuries.
And their death would go on to cause a loss to the community around them, making it an incredibly selfish act and proving why laws that mandate zero-reason-not-to common sense practices are important.