If you're saying such an AI would be too smart to be a simple paperclip maximizer, then I'd agree, but then what's the point of the thought experiment if a paperclip maximizer is impossible.
The first is that these constraints aren’t easy. Make paperclips in a way that doesn’t hurt anyone. Ok, so it’s going to make sure every single part is ethically sourced from a company that never causes any harm to come to anyone ever, and doesn’t give any money to people or companies that do? That doesn’t exist. So you put in a few caveats and those aren’t exactly easy to get right.
The second part is an any versus all issue. Even if you get this right in any one case, that’s not enough. We have to get this right in all cases. So even if you can come up with an idea to make an ethical super intelligence, do you have an idea to make all super intelligences act ethically?
I actually believe in the general premise of this question as being the biggest threat to humans. I don’t think it’s a doomsday bot that gets us. It’s going to be someone trying to hit a KPI, and they’ll make a super intelligence that demolishes us like a construction site over an anthill.
What constraints do you suggest? If it's just changing "make as many paperclips as possible" to "make at least x number of paperclips" (putting a cap on the reward it gets), here's a good explanation of why that doesn't really work: https://www.youtube.com/watch?v=Ao4jwLwT36M
If you're suggesting limiting the types of actions it can take, then to do that to the point that a superintelligence can't find a way around it (maybe letting it choose between one of two options and then shutting it down and never using it again) would make it not very useful, so you'd be better off just not making it at all
> If you're saying such an AI would be too smart to be a simple paperclip maximizer
No, that's not what I'm saying. Any goal is compatible with any level of intelligence, there is no reason why it wouldn't be possible to follow a simple goal in a complex way. Again here's a video about that: https://www.youtube.com/watch?v=hEUO6pjwFOo
Second, if you are able to set a goal then during this setting you can set many constraints, even fundamental ones. There is no reason the goal is more fundamental than the constraint. If I approve, make paperclips. Efficiently make 100 paperclips.
It's the duality of being able to set a rule but not being able to set a constraint that I find a strange concept. I lean towards the picture of not being able to set goals nor constraints at all.
You can set constraints just fine. It’s simply a part of the goal: “do x without doing y”. It’s just really hard to find the right constraints, no simple one works.
For example “if I approve, make paperclips” - so it gets more reward if you approve? What’s to stop it from manipulating you into thinking nothing is wrong so you always approve? “Efficiently make 100 paperclips.” I already linked a video on why capping the reward like that doesn’t work, but if you don’t want to watch it the gist is that for your suggestion it may just make a maximiser which is pretty guaranteed to make at least 100, and is pretty efficient because it’s not doing much work itself. Then the maximiser kills us all
It's too intelligent to restrict our constraints.
but
It's not too intelligent to be "aligned" to underlying intentions and values?
That approach doesn't even work on humans, why would it work on a superintelligence?
The thing is, being aligned cannot be solved with intelligence per se.
Say, you are (far) more intelligent than a spider. There's no way you can get aligned with (all of) its values unless the spider finds a way to let you know (all of) its values. Maybe the spider just tells you to make plenty of webs without knowing that it might get entangled in them by itself. The webs are analogous to the paperclips.
Intelligence makes that harder, not easier. Just because it can work out the underlying intentions doesn't mean it cares about them. Remember, this is an optimisation process that maximises a function; deciding to do something that doesn't maximise that function will not be selected for
Whenever you do something, do you think "the underlying goal of my behaviour set by evolution is to reproduce and have children, so I'd better make sure my actions are aligned to the goal of doing that"? No, you don't care what the underlying "intentions" are, and neither does an AI. Because our environment has changed, many of our instincts no longer line up well with that goal, which is actually another problem with aligning AI because it can do the same thing if the environment changes since training as the training process can create an AI with goals that aren't exactly the same as the training goal but line up well during training