Artificial intelligence will do what we ask, and that's a problem
quantamagazine.org
quantamagazine.org
The most obvious analogy, and I acknowledge the potentially controversial nature of it, is that of a date rapist attempting to defend their actions by pointing out the victim was dressed provocatively, gave all the "signals" that they wanted sex, and when they expressly said "no" the rapist knew from their prior behavior that they actually did want it (or had some inner heretofore unexpressed need for it).
The true underlying motivation for such a rewrite of Asimov's laws lies in the paragraph that follows Russell's new list. "...developing innovative ways to clue AI systems in to our preferences, without ever having to specify those preferences." Perhaps he can start by ceasing to write and speak and instead clue in his audience to his preferences via behavior, since specifying them is somehow undesired.
Asimov's Three Laws exist to place boundaries around the genie's means of achieving our wishes. Russell's laws remove not only the genie's boundaries to means but also lets the genie make the wish as well. After all, it knows better than you what you want.
It's worth noting that most of Asimov's robot stories are him basically exploring the problems with the Three Laws; The Naked Sun is basically an entire novel showing how Three Laws robots can be capable of murder.
It's not worse.
> a date rapist
A rapist shares 99+% of definitions and values with you as a fellow human being. AI won't (unless you somehow program them in). Rapist has his evil motivations, but won't suddenly invent a virus that kills the whole human species because you asked him to stop neighbor kids trampling your garden. AI might. If you tell it not to kill anybody it might destroy our civilization to prevent us from killing ourselves with global warming. Why not - it only makes sense.
Simplest way to ensure safety for the maximum number of human beings is to anesthetize everybody and put them on life support till their natural death. Perfect record is possible - you might cure addicts and prevent crimes and wars. If you specify people have to be awake as often as they usually are - it can keep them awake but restrained. You want people to have freedom of movement? Ok - you just released the whole prison population :) Keeping criminals in prison is OK? Then it may make everybody a criminal for a quick fix. Or just drug everybody to WANT to be restrained. And so on, and so on.
There's infinite number of possible courses of action that we discard without consciously thinking about them because of our assumptions. You have to put these assumptions in the AI, each and every one of them, and they are very subtle and invisible for us most of the time. And they often border on philosophy and morality, and defining them is political by definition.
It's probably impossible to code all our values and assumptions in by hand. That's a much bigger problem for safe general AI than a post-factum explanations that rub your morality the wrong way.
> Asimov's Three Laws
Are self-contradictory and useless for anything except literature.
Russell's laws encode NO limits and doesn't even demand the AI check its judgment with the so-called beneficiaries of its decisions. That is most assuredly worse.
There was this thing called "cold war".
Have you heard of https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alar...
Thanks to the human element the system has shared human values and decided not to start a thermonuclear war despite it being the recommended course of action.
If it was a completely AI system it would probably just start the war.
Look at the Schindler's List Netflix example. Also, "A person can do what they want, but not want what they want."
1) AI-to-Zuck alignment problem: Align AI to it's master(s).
2) AI+Zuck alignment problem: Align AI and it's master(s) to the rest of the humanity.
Zuck is just stand-in variable name for any tech billionaire CEO, corporation or governing body, could be people running OpenAI, Google, Facebook, Microsoft, China or Pentagon. 2) seems like the problem for humanity and 1) as the problem for Zuck.
Very badly chosen variable name, because it singles out one special case instead of reflecting the broader concept, which could confuse other devs picking up from here.
It's like defining a variable that stands for "vegetable" and naming it "carrot".
Selecting easy to remember overly specific representatives is just way human mind works.
Great so the robot would see climate activist flying around in private jets pumping out several thousand times the average persons output, and promptly destroy the environment by assuming our actual goal is to destroy the environment with more CO2.
Do not mistake the two.
To bring it to a more human scale, an alcoholic can both sincerely desire to stop drinking alcohol and also drink themselves into a stupor every night, and a "sufficiently advanced intelligence" won't sit there spinning trying to figure out exactly which statement "The human loves alcohol" and "The human hates alcohol" is true to the total exclusion of the other. (And I mean this as merely one example dimension, of which the "global warming" problem contains hundreds or thousands of such issues at a minimum.)
I support the idea of being physically fit. My behavior belies that.
Robotics laws are unbreakable limits hardwired into every positronic brains. On top of those you can put some other programming (servant, mining machine, space ship...). It is not possible to create brain without those rules.
Asimov universe explores machine learning, autonomous weapons, humanity etc using those principles.
New proposed laws are just tautology, "do what people want". No limits, no orders...
What should machine do if no people are around? Just sit idle?
What if majority behaves in "wrong" way (vote extremist politician)?
I would expect some other laws, for example there should be always an option for people to opt-out or leave country.
a) The tyranny of the majority, or
b) The tyranny of those with strong preferences.
Let's say 51% of a population wants to get rid of the other 49% and wishes they were dead. What would stop the machine making it happen?
Let's say a super-person is born, who just wants stuff more than anyone else alive. Maybe he's had a bad childhood and been through terrible things, so now those things that he wants reach the level of desperation for him. Wouldn't this make the worth of his values more than other people to the machines?
Finally:
> Still, Russell feels optimistic. Although more algorithms and game theory research are needed, he said his gut feeling is that harmful preferences could be successfully down-weighted by programmers
... so we're back to square one. We've created a learning machine that can come up with its own morality. But... we're going to have to down-weight certain behaviours just to be sure. Doesn't that simply recurse to:
a) Down-weighting removes the learning, and
b) We have to trust ourselves to correctly program and define the down-weighting.
I’ve blogged about this recently.
Morality, thy discount is hyperbolic: https://kitsunesoftware.wordpress.com/2020/01/08/morality-th...
Normalised, n-dimensional, utility monster: https://kitsunesoftware.wordpress.com/2018/01/21/normalised-...
Disclaimer: although I do have a formal qualification in philosophy, I did not get a very good grade.
Have you seen this cartoon? It's something I remembered from way back, and finally found it again!
That SMBC looks familiar, but I can’t be sure — the comic has so much interesting philosophical and transhumanist content, it can sometimes blend together.
1. The machine’s only objective is to maximize the realization of human preferences.
Not all humans are the same. How would a robot deal with incongruent preferences? Or cultural differences that conflict?
2. The machine is initially uncertain about what those preferences are.
So am I... We often don't even know what our own preferences are. This has been studied. Given too many options of snow cone flavors, humans are less likely to even pick one! [0]
3. The ultimate source of information about human preferences is human behavior.
What?! The phrase "Do as I say, not as I do" comes to mind. Many people behave against their own intuition, best interests, or even their preferences, given the situation, peer pressure, or blackmail.
This "rewrite" is less preferable to Asimov's. Humans are fallible. I wouldn't want a robot to follow our lead.
I think the part about "robots could learn what Russell calls our meta-preferences: 'preferences about what kinds of preference-change processes might be acceptable or unacceptable.'" is what would be used to resolve preference conflict issues. People tend to be biased in consistent/similar ways so it doesn't seem implausible that a machine that could infer preference from action could take the extra step to infer circumstances affecting that preference.
Sort of like Google, Facebook, Amazon, etc. Millions of people use them, but they're only subservient to their owners and their preferences.
Imagine an AI that took on the preferences of a serial killer for instance, or an AI that takes on preferences of a genocidal warmongering culture.
Cherry-picking a detail from an article and putting it in the HN title is the quintessential kind of editorializing. Because threads are so sensitive to initial conditions, it ends up skewing an entire discussion. It also causes comments to make less sense when moderators come along and revert the title, as we've done here.
Submitting a story on HN doesn't convey any special right to frame it for other readers. If you want to say what you think is important about an article, please do that in the comments. Then your view will be on a level playing field with everyone else's.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
I haven't thought about this before on HN, but it makes a lot of sense. I'm curious if you or others have written about this intuition and your experience with it -- I'd like to understand it more.
Nothing in here seems to be an advancement towards AGI. I'm sure their robot can learn the optimal rules for driving in a simulator, but what about the unexpected in the real world? We don't even have a clue as to how human creativity actually works, much less how to make a creative machine (ie an AGI).
Simply watching a video is not an expression of agreement nor an endorsement of the content. Any of us read articles and books with which we may disagree if for no other reason than to fully understand the opposing viewpoint, and this is no less true for videos, audio, etc etc.
Want to know if the user actually likes the content? Ask them. Stop thinking that statistical inference is better or even equal to directly measured data from a specific entity, especially when dealing with qualitative concepts rather than quantitative.
Frankly, I feel there's major problems behind the concepts of AI and even intelligence itself - and it's difficult to articulate why. It’s as if these terms require aggrandizing to the point of impossibility or they lose all their apparent meaning. Which is why I feel we'll never achieve what we call (Strong/General) AI, or if we do, we will find ways to be unimpressed by it...
It is a recent trend to brush off Isaac Asimov's laws without an actual critique. Which betrays a lack of concrete thought in the matter.
It's not that Asimov's laws are flawless, but that countering them seems easier than it really is. I have seen various bad, hand-wave dismissals but I can't recall any careful critiques.
I have the new book, Human Compatible, and I buy Russell's argument but note that his rules are much more abstract, and therefore it is perhaps harder to counter.
Every single one of them fails in dramatic ways and that's what drives the story forward.
AIs are programmed to do what people ask unless it's bad, and one day, that "unless" mechanism fails. Joe starts doing everything people ask of it. Stalking that would make Zuckerberg proud....
then what?