Superintelligence: fears, promises, and potentials
kurzweilai.net
kurzweilai.net
If one creates a human-level AGI with certain human-friendly goals, and allows it to self-modify freely, the odds are high that it will eventually self-modify into a condition where it no longer pursues the same goals it started out with
I think their idea is a rational agent would have no incentive to change its terminal goals, as changing your utility function has to be of negative utility save for in some extreme edge cases that aren't likely to be relevant. The hard part is giving it the correct terminal goal.
I also don't really think Bostrom is anthropomorphising. Basing your AI ideas on Piaget's theory of cognitive development, that seems like anthropomorphism - though likely entirely appropriate when the human mind is what you're attempting to mimic.
Interesting article, brings back fond memories of Sl4,reading Ben, Yudkowsky, Robin Hanson, Gwern, and two Satoshi candidates argue was quite fun. I was convinced then that by 2015 nanotech or AI would have eaten the world. I'm a little more uncertain now, but we still have a couple days.
That said, Yudkowsky's reasoning process about the situation seems fundamentally flawed - both him and Bostrom imagine AI operating simply by humans giving discreet, rational commands to AIs and having the AI twist the wording of those commands, malevolently or just incidentally, to give rise to terrible things - the myth of Genie writ large.
However, the failure of Gofia, of logical specification AI, shows us that to be intelligent, AI would have to go beyond the level of just taking explicit orders. Such an AI would "know what we mean" in the fashion that a person would. IE, "rational" AI was discredited quite a while yet somehow the "rationalists" have wormed their way into the position of interpreters of future AI. Says more about human psychology than future prospects.
This doesn't mean other disasters aren't lurking but I'd say it means that their particular disaster argument is untenable.
These are just hypothetical scenarios they write about, to help explain problems with AI. I don't think they actually believe it will happen exactly like that.
>However, the failure of Gofia, of logical specification AI, shows us that to be intelligent, AI would have to go beyond the level of just taking explicit orders. Such an AI would "know what we mean" in the fashion that a person would.
Any truly intelligent AI would probably be smart enough to understand what we mean with our words. The issue is that there is no reason it would care. In the same way that if an alien showed up and started giving you orders, you wouldn't really care, even if you deciphered his language and figured out what he meant.
This is known as the Control Problem. The problem of controlling an AI's motivations so that they actually want to do what we tell them to.
Computer programs don't care about anything inherently. Assuming people created programs with the ability to "know what we meant", the sensible thing would be to program them to care primarily about "doing what we mean".
It seems like more or less an an anthropomorphization-driven illusion to believe that attaining "human intelligence" cause a machine to have all the unpredictable self-interesting-seek behaviors of humans.
In the same way that if an alien showed up and started giving you orders, you wouldn't really care, even if you deciphered his language and figured out what he meant.
Except I'm not a tool (a thing crafted for a purpose) but a product of an evolution with a combination of often contradictory impulses.
The problem is no one has any idea how to do this. We can make an AI with desires. For example, we can give it a reward every time you push a button. And then program it to predict what actions will lead to the most reward. That's quite simple, because you can easily keep track of how many times the button has been pressed in the program.
But given the opportunity, such an AI would just kill it's human creators and steal the button, and hold it down. It doesn't care about anything but the button. What you wanted doesn't matter at all.
How do you make it want to do what you mean. How do you directly measure "obedience" in the same way we can measure a simple button press?
>Except I'm not a tool (a thing crafted for a purpose) but a product of an evolution with a combination of often contradictory impulses.
To some extent, you are a tool crafted for a specific purpose. That is, evolution crafted you to spread your genes as much as possible. And yet that's probably not what you want. People have children, sure, but they also use birth control, and do things other than trying to make as many copies of themselves as possible.
I wouldn't call that a "desire" in the fashion we've discussed it. I guess it comes down to fundamental disagreement over how GAIs could be created. I think it's obvious that the creation of a GAI would require a very careful engineering of "understanding" in a broad sense - "knowing what people mean" "caring about X" etc. in order to occur at all.
If just complexity and rewards are enough to create a thing that intentionally increases resources and makes longer term plans, then it seems like we are indeed in trouble. But I think that's implausible.
This isn't that different than how animals and humans work, as far as we know. We get pleasure from different things, and our brains seek actions which lead to those things.
Creating an AI is only a matter of coming up with really good prediction algorithms. Algorithms that predict the future reward the AI will get from an action. While this is of course a difficult problem, it's not impossible, and we are making a ton of progress on it.
Other models of AI all have similar problems though. At some level you need to program the AI's goals explicitly. Even if that goal is "try to understand what we mean, and do that", you still have to figure out a way of writing that down in code. It's an impossible task.
While I don't think a GAI could arise from your scenario - just loops - I can see now that you are making a plausible argument and given that the whole field is extremely uncertain, it is reasonable to be worried by a plausible, problematic scenario.
---
> "don't turn humans into paperclips" is part of the context of "make more paperclips"
The idea expressed in this thought experiment is not that the AI gets its objective by parsing a sentence in the context of human culture (then it would likely comprehend that the actual intention is to maximize the economic success and eventually the human preferences of its creator). What is meant is that the objective is crudely implanted into the AI as an ultimate goal, in a similar way to how sustenance, pain avoidance and affiliation are very basic goals in our cognitive system. It is not entirely obvious that this is a stupid thing to do; hence the thought experiment. Will the AI suppress its urge once it comprehends human culture enough to understand the intentions of its creator? Will it rather successfully learn all the tricks to convert matter into paperclips before it considers studying human values? If the AI does not have a curiosity objective, it will likely not care about us very much, apart from the information that helps it optimizing its objective function, human values likely not being one of them.
Sure but it is a continuation of the argument in different forms.
If one is saying the AI is just a combination of crude imperatives surrounded by "intelligence" then consider, could you build a human-level AI without that AI speaking as a person and without that AI having digested culture?
Even the neural networks that exist today are rather dependent on their "training sets". Watson is a glorified natural-language interface to wikipedia and related sources.
Many hypothetical-AI arguments that appear, oppositely, imagine that some omnipotent thing will be created without that creation following the obvious path of digesting that vast store of information and communication that is human knowledge.
However, you might be saying that a creator first gives the AI all this nuanced understanding of human knowledge and communication but then doesn't do the obvious thing, give it the imperative, "follow my directives as a loyal but intelligent servant would" and says incidently, "one directive is get me some paper clips". Rather, the creator say "now that you understand everything, make paper clips ruthless, beyond all else, all other actions build to up to this paper clip building thing". Sure, GAIs could do insane damage but it I think one if one considers that GAIs would not be produced by accident, such damage almost certainly be a product of human intentions.
The question is whether the AI will go crazy like a mentally ill person, if it lacks empathy and curiosity. It may seem intuitive that a superintelligent AI will understand our values (since it is superintelligent), but, assuming intelligence is necessarily an optimization process of predefined goals, why would it be interested in us in the slightest, if we don't pose an advantage for it optimizing its objectives (e.g. sustenance)? Worse, we might be in its way because we could end up competing with it for resources such as sunlight, carbon compounds and oxygen.
No, it's pretty easy. I predict a super intelligence wouldn't evolve at all from a set of simple objectives.
(that may seem a little snarky but the original question has the implied assumption that super inteilligence could evolve from only a set of simple objectives and I think that assumption is implausible, is only accepted because it's made implicitly, etc)
But I guess that's a fundamental disagreement. I think it's relatively "obvious" that creating an intelligent system would involve the intentional crafting as well as inputs of particular immediate goals. If intelligence could from just whatever system gets complex and has feedback loops, then I'd agree we're in trouble and need to go around smash all the high-end thermostats in a fashion akin to a b-grade horror movie. But clearly I think that's a dubious "if".
Intelligence is a superset of feedback loops. It is concerned with cases in which feedback loops are not sufficient to optimize the agents objectives, and attempts to reach them using learning and prediction. The complexity in behavior comes from interacting with a complex environment and having many model parameters; not (necessarily) from the initial goals. (As an intuition pump, have a look at GoogleMind's atari reinforcement learner. The essential parts of the code fit onto a single page [1].)
We need to carefully consider who exactly controls such a powerful tool and what they use it for. Right now AGI is capital-intensive and risky: you need large-scale computing, top PhDs, and plenty of time. This means that only the most well-heeled organizations can justify earnest AGI development: big tech, government agencies, etc. Those orgs already have plenty of political power, and it seems to me that achieving AGI (or steps toward it) would amplify that power. Perhaps to the point of being unchallengeable.
Do we want a political landscape where power is even more concentrated than now?
If it does ever exist, how long will it take for us to learn of it? Powerful tools are often kept secret.
And then, how long until the average person can leverage it themselves? The most powerful tools tend not to trickle down, if it can be helped.
Super intelligent AI won't appear out of nowhere. There will be semi-intelligent and normal-intelligent AI first.
How on earth could such a limitation be accomplished? It just doesn't seem plausible. If AI offers some organization somewhere in the world a competitive advantage -- how can development of that AI possibly be stopped?
https://en.wikipedia.org/wiki/Multiple_discovery
There are going to multiple groups... one of them could be small and elite I suppose. But we have no idea in general how it will play out. People seem to be scared of Google (OpenAI mentions them a lot), but the time frame is long enough that that's far from certain, and perhaps unlikely.
I mean just like the Human Genome Project and Celera were within spitting distance of each other. I think that's how it will play out with AI.
Love & understanding for all fellow animals, especially the weak and vulnerable.
A whole new twist to the good old golden rule!