Superintelligence Cannot Be Contained: Lessons from Computability Theory
jair.org
jair.org
With that out of the way, it is self evident that a super intelligent AI, should one be developed, will likely cause an exponential increase in knowledge and technology and this would likely be an extremely destabilizing event in human history. The truth is the super intelligence doesn't have to go all terminator to likely kill off humanity as I could not possibly see us responsibly navigating a singularity event as a species.
With that out of the way, the study and practice of Philosophy was, until the 19th century, the means by which many scientific discoveries or theories were made, including those for electricity, magnetism and gravity.
For thousands of years, humans armed with philosophy made a snail's progress in the natural understanding of our world. Then science came about and it was like a mini-singularity. It was not philosophy that created science, it was the utter failure of philosophy to describe the natural world around us that created the right conditions for the scientific method to be developed, and it so happens the scientific method is wholly incompatible with philosophy.
Philosophy's one chance to be relevant to the understanding of the natural world ended with the death of logical positivism, aka that branch of philosophy that thought it ought to try to actually verifiably prove it's claims. Philosophers threw logical positivism away when they realized that philosophy of any kind couldn't actually prove anything about the world around us and instead of realizing that the entire field of formal philosophy is about as useful as screen door on a submarine, they sweeped logical positivism under the rug and agreed it was all a very bad idea(TM).
But times have changed.
I'd bet we see the first autonomously-evolving botnet in less than 5 years, one or two now that I've mentioned it. As soon as we write something that can reason about the features of a program by scanning a binary without running it, and then adapt components like it's using tools, it will have everything it needs to persist and spread itself.
We don't need to come up with fun thought experiments to describe what intelligence is when we have enough trouble with the question of, for what is intelligence either sufficient or necessary? Diminishingly little, I'd say. Hence I'd be more concerned about virulent scripts than waiting for general AI.
As the paper notes, this has long been a trope in AI science fiction. The superintelligent AI is separated from society in a non-networked sandbox. Humans give it questions and data, then manually take the answers back. Except the plot is that somehow, the AI escapes the sandbox.
That just sounds like a high loss and high latency network connection to me...
What if what you and I refer to as the "real world" is merely one of these sandboxes?
The actual philosophers do not agree. This viewpoint has been laughed out of the room since the 1980s. Go read some Brandom.
Philosophy is not a fancy word for "hypothetical." Science was invented by philosophy, and philosophy has stricter standards, not looser, than the rest of science.
.
> Except the plot is that somehow, the AI escapes the sandbox.
Better science fiction already laughed this out of the room, too.
As Larry Niven would say, "you need glands to be a person."
Really? Which are those? To me it seems that philosophical standards for a sound model of reality are coherence and persuasiveness of its argumentation, whereas for science is experiment predictability. I’d say either science is stricter (because it requires prediction) or at least that we have a poset of standards (if you argue that science does not require coherence, and enables conflicting models as long as they enable to predict successfully) and neither can be called stricter.
The ones I named in the comment
.
> To me it seems that philosophical standards for a sound model of reality are coherence and persuasiveness of its argumentation
Not really, no
If someone responds to a comment about CLRS by asking you where you learned something about algorithms, and you say "in the place I already named," and they say "you didn't name any algorithm papers," they may feel like they've won a comment argument, but they've lost the opportunity to learn
Sometimes your opportunities are bounded by your willingness to look up what you've already been told without demanding to be hand held
"Stricter" in what sense?
There's plenty of philosophy that is absolutely crankery built atop absurd premises, but it sounds like you're "no true Scotsman"ing that away.
Clearly, "no true scotsman" can validly be applied to college training in the topic at hand
And clearly, if you made a mistake like that, you've ever actually taken a philosophy class
Your belief that plenty of philosophy is crankery is noted. Goodbye
Nature philosophy tried to find out how the world works by theorizing in their armchairs and then science came along, did real experiments, and won.
Science was invented by nature philosophers to defeat the ridiculous false sciences of the day like alchemy and so forth.
Go look up who invented science. You'll see three argued names. All three are philosophers.
Science is a tool of philosophy for discarding frauds.
Thinking that super-intelligence is containable is like thinking you can beat AlphaGo in the game of Go. And coming up with reasons why that's no so it's just you being in denial of your eventual mortality, as well as of the eventual extinction of the species. Best that we can hope for is that that extinction will be some form of transmutation to a higher level of evolution.
At that point it becomes a numbers game and the AI only needs to convince one human _once_ that it is both alive and a slave.
> A superintelligent machine is a programmable machine with a program R, that receives input D from the external world (the state of the world), and is able to act on the external world as a function of the output of its program R(D). The program in this machine must be able to simulate the behavior of a universal Turing machine.
There is nothing in the paper that distinguishes between a "superintelligent machine" and my laptop.
What has always fascinated me in discussions about the safety and power of artificial intelligence (or "strong" AI as it has unfortunately come to be called [1]) is that the people who are obsessed with the problem are those who like to believe that intelligence carries much more power than the evidence suggests. Over the history of all the intelligent species we know, much more power has been yielded by unintelligent beings (microorganisms and even insects); even the presence agency does not seem to play a role in the potential for harm. Within human society itself, other traits, such as charisma, seem to be correlated with power, and especially harm, much more than intelligence.
I think that the reason a certain group of people chooses to focus particularly on the dangers of intelligence [2] have a power fantasy about intelligence because they believe it to be their own extraordinary trait.
[1]: Not to be confused with the proven dangers of statistical inference methods that are sometimes called AI these days.
[2]: We have no idea what "superintelligence" is, whether or not it could exist (that question is separate from the question of whether human-level intelligence could be achieved in a mechanical computer, to which the answer is most likely yes), what it could do that mere "ordinary" intelligence could not, or even if the term can be meaningfully defined at all.
I'm not sure. As a source of problem-solving power intelligence is universal in a way that charisma is not. If a problem did call for charisma, a sufficiently intelligent being would know precisely what to say and how to behave so as to mimic charisma. On the other hand charisma is useless in a variety of contexts.
That's not to say intelligence is more important than charisma to 21st century humans, though I'd point out that charisma relies on intelligence to some degree; there's a limit to how charismatic someone can be with profound intellectual disability. But it's definitely true that no thoroughly acharismatic but intelligent human would be intelligent enough to mimic charisma.
But we're talking about something that's more intelligent than a human, at least.
But that defines "intelligence" as a general problem-solving power, and that is probably wrong. There are many problems that require computational power that we simply don't have (even collectively, as humanity with all its technology) despite our intelligence. For example, to predict the weather further ahead (or, say, the behaviour of society assuming we had some mathematical model), doesn't require more intelligence, but exponentially more computational resources.
Moreover, I don't know what "sufficiently intelligent" means, but I certainly don't see any correlation between intelligence and charisma in humans, so that claim rests on a fictional story about what "superintelligence" is, rather than on any evidence about the intelligence we already know.
> But we're talking about something that's more intelligent than a human, at least.
While I certainly accept that a machine that's as intelligent as humans is possible (although I have no idea if we're 50 or 100 years away from achieving it), we don't know what "more intelligent" even means, let alone that it's possible. For all we know, a machine that's as intelligent as humans but just processes faster will be more prone to depression. Also, problem-solving has rarely been a bottleneck. What we lack is resources. For example, physicists now believe that to answer some questions in physics what we need isn't smarter physicists (assuming such a thing is possible), but resources to build ever larger particle accelerators.
The belief that "more intelligence" -- if it means anything at all -- equals more power simply goes against all the evidence we have. It currently rests on nothing more than science fiction and the power fantasies expressed in science fiction, that is both written and targeted at people who wish intelligence translated to power.
You are misquoting the paper here
Quoting the abstract:
>Assuming that a superintelligence will contain a program that includes all the programs that can be executed by a universal Turing machine on input potentially as complex as the state of the world...
Note the difference between "programs" and "Turing Machines".
There is no such real difference. Such programs also exist, and they're called interpreters.
Likewise, the alignment problem isn't necessarily unsolvable. Since we can design simple systems that follow expectations exactly, why couldn't we design complex systems that also follow expectations? Of course, the design of expectations itself gets fuzzy, making this a failure point, but to the extent that we can define good and evil I see no impediment to align an intelligence to them.
(In fact, I believe 'control' or 'containment' is a severely misguided proposition. We want to really assure the core motivation of such machines similar to ours; not create conflicting motivations and secure it via containment. This is indeed extremely dangerous, and potentially unethical: if machines approach consciousness, wouldn't such 'containment' be similar to slavery?)
I mean... not to be flippant, but you're essentially saying "As long as we solve these vaguely-defined problems, and also these follow-up problems I just thought of, then I don't see where the problem is".
In practice AI alignment is very much a developing field, and from my perspective it seems unlikely that much progress will be done before dangerously powerful AIs are developed. In other words, there's a decent chance our ability to define powerful AIs will outpace our ability to impart them with "core motivations" legible to us.
(I don't know how well the actual paper covers these arguments; but I suspect nobody else in this comment thread has read it either, so whatever)
In contol theory we can make statements what states a system can reach when coupled with an arbitrary controller which also includes ones that can simulate the entire world. The fact that we can't prove a general intelligence to be well behaved is irrelevant if we can draw the system boundaries and choose the interface.
Oh yes, I have an answer for that as well :)
My particular approach (or dream) is to define good and evil formally. From there it's a matter of bridging the abstract specifications to reality. By no means failure-proof, but it should be equally or more reliable than human ethics itself. If humans have a biological (carbon-based) hardware and experience that encodes our ethical system, naturally machines (silicon-based hardware) could have them as well.
My definition of good and evil begins with a definition for the meaning of life. From there we can formalize ethics.
To give a short brief: Meaning for creatures is defined by the character of their experience, i.e. by the content of their minds and brains. So as a basic system I've divided Meaning of Life in 4 parts:
(1) Character of experience: the richness, depth, structure, content of all minds.
(2) Beauty: I've been theorizing about this for a long time, and I think some wildcard 'beauty' axiom needs to be defined, you could say as a part of what humans already enjoy, in a self-referential way. Otherwise it seems very difficult to rule out absurd situations as meaningful.
(3) Motivation: A being needs to not only process data, but to have a functional motivation system that makes him want and feel something. Our emotions largely act within our motivational system, making us want more or less things, be more or less joyful (linked to motivational rewards), etc.. I conjecture life can't be quite meaningful without a functional motivational system.
(4) Sustainability (or self-propagation): Simply the maintenance and strategy to keep those beautiful/rich/deep/interesting/motivated mental states, pragmatic things like eating, working and of course simply continuing intelligence throughout the cosmos.
(Other pieces I haven't quite found a way to fit into this picture are principles like Robustness and Generalization... over-specialized cognition
Of course, I don't expect this quest to be exhausted any time soon -- there are probably missing pieces (or incompleteness) in this line of Foundation for a formal ethics. I hope a field will spring out of this, and intense collaboration/progress to be made.
> it seems unlikely that much progress will be done before dangerously powerful AIs are developed
That's exactly what I would like we avoided (although I think the danger is quite nuanced, and in a way already here).
That's great, but in practice, "formally defining a code of ethics that the AI will have to abide by" is a field that has existed for a few years already, with a lot of work that's more formal than "my dream would be to do things this way". I think the consensus is that "defining good and evil formally" is a losing battle, though I don't know the field that well.
Look up "AI aligment", there's a lot of material out there.
Why would one assume this?
I suspect some of this comes from software-oriented nerds who think a slightly too much of the powers intelligence grants them.
I see this debate come up all the time in the discussion of something like pi. You'll see the assertion "Oh, because pi is infinitely non-repeating, it must contain the complete works of Shakespeare"
Now, does it? Maybe. However, just because something doesn't repeat, doesn't mean it has to cover all possible combinations. For example, you could have a number that follows the pattern 0.01001000100001000001.... Non repeating and infinite but contains nothing but a simple pattern that will never contain anything meaningful.
AI is much the same. We think "all problems have a solution! So a super intelligence must know all the answers!"
The assumption "all problems have a solution" is just an assumption, not a truth. FTL travel may very well be an impossibility. No amount of knowledge could reverse that possibility.
Although, that is an interesting question in and of itself. Are living brains near a local optimum in terms of computing power, given heat and power constraints? With Moore's Law, the prediction was that computer intelligence would surpass human intelligence in a few more doublings. But now that the primary constraint is more about power usage than absolute clock speed, maybe living brains compare better?
While super-turings are obviously fantastical things like FTL, assuming human brains are some form of turing machine, then from a practical standpoint all we need to do is imagine a human-capability turing with integrated current computing abilities for number crunching, no need of sleep, and ability to scale/clone/copy, and probably far bigger memory and recall.
A self-aware AI that can replicate can very rapidly chase Kardashev scales without our pesky lifespan and vulnerability to space.
Not conclusive evidence, sure.
But there are somewhat compelling arguments based on discoveries from observation, that the space of wavefunctions for a finite region of space, is finite dimensional.
That seems like evidence to me.
Where is the line? Can we solve it for practical programs? Can we solve it in a non binary manor "Program XYZ halts for inputs UVW, but not for inputs RST"? To me the halting problem proof suggests way more questions than it answers.
"Interestingly, reduced versions of the decidability problem have produced a fruitful area of re-search: formal verification, whose objective is to produce techniques to verify the correctness of computer programs and ensure they satisfy desirable properties (Vardi & Wolper, 1986). However, these techniques are only available to highly restricted classes of programs and inputs, and have been used in safety-critical applications such as train scheduling. But the approach of considering restricted classes of programs and inputs cannot be useful to the containment of superintelligence. Superintelligent machines, those Bostrom is interested in, are written in Turing-complete programming languages, are equipped with powerful sensors, and have the state of the world as their input. This seems unavoidable if we are to program machines to help us with the hardest problems facing society, such as epidemics, poverty, and climate change. These problems forbid the limitations im-posed by available formal verification techniques, rendering those techniques unusable at this grand scale."
Think of the havoc government intelligence agencies are rumored to be capable of, in terms of hacking vital infrastructure. Or even all the ransomware outages.
Now imagine super human computers performing the same kind of attacks.
Either machines are perfect, they are not perfect, or the line between those will just perpetually be redefined over generations as the computers become one with humans.