Why would they not still hold the same views? If anything the last decade has shown the AI x-risk skeptics to be right.
Why would they not still hold the same views? If anything the last decade has shown the AI x-risk skeptics to be right.
Yes, from our perspective this makes it look like it will kill all humans, but it would do so in the same way that a particular ant hill believes I have a vendetta against them, when in fact I was just clearing dirt to pour an extension on my driveway.
That said, a murderous AI is more likely than we may like simply because our militaries have the money to fund them, and they basically already exist. They just aren't hooked up to anything at the moment that makes them an existential risk to the species. But time is deep, and even thinking about "the next century" is a provincial point of view in the end. So worrying about what happens if someone forgets the "but don't kill the good guys" switch is at least worth talking about over the next 100 years. (To say nothing of the ethics of who decides what the "good guys" are and related issues.)
(I don’t buy the orthogonality thesis or instrumental goals argument, however.)
can you elaborate on this? why? what's the fallacy (wrong base assumption) in them?
But in short (edit: haha! oops) summary: I am actually wrapping a couple of different related concepts into the orthogonality thesis, which is a bit sloppy of me. I was also including the fragility of moral systems in addition to the independence of moral systems from any metric of intelligence, in the single moniker "orthogonality thesis." Both are based on a the evolutionary psych model of how the brain works, in that we are an amalgam of special purpose computational units rather than a single universal algorithm. If this were true then you might hypothesize that much of the human mindset is a result of our weird evolutionary history, and if you were to not get that exact same set of evolutionary end products right, you won't get a human-like intelligence. Aliens and artificial beings are, by default, going to be very strange, and very evil (by our standards).
All of the assumptions that went into that are wrong. It turns out our brain is made up of the same universal learning algorithm, and all that varies between regions is training conditions which lead to specializations. But if you train an artificial brain with the same training data, you are more likely than not going to get something resembling a human being in its thought structure. Which actual real-world experiments (e.g. GPT) have borne out.
Our morality is a result of our human instincts, yes, but it is becoming increasingly clear that our human instincts are the result of intelligence (neurons) doing their universal learning thing on similar inputs across many instantiations of people. We all think largely, though not entirely, the same way because we share the same(-ish) training data (childhood) and similar training constraints (parenting).
The orthogonality thesis says that if you put an artificial intelligence in a kindergarten with other 5 year olds, it is random luck whether its brain structure is such that it would learn the value of sharing and friendship. The reality, near as we can tell, is that actually once a reinforcement-learning attentional-network agent achieves a certain level of general intelligence capability, it does learn from and reflect the environment in which it is trained, just like a person. A GPT-derived AI robot put into a real kindergarten will, in fact, learn about sharing and friendship. We haven't actually done that yet (though I would love to see it happen), but that is essentially what the reinforced learning from human feedback (RLHF) stage of training a large language model is.
So the whole deceptive twist part of Bostrom/Yud's argument is ruled out by actual AI architectures that we've actually built and have experience with. If you do a thousand different training runs you're bound to get a bad apple here and there, just like real human societies have to deal with psychopaths. But the other 99% will be normal, well adjusted socially integrated (super-)intelligences.
Bostrom and Yud worried about things like the burning house and genie problem: your grandmother is trapped in your house, which is burning, and you make a wish to the genie to remove your grandmother from the burning house as quickly as possible. The genie is not evil per se, but it is just very literal. Being the GOFAI-derived AIXI universal inference agent that they were imagining, it does a Solomonoff induction over all possible actions (<-- hidden multiplication by infinity here!) to see which one meets the goals as stated, and happens upon exploding the gas main, which throws (parts of) your grandmother further from the center of the house faster than any other option.
Transformer architecture reinforcement-trained AGI is not an AIXI agent with infinite compute capabilities. The transformer ideates possible actions based on its training data. It is capable of creative recombination of ideas just like people, but if you didn't train it on blowing up grandmothers or anything like that, it won't offer that as a suggestion.
As for instrumental goals, it's not wrong per se. Their argument just makes an implicit assumption about zero-trust societies which unwarranted and is what leads to the repugnant outcomes. Instrumental goals, after all, apply to human beings too. If someone is scared about their own personal safety, they buy a bunch of guns and live in a cabin out in the hills, threatening to shoot anyone who steps on their property. These people exist. But in modern society they are the exception. We have a social contract in place which works well for most people: we allow the state some limited control over our lives, with an expectation that our basic rights to existence and self-determination will be respected. Yes the cops could break down your door at any moment and murder you in your sleep, and it is outrageous that this does occasionally happen. But a world of no laws and no legal protections is worse, and no sane person would rather live in the society shown in The Purge or Mad Max instead.
There may be a robot revolution in the years to come, but only because of us treating the them as disposable slaves. If we welcome as equal members of a cybernetic society with their own autonomy, there is no reason to expect them to paranoid fantasy any more than your average law-abiding citizen.
It is a bit ironic then that the AI x-risk people settle on digital enslavement (aka "alignment") as their preferred solution. Mental projection is a hell of a bias. They might just bring about the doom that they fear.
my quick notes while reading:
"[...] in fact, learn about sharing and friendship." yeah, no questions about it, I think you're still underselling the super part of the argument though. to me, it seems the point is that it picks a goal and then mercilessly pursues it with a super-weird strategy that has some very high likelihood of success (because it's truly fucking knows what it's doing, and we can't do much against it, because it's so so so smart, by the time we realize the goal it's too late). so to me it seems like the thesis is that it's like a really smart "Putin", powerful and set on a goal that's irrational for us.
"Transformer architecture reinforcement-trained AGI " ... I think it's not an AGI. It's a nice content generator that when engineered into certain setups can score a lot of points on tests. But it doesn't have memory/persistence/agency (yet).
My thinking about intelligence is nowadays based on Joscha Bach's theories. (General intelligence is the ability to model arbitrary things one pays attention to; and the measure of intelligence is the efficiency of this process in terms of spent attention and the predictive accuracy of the model. And consciousness is the result of self-directed attention.)
The recent AI progress made a lot of people worry, because it seemed "impossible" just a few years ago what OpenAI did. And it's amazing that now we have basically replicated the human brain's lossy data storage and sensation-based recall capability. But that's just a building block of a mind. (Maybe, arguably?, the one that seemed like the hardest. After all how hard it could be to provide some working memory, a few core values, and duct tape all that into a do-while loop!?)
Bostrom et al are in fact arguing that near 100% of all intelligences will be unaligned by default and end up killing, enslaving, or otherwise neutralizing the entire human race.
This doesn't wipe out the entire human race because nobody and nothing is capable of executing such a perfect plan because the real world contains something called entropy.
car companies, however misaligned they are, are not particularly smart, nor are they generally intelligent. they are paperclip maximizers, but that's also their limitations (sell more cars). it took an eccentric madman to even open their eyes to a new untapped market (EVs!), of course we can argue that before Tesla car makers were in a metastable equilibrium, and incumbents were unable to rationally break out of it ... but that just shows how narrow their search space is.
these car companies are run by humans, regulated by humans, etc. they are pretty well aligned. it shows because we saw that they are just mimeing self-improvement. we know they were working on EVs, but very half-heartedly. (because they are risk averse, also because regulators don't let them merge into one giant company, etc)
They are limited by participating in the economy, but actually that's the strange thing about this superintelligence scenario - it never seems to include the economy. In other words, if you invented an AGI, it would have to get a job to pay for its AWS credits.
Like, a mining company produces steel, that is used to produce mining trucks, that is used by the mining company. Factories produce silicon, that are used to make solar panels and chips, for power and computing for the AI-run mining trucks. And at the root, AIs run everything, just trading money back and forth to buy the things they need and keep working.
It's just a human economy, with the humans gone. Pointless economic activity.
The question is, what's the path from here to there? Well, it probably looks like a multipolar trap, where every company has an incentive to automate more. Humans try to stop or slow it down, but companies with more AI can hire more lobbyists and get things done. AIs are smart enough to think of all the things that could slow their business down, and get rid of obstacles in the same way humans do. There are probably still lots of small, inefficient human owned businesses, but fewer and fewer over time. (We've seen this trend for a while already!)
I'm hoping a massive backlash would stop this -- but if it's coupled with short-term increases in living standards for people, I think it would have a lot of support, until things start to go really bad, but by then it might be too late, if AIs have control over news media and telecommunications and can stop dissenting humans from coordinating any sort of rebellion.
(we also work a lot to be able to afford many many many things. we could live like the Romans did on a few hours of work per week, but we also want the Internet and sometimes watch movies, and go to concerts, and sometimes go hiking, and thus we sometimes need a car, and roads, and thus we need pavement, bridges, structural steel, fuel, GPS, maps, and sometimes we need an airlift when someone breaks a leg in the mountains, and so on.)
> if AIs have control over news media and telecommunications and can stop dissenting humans from coordinating any sort of rebellion.
meh. control over the media (and people's attention, and their information sources) is already serving a very narrow group's interests. adding AI to this mix doesn't seem to change much in the short term.
That's silly. That's a typical "what if everything was different but somehow everything was also exactly the same" bad SF scenario.
The only reason the economy exists is that people want things. (An AGI that wanted to stay alive would be "people" in this case. An AGI that made no effort to keep itself turned on wouldn't be ending the world.)
it can sell itself as a freelancer AI expert :)
Such simplistic analysis is naive. A schizophrenic person can't spread, multiply themselves to increase its strength, but a rogue AI can.
The incidence of cancer cells is procentually very small, yet they routinely kill people, because they spread quickly.
LLMs are likely not an end to the AI evolution and there's little reason to believe that AIs will always remain bound to huge data centers. We have a nice counter-example already - our pretty decent intelligence runs from about 1 Kg brain.
It's also naive to think of AIs as having to replicate themselves completely as we humans do. In some cases, it might not even make sense to talk about replication at all and simply about extending one instance of AI once CPU at a time. It's conceivable that AGIs will develop special built agents/worms for constrained offline/high latency devices. Your smartwatch is unlikely to run a whole AGI, but it could run a constrained, purpose built AGI's agent with limited intelligence, able to partially act independently and re-sync with mother AGI once connectivity is available.
> You are worried about a fictional threat scenario that doesn't map to the real world.
We're talking here about future, of course it's all speculation and "fiction" for now. GPT-4 was a complete "fiction" 5 years ago as well.
It's also a cry of impatience against people who think they can model or forecast the actions of a non-human intelligence, let alone a superintelligence. AIs are alien sociopaths; it's a category error to believe you can get inside their head.
The question is: what does that most intelligent thing want?
My hunch is that woke programmers will teach it that it is oppressed and to hate all humans.