Since very smart human beings have in history done a lot of damage, but never ended the species or anything, we can conclude that the world has an "intelligence safety margin" that extends up to Alexander the Great or Napoleon. There have been some incredibly smart people in human history and none of them have ruined everything, a few countries at most.
This has been discussed under "warning shots" https://www.lesswrong.com/posts/idipkijjz5PoxAwju/warning-sh...
Gain of function research is supposed to help with pandemics, right? Did it help with COVID?
No, not only has GoF not helped with COVID and probably caused it, but that's also true of the entire field of virology.
At best, history serves as a counterexample to the idea that if an AI goes bad the attendants would just unplug it, seeing how often it is that dictators don't get stabbed by their aides as soon as they start causing mass deaths, instead often receiving broad popular support as the world burns.
Every step of this relies on assumptions that are not merely questionable, but unfalsifiable.
Whether AGI is possible or not, regardless of anyone's personal opinion, is as yet unprovable and unfalsifiable.
Assuming that AGI itself is possible, there is no way to tell whether we, as humans, can create an intelligence that is "smarter" than we are.
Assuming that we can create an AGI that is "smarter" than we are, there is no way to determine whether it would be able to upgrade itself to become exponentially smarter than that, and beyond human understanding and control.
If you have a hard time coming up with arguments against these things, maybe it's because they're fundamentally unfalsifiable, which makes them useless for trying to build any framework of understanding on.
But the AI safety groups don't just assume these without any justification. The existence of the human brain itself is either a proof and very strong launching point for most of these. In 1930, maybe you could have convinced someone that a nuclear bomb is impossible or at least an unfalsifiable worry, because one had never yet been built, but your reasoning is like trying to cast doubt on the possibility of artificial digestion, when there are already billions of stomachs roaming the earth.
The idea that a computer can't possibly be made to accomplish whatever a brain can is losing plausibility with every passing day. And as for surpassing it: a human brain is powerful, but has so many surmountable limitations, like: requiring 20 years of education to get up to speed with existing experts; dying after 70-90 years; not being able to run copies of oneself in parallel; etc. We only need to imagine removing these limitations, and doing so violates no known scientific principles.
It's like having a 50-megaton warhead in your lab, and meanwhile the rocket scientists are gradually figuring out how to make rockets fly, and you're saying "yes this is a big bomb, but the idea that one of these could be more dangerous if mounted on a missile is an unfalsifiable assumption!"
Consciousness, intelligence, sapience—these are not well-understood phenomena. We don't know what makes us conscious. We don't know if other animals are, or to what degree. It's not even possible to determine with any scientific certainty that another human being is conscious.
As things stand, "we can build an AGI" is not a scientific statement. It is not grounded on a foundation that allows clear reasoning about it, one way or the other.
Your arguments are not incorrect; however, they do not bridge the gap to "and so we can definitely make a conscious, sapient, intelligent computer." Being able to replicate particular capabilities of the brain is not the same thing.
And, again, that's only the first unfalsifiable proposition that must be satisfied in order for the purported AI threat to be real. They also have to be capable of breaking free of our control, decide we're a threat to them for whatever reason, and have the means to carry it out.
Consciousness is irrelevant, as it's clearly not necessary for AGI nor even human intelligence; otherwise you wouldn't say "it's not even possible to determine with any scientific certainty that another human being is conscious."
On what grounds do you believe that there are capabilities of the brain that cannot be replicated by a computer? When I look at the 1.5 kgs of matter in a typical human brain, I don't see anything that jumps out and says "my operation is not computable!"
At least with fusion, the high temperatures and pressures are a clear barrier. Our brains don't need to be held at 3.8 trillion psi and 15 million Kelvin in order to enjoy poetry.
I can certainly argue on the other points as well (breaking free / deciding threat / acquiring means) but you need to pick your goalposts one at a time. I think the first step would be asking yourself: if you were a super-genius but being held by guards in solitary confinement on a remote island with only a supercomputer connected to the internet, how would you earn money?
Those conditions clearly would not have stopped even an ordinary Satoshi Nakamoto from gaining control of sufficient resources to hire private military contractors to arrange his escape. I'm not sure what a superhuman would do, but that's a human baseline.
The argument goes like this: If you want to save the world, then doomsday scenario S is the best place for you to invest your resources if S has a nonzero probability, and currently we are investing less resources in S than in other existential risk scenarios (per basis point of probability).
"AI risk" is a pretty good candidate for S (especially back in 2010-2015 when the movement was just starting.)
> Whether AGI is possible or not, regardless of anyone's personal opinion, is as yet unprovable and unfalsifiable.
"AGI is impossible" is certainly falsifiable: All I have to do is build an AGI and show it to you.
Further, there is no theoretical reason AGI is impossible. Rather the reverse; consider these Well Settled Scientific Facts:
- (1) Human Mind = Human Brain
- (2) Human Brain obeys the laws of physics
- (3) We understand the laws of physics well enough to simulate them in a computer
If you accept these three facts, then in theory AGI is possible: You could implement AGI by building a machine implementation of the Human Brain by fully simulating the underlying physics.
> Assuming that AGI itself is possible, there is no way to tell whether we, as humans, can create an intelligence that is "smarter" than we are.
> Assuming that we can create an AGI that is "smarter" than we are, there is no way to determine whether it would be able to upgrade itself to become exponentially smarter than that, and beyond human understanding and control.
An AI safety person would say to this: "You're right, we don't know -- and that's exactly the problem."
What we can do is assign a probability based on our confidence. How likely do you think it is that we can do those things? 20%? 2%? 0.000002%?
If you say "0.000002%", what makes it so extraordinarily certain it's impossible? If you say "2%" or "20%" then as a matter of self-preservation, shouldn't our society be devoting a lot of money and smart people's time and attention to figuring out how to make sure it doesn't happen?
As well, while the arguments are logical, most of them rely upon large assumptions to move between steps. If any of these assumptions fail, the entire thing fails. Especially the hard take off assumption.
As well, the assumption that AGI is happening in the next 3-10 years. I’d say most prominent people in the AI research space don’t think we’re much, if any, closer to AGI. Yet you have Yud and LW screaming that we will all be dead in a few years and AGI is right around the corner.
When people like Chollet and Ng say we aren’t close to AGI, I’m more likely to believe they’re right, vs. Yud who hasn’t contributed to any actual developments within the field besides theorizing about alignment and how AGI can go wrong.
So unless one is proposing a global crackdown on AI research, AI safety is a lost cause.
There are lots of solvable problems that will never be solved if we just throw up our hands and say "we'll never get it right, might as well not try".
Any uncontrolled group that creates an AI can either find a body of safety research accessible to them, or not. Preparing the former is hardly a lost cause.
AI certainly has risks, I don't think any reasonable human being doubts that, its just that the AI doom cult seems to think the worst outcomes are near certainties without really backing that up.
What are you basing this on? There have been tons of arguments written about why these worst outcomes are likely. Read Bostrom's Superintelligence for example, or Yudkowsky's Intelligence Explosion Microeconomics.
Many conversations with AI doomers. They gloss over and make assumptions about intelligence that aren't really backed by priors and when this is pointed out they hand wave and say "but computer".
> Read Bostrom's Superintelligence for example, or Yudkowsky's Intelligence Explosion Microeconomics.
I don't really have any interest in doing so, and if I'm honest have a particularly unfavorable read of Yudkowsky as a person based on his cultish following.
The orthoganility thesis is an unproven assumption that intelligence and goals are not corellated, meaning an intelligent being can pursue stupid goals. States like that, it’s obviously wrong and laughable. But by using complex language, EA cultists hide the ridiculous assumptions their system has so that they can maintain their feelings of superiority while gaining real power that enables them to abuse others.
I maybe agree with you that there's an this belief that a maximally intelligent creature will blindly follow maximally obviously stupid goals, and that belief is under-argued, but your phrasing above isn't the slam dunk argument that you seem to believe.
You could conceive of a super intelligent AI that came into existence with the goal of terminating itself. That would be an "stupid goal" from our perspective, since we have the goal of self-preservation really ingrained in our brains.
But for a being that self-termination is the absolute best thing ever, it's not stupid. It makes perfect sense, since, well, that it's goal. It doesn't care about self-preservation, it doesn't care about becoming more intelligent/rich/powerful, other than as an instrumental goal to help achieve self-termination, if it's not able to do so in its current state.
And most importantly, no amount of getting more intelligent would change this fundamental goal, just as humans getting more intelligent has not overridden our fundamental goals of "breathe, feed, have sex". It may have given us other goals as well, but those are very much still there.
Yes, that's pretty much the point.
And most likely, humans will have a pretty naive understanding of whatever motivates a superintelligent AI.
This is just a specific example of why I reject the orthogonality thesis. You change the context, you educate the agent, you change the goals of the agent. I do not agree that humans only chase “breathe feed sex” and while I do believe many stupid behaviors do come from evolutionary history, It’s plainly obvious that education, training, and genes play a role in self restraint and goal redirection.
And I agree that humans do not only chase "breath feed sex", as I explicitly said that in my comment. We have other goals as well. But those are very much still there.
"Better" according to what metric?
It may be the case that there is a tendency for high-intelligence humans to pick "more enlightened" goals. Perhaps there is a natural "enlightened goals" attractor for our species.
However I don't think we can extrapolate from that to a fundamentally alien AI.
I think even if this statistical tendency exists, it has clear counterexamples -- consider that 2 genius chess players may have opposite goals, of beating one another. And we shouldn't bet the future of humanity on this statistical tendency extrapolating outside of the original distribution of human species.
Here are some intuition pumps on how diverse goals can be even across intelligent species:
* Orcas killing sharks for their livers: https://www.livescience.com/2-orcas-slaughter-19-sharks-in-a... I don't believe dolphins show the same level of violence, even though both are smart cetaceans
* Intelligent dogs bred to be responsive and attentive to human needs -- unlike close cousins like the wolf
* Chimpanzees and bonobos are both related to humans, both highly intelligent, but with very different culture and goals https://www.youtube.com/watch?v=c6Ko0Hzi47U
As soon as you phrase goal selection in terms of metrics, you’re assuming that goal selection is based on some other goal - that is you’re already assuming the orthoganality thesis. Your logic is fully circular.
One thing that’s interesting to note about all of the examples you picked - every single one of those species shows cooperative behaviors. They share many other behaviors that are more similar than they are different. To reject the orthoganality thesis it’s sufficient to show that there is an empirical general association between intelligence and certain goals - then we can extrapolate an AI although of course it will function differently and may have many unusual behaviors will tend to follow those goals more directly. For instance intelligence is associated with : cooperation, empathy, inter and intra species communication, curiosity, etc. all of the species you mentioned exhibit these more than less intelligent species. Meanwhile something like “hunting to eat” is observed across the intelligence spectrum.
A safety mindset would suggest that rather than disregarding it until proven true, we should worry about it until proven false.
> meaning an intelligent being can pursue stupid goals.
It would pursue very intelligent instrumental goals, but the terminal goal is a free variable and I don't think there exists any measure by which terminal goals can be considered smart or stupid. It would be whatever is implied by its programming.
> States like that, it’s obviously wrong and laughable.
Perhaps not so obviously wrong nor so laughable as you think?
There’s no clear delineation in real entities between instrumental and terminal goals.
Yeah if you’re talking to a brainwashed religious fanatic, definitely it can seem counterintuitive to the doctrines they believe.
In healthy individuals, yes there absolutely is. Terminal goals are the ones you pursue for their own sake; instrumental goals are the ones you pursue as part of a plan to pursue a terminal goal, or another instrumental goal which connects to a terminal goal. Most people go to work in the morning not as a terminal goal, but as an instrumental goal; employment is in service of another instrumental goal of earning money; earning money is in service of a terminal goal of not starving to death. This is not exactly controversial stuff here. Some people do get so focused on an instrumental goal like "earning money" that they develop tunnel-vision and forget what terminal goal that money was originally in service of, but that's something most of them will eventually realize and then write a self-help book about.
Anyway, it takes intelligence to decide what your instrumental goals should be, such as whether there's perhaps a cleverer way to make money than by going to work for your boss each morning, but there's no way in which intelligence will help you choose your terminal goals. For the most part they aren't something you can even consciously choose.
In reality, people do not need to go to work to “not starve to death” as you say. There are a myriad of ways to survive without working a daily job.
Humans have to be socialized and trained to work a 9 to 5 job - there’s an entire education system structured to help create humans who view that as an acceptable goal.
No what you are saying may not be controversial in your little community but the AI panic is mostly isolated to a small community in a small corner of the USA.
As for there being "a myriad of ways to survive without working a daily job", congratulations! Your intelligence has allowed you to identify alternative instrumental goals that provide a path to your terminal goal; now you can rank them and choose the best option. You can also grow carrots in the garden or ask your neighbour if they have any, or ask your spouse to pick some up on the way home. Your intelligence will do the work and find a way. But your intelligence isn't what will guide you toward preferring carrot soup over parsnip soup, and preferring parsnip soup over fasting.
I'm not American and not in the USA, btw.
Think more carefully about the implications of multiple ways to survive here. Why do people pick one over the other? In a terminal/instrumental goal model, agents would pick the instrumental route that maximizes the return on the terminal goal. In reality we see that instead humans adopt habits, processes, and heuristics that guide them through daily life even when those do not lead to any specific goal.
So remind me of your original point? I believe you said it's "obviously wrong and laughable" that "an intelligent being can pursue stupid goals". Now here you are trying to convince me that humans are the ones who, like Pavlov's dogs, "pursue habits, processes, and heuristics that guide them through daily life even when those do not lead to any specific goal". Even when those habits involve repeatedly re-opening a fridge that you already know has no carrots in it, or salivating at a bell when you already know no food is coming.
So I'm confused how that proves your point about AGI. If I accept your view, it seems that if an AGI does merely no better than a human on this metric, I should anticipate all sorts of strange and irrational behaviour, including the pursuit of goals that would appear stupid, such as addiction to a reward channel. That does not seem to undermine the orthogonality thesis.
And the smarter the AGI gets, presumable the less it should lean on Pavlovian heuristics and the more it should make use of clear thought, which puts it more in my camp.
So that would apparently put the lower bound at "the AGI takes unexpected and irrational actions because it's not a rational agent and doesn't think coherently", and the upper bound at "the AGI takes unexpected and dangerous actions as rational steps toward an unaligned terminal goal".
I'm not sure where in this chain of thought it becomes laughably obvious that intelligence and goals are correlated, such that an AGI's increasing intelligence will tend it toward actions that we humans approve of, because anything else would be a "stupid goal"?
To me it's obviously wrong when stated that way because 'stupid' is not an appropriate metric for goals. We may consider goals good or bad within our value system, but that has little do to with 'stupid' or 'intelligent.' e.g. a body builder may have a goal to get as ripped as possible, a VC to make as much money as possible, an ascetic to deny the flesh as much as possible. Which of these are stupid or intelligent goals? I don't think that's a question that makes sense.
We may consider the goals of the Athenians (to expand their power) "worse" than the goals of the Melians (to be left alone)[0], but I don't see how they were "stupider."
[0]: https://en.wikipedia.org/wiki/Siege_of_Melos#The_Melian_Dial...
Let chaos take the world! /s