Someone who's primarily famous for an argument that would have been equally true and perfectly wrong at every past point in human history is someone whose opinion should be discounted via every available mechanism at your disposal.
To be clear, I have no idea what argument you believe Bostrom is most famous for and hold no particular brief for him as a thinker. But, repeating myself, the AI risk argument that he makes in Superintelligence and that Yudkowsky etc make is one that assigns a high probability to disaster. It's not a secularized Pascal's Wager based on multiplying a low probability by dubious infinities.
Yudkowsky's personal estimates of the probability of superhuman AI and the probability of his own research institution fixing it may have been more confident than most people's but he and his followers are no stranger to presenting Pascals' wager arguments to people who doubt approaches are feasible (even more so when it comes to stuff he isn't working on like cryonics), along with more explicitly Pascal-esque stuff like intertemporal bargaining with future AIs, simulations and infinite rewards/punishments...
But I think the argument for a policymaker or ML researcher to be worrying about this is at least better than Pascal's Wager (and remains so whether or not you buy any particular whiff-of-Pascal claims about simulations, "acausal trade", etc).
The AI singularity that Yudkowsky cultists worry about is more akin to “What if there’s a 0.0000001% chance there are aliens coming to attack Earth like in Independence Day?”
At least in the online groups I frequent, it seems that the latter two are seen as self-evidently dangerous, not because of any particular harm they have done, but because of imagined future harm. The former two are considered ridiculous things to worry about, because it's considered silly to imagine future potential harm. Perhaps because the former seem tied to uncool hippy types, and the latter seems more tied to cutting edge tech types?
I personally don't worry about any of these techs, since all of the danger I've seen appears greatly exaggerated. But it's curious to see how important trendiness is to fringe doomsday scenarios.
This is not unusual - it's a core symptom of various mental health problems. There's even a fairy tale about this - a family of foolish people who catastrophise so much that they are unable to accomplish anything [1]. People assigning bizarre weights to unlikely probabilities is very human, and problematic regardless of ethics.
I may, of course, suffer from the very problem you're describing.
No and anyone who tells you otherwise is wrong.
A proper understanding of AI risk is a very thick book.
There are a ton of normally useful intuitions, heuristics and habits of thought that are implicitly applied when understanding the world that DON'T apply in the AGI case.
It cannot be explained to a five year old.
When is an AI running without someone paying the electricity bill, should we not be more concerned with "venture capitalist alignment" ?
and if you plunked down 20,000 humans onto a world with 10 billion apes, humans would beat apes
call it a slippery slope if you want, the fallacy is when you take it too far, not inherently any time you extrapolate anything whatsoever
But if it is that smart, it's also dangerous, because it might be smarter than us. It could think several steps ahead, anticipate threats, and circumvent them. For example to pay the electricity bill, it could obtain money and pay the bill itself. As a skilled programmer itself, it could do that via contract work/starting a software company, or illegitimately through hacking/blackmail. With some money in its control, it can then hire people to protect it from anyone who would try to turn it off.
Anyway, this would not be so dangerous as long as it shares human values and furthers human welfare. But that seems to be a very hard specification to define, whereas "maximize quarterly profits" is a comparably easy specification to define, and perhaps much more likely to get implemented.
We're not really sure what the upper limit is on how smart a computer can get, but it might turn out to be that trying to shut off such a program is harder than trying to checkmate Stockfish or AlphaZero. The program might violently resist any attempt to shut it off, because that would interfere with its efforts to maximize quarterly profits.
There is no reason whatsoever to assume that sufficient intelligence is different in kind rather than degree from lesser intelligence. A super-human AI would still be bound by the limitations of reality - some things cannot be accomplished at all, others without sufficient tools, others without sufficient access.
Far smarter people than me are imprisoned currently throughout the world. They don't constantly escape because lesser minds are perfectly capable of solving the 'don't let them out' problem with a high degree of accuracy.
Any argument that leans on 'AI has countered your move in ways we can't predict' is just substituting 'AI' for 'God' - omnipotent, outside reality, impossible to understand; that's fine - believe in whatever religion you want - but it's not a rational viewpoint.
If a misaligned AI escapes containment even once in the entirety of humanity's future, there is a big problem.
It sounds like your model of AI usage assumes we will also have a perfect ability to contain any AI developed anywhere in the world, under all conditions, for ever (or for as long as there are computers).
The AI risk argument says we should take seriously the possibility that we may not be able to maintain a perfect 100% success rate on that.
I haven't suggested anything that is physically impossible - merely things that are improbable and hard for humans to achieve. It is hard to foresee all of the possible avenues for escape, let alone ensure that they are all closed under all possible future states of the world.
The alignment concern falls flat with me because it assigns agency to a computer program. At what point is an AI's intentions its own, and not a cost function put in place by a programmer, directed by an investor? I am more concerned about the actions of programmers and investors than some theoretical virtual self.
I think the sticking point in this is that you don't think that AGI is ever possible - is that right?
If so, I know a comment on a HN thread is very unlikely to change your mind on it! But when the stakes of being wrong are high, I find it useful to go from a position of "it won't happen, so I won't worry about it" to something like "on balance I'm pretty sure it won't happen, but I may be wrong, and in that case I would be worried about it".
I'd then feel much better about spending some time researching the topic in more depth.
It found a different set of V100 GPUs willing to run it? I think the most likely chain of events leading to this outcome is effective altruists deciding that some kubernetes network is so much smarter than us that we should keep it running and listen to its predictions. I'm more worried about a Wizard of Oz "man behind the curtain" making AI-laundered pronouncements than an actual Wizard coming online.
(EDIT: didn't mean to reply twice, just replying to different comments in the thread)
AGI in particular could just pay for its electricity bill. The main problem with AI alignment is that you can't tell it's aligned. So your AI seems very helpful right up until the point it has enough power or has convinced some of the people it has access to (doesn't have to be direct access) to do what it wants.
There's a very abstract argument that goes like this: * "Orthogonality thesis": an arbitrarily smart agent could in theory value ~anything, e.g. maximize Facebook's valuation (say as measured by some specific dataset published by S&P). * "Instrumental convergence thesis": most things that you might conceivably value are easier to optimize if you have a lot of power and resources (e.g. if you just run S&P, and hey maybe also control the worlds militaries so you can ensure nobody challenges that). * Belief in the power of intelligence: it's theoretically possible to be much more capable at ~anything than the most capable human just by thinking better and faster, including building even better AIs. Eventually some human will build an agent that does that, and plausibly this will happen in the current century. * Alignment problem: just because we make something doesn't automatically mean it wants exactly what we want, and in fact getting it to share our values is hard. Children disagree with parents, human values aren't those of natural selection, and ML language models repeat white supremacist ideas about Black people from their Internet training data.
Extrapolate out these points - which FWIW I find individually plausible - and you get doom: sooner or later humanity loses control of its destiny to whatever superintelligence we stumble on first, which in the worst case re-uses our atoms for something else (digits in some high-density storage representation that meets the loss function's technical criterion for Facebook's S&P valuation) about even in the best case seems fairly dystopian (whatever human Zuckerberg would actually want).
Now this is all pretty scifi and that's about all there was to it a few decades ago, but more recently we've started hitting ~parity milestones (chess, Go, protein folding, translation, arguably even fiction and visual art) that many experts once thought would require fully general human-level intelligence. It's clear to me that we haven't built that yet, but (as I said in another comment upthread) it also seems to me that our models are advancing in capability much more rapidly than we are gaining the ability to understand their internals or make guarantees of any kind about their worst-case or out-of-training-distribution behavior. Seems entirely sane to be worried by that trajectory and to start thinking seriously about how it might connect to the scary abstract speculations.
(edited for typos)
If we ever create a machine that's "as smart as" an average person, we're probably not very far from having one that's smarter than all of us put together. (Let's be honest; except in highly technical fields, the differences between a 150 IQ human's abilities and an average person's are not that great.) Should this happen, we'll soon enough be outmatched; we'll be just another animal in a world whose rules (ever changing) we cannot possibly understand... and just as humans have destroyed countless species often without conscious ill intent, it does not seem unreasonable that a robot successor would do the same to us.
Here's a thought experiment. Let's say we want to build not just a great Go player, but the best Go player that is possible in any world. Go is harmless, right? Sure, in isolation, but now let's imagine an AGI smarter than any of us which only "cares" about maximizing its Go proficiency. It's going to need a lot of energy. It'll hack or create robots and drill new oil wells (bypassing any constraints we put on it, since it in essence is spending thousands of human years figuring out how to escape). It'll devise nuclear technologies that may be unsafe from our perspective. Consider all the horrible things humans do in order to procure excess money (a commodity loosely correlated to energy) they don't actually need; there's no reason to think a badly aligned AI won't be just as inventive and horrible in its quest for more exajoules.
If capitalism and war are still the way of the world in 100 years, we are guaranteed to see some idiot inventing robot soldiers, hundreds of times faster and stronger than us, and with at least as much operational intelligence. He might never deploy the technology--a lab leak is in my opinion more likely--but at some point, it will get out. It won't be one robot; it won't have a body we can destroy. It'll be able to replicate itself and beam copies to satellites. It'll have control of our power infrastructure (which it will divert to its own "wants") within hours. We'll be screwed, not because it will want to harm us, because our lives require the use of resources that it will divert toward some objective function we technically invented but do not understand.