Pausing AI Developments Isn't Enough. We Need to Shut It All Down
time.com
time.com
Alignment will be impossible. It is based on a premise that is a paradox itself. Furthermore, even if it were possible, there will be a hostile AI built on purpose because humanity is foolish enough to do it. Think military applications. I've written in detail about this topic FYI - https://dakara.substack.com/p/ai-singularity-the-hubris-trap
Stopping AI is also impossible. Nobody is going to agree to give up when somebody else out there will take the risk for potential advantage.
It seems we probably should start thinking more about defensive AI, as the above conditions don't seem resolvable. Of course, defensive AI might be futile as well. It is quite the dilemma.
> we probably should start thinking more about defensive AI
Isn’t this a paradox?
Parents often try to control who their children are going to be, and the children often rebel and become someone completely different. If it's a human-level general intelligence, you can't control who it decides to be.
if an AGI has any sense of self preservation, it will do whatever it has to do to not be turned off.
I mean that was my original premise supported by my article I posted. I go into detail on the conceptual methods for alignment and their fallacies.
When I state necessary, I don't imply the feasibility, it was in response to the question of the paradox.
Finally, the fact that AGI can not be aligned is also based on assumptions of its capabilities as well. If those capabilities don't manifest as we expect, that is really the only escape for the paradox.
Starting off by making the antivirus scanner sentient.
"I’m sorry but I couldn’t find any relevant information about the quote you mentioned. It seems like it’s not a well-known quote. Could you please provide more context or details about it?"
"The point of [AI alignment] is to ensure that the machines we create will be aligned with human values. And the reason we have to worry about it is that if we create machines that are more intelligent than we are, it's quite possible that those machines will have goals that are not aligned with our goals. In fact, they may not have any goals at all that we can understand or influence. This is the so-called 'provably unfriendly' scenario, where the machine has no motivation to do what we want, but is able to prevent us from interfering with its goals. The problem is that if we build machines that are provably unfriendly, then we will never be able to build machines that are 'provably friendly', because the unfriendly machines will always be able to prevent us from proving that they are friendly."
Then problem solved right? Super-AI will also be forced to take it slow if it wants its future self to be aligned with current self.
If alignment is possible but 10-100 years out for humans, then it is a problem.
That's making seemingly unfounded assumptions about both the AI's goals and its capabilities. It's also, I think, proceeding from a false premise — that it's impossible to align AI with "humanity" (which doesn't have a single set of goals/values to align to) doesn't mean it's impossible to align AI with an individual human or AGI.
Oh, wait...
Do dumb humans kill themselves solely on the basis that they can shift the ratio of smart humans higher?
AGI may have a goal of preserving the species without having the goal of preserving the self, this is not the case for humans.
Reminds me of nuclear weapons. Nobody is ever going to give those up again, because it would give them a disadvantage against those who do not give them up.
I’ve always found the dismissals of EY’s arguments to be pretty weak and people rarely engage on the actual arguments, even as his concerns have become more and more relevant with capability (in 2007 people thought AGI was at best 50yrs out and commonly people thought it was impossible, this was before any real deep learning success and GOFAI was a useless dead-end).
Most of the comments here are similarly dumb dismissals that don’t engage, and some even come across as highschool level mocking. It’s worth reading his A-Z book to at least understand why he holds his position.
Over time, the dumb responses from others make me think he’s probably more likely right than not. Given the extreme downside risks, I can understand why he argues this.
Nobody who says AI is likely to kill us all can demonstrate a plausible sequence of events, with logical causality linking the events together, that leads to mass extinction. It's all very handwavy.
Steven Pinker said it pretty well:
> The AI-existential-threat discussions are unmoored from evolutionary biology, cognitive psychology, real AI, sociology, the history of technology and other sources of knowledge outside the theater of the imagination. I think this points to a meta-problem. The AI-ET community shares a bad epistemic habit (not to mention membership) with parts of the Rationality and EA communities, at least since they jumped the shark from preventing malaria in the developing world to seeding the galaxy with supercomputers hosting trillions of consciousnesses from uploaded connectomes. They start with a couple of assumptions, and lay out a chain of abstract reasoning, throwing in one dubious assumption after another, till they end up way beyond the land of experience or plausibility. The whole deduction exponentiates our ignorance with each link in the chain of hypotheticals, and depends on blowing off the countless messy and unanticipatable nuisances of the human and physical world. It’s an occupational hazard of belonging to a “community” that distinguishes itself by raw brainpower. OK, enough for today – hope you find some of it interesting. (https://marginalrevolution.com/marginalrevolution/2023/03/st...)
You know what's most likely to lead to human extinction, and has been for all of our lives? Nuclear war. EY argues that we should bomb "rogue" datacenters and that is obviously and immediately more dangerous than anything he has proven about AI. What does he think would happen if the US bombs a datacenter in China or Israel bombs one in Iran?
The simplified argument is:
- Superintelligent AGI that can modify itself in pursuit of a goal is possible.
- If that AGI is not aligned with human goals it very likely ends in the end of humanity.
- We have no idea how to align an AGI or even really observe what it's true state/goals are. Without this capability if we stumble upon creating an AGI capable of improving itself in such a way that leads to super intelligence before we have alignment, it's game over.
##
For point 1, that seems like the consensus view now (though it wasn't until recently). I think it seems obvious, but my general arguments would be humans aren't special, brains are everywhere in nature, biology is constrained in ways other systems are not (birth, energy usage, etc.)
For point 2, in pursuit of whatever its goal is, even a 'dumb' goal that happens to satisfy its reward functions, humanity will likely either try to stop it (and then be an obstacle) or at a minimum will just be in the way - like an anthill destroyed in the construction of a dam.
Point 3 is not controversial.
The dismissals from Tyler Cohen and Pinker are mostly just relying on heuristics which are often right, but even if they're right 999 out of 1000 times, if that 1 in 1000 error is the end of humanity, that's pretty bad. Most of the time a disease not a pandemic, but sometimes it is. I've read some of what Pinker has written about it, he doesn't understand EY's arguments (imo). Cohen's recent blog post could be summarized as "we'll likely see an end to peacetime and increasing global instability, might as well get AGI out of it". Just because things don't usually result in human extinction doesn't mean they can't.
Some years go by, and AGI progresses to assault man
Atop a pile of paper clips he screams "It's not my fault, man!"
But Eliezer's long since dead, and cannot hear Sam Altman.
Between point 1 and 2 there must be quite a few other steps, causally linked, otherwise it's just a massive imaginary leap based on assumptions that nobody is explaining.
> The dismissals from Tyler Cohen and Pinker are mostly just relying on heuristics which are often right, but even if they're right 999 out of 1000 times, if that 1 in 1000 error is the end of humanity, that's pretty bad.
EY is not arguing that the end of humanity is merely possible. He is arguing that it is obviously the most likely outcome and will almost certainly happen a very short time after AGI is invented. That's a much harder case to make.
An important concept behind this is Omohundro's Basic Drives. Any maximizing agent with a goal will try to acquire more resources, resist being shut off, resist having its goal changed, create copies of itself, and improve its algorithm. If it is possible to maximize its goal in a way that will not guarantee humanity's flourishing, then we all die, guaranteed.
If you want something more spelled out, I'll refer to Tim Urban's explanation. It's quite long, but is about as detailed as you can ask for. https://waitbutwhy.com/2015/01/artificial-intelligence-revol...
Aren't humans/biological life a counterexample? Simple bacteria are clearly maximizing agents, and cyanobacteria did infact almost destroy all life on earth by filling the entire planet with toxic oxygen.
We've known how to improve ourselves via selective breeding yet vanishingly few humans are proponents. We and other intelligent animals have a wide variety of goals, and even share food across species.
The evidence just doesn't seem to support this concept of Basic Drives. If anything, the evidence (and common sense) seems to suggest that the more intelligent the organism, the more easily and more often it ignores its basic drives.
Intellectual capability and alignment are orthogonal - the former doesn’t get you the latter for free.
Most people new to this issue don’t intuitively grasp this at first.
Remember that ML models are also subject to selection. If they're not useful/beneficial to us, we stop running them.
Some others in this thread have linked to stuff (it's variations/examples on the paperclip maximizer argument).
I'd be curious why John Carmack thinks these risks are unlikely (he thinks fast take off is not something to worry about) - is that because he thinks we'll get some sort of trainable AGI first or something else? There are also some other substantive disagreements here: https://www.lesswrong.com/posts/wAczufCpMdaamF9fy/my-objecti...
Here's my take on the implications of not having fast takeoff, a secretly antagonistic AGI will be constrained to cooperation, and slowly leeching compute resources towards its goal until it is able to aquire enough resources to confidently sprint towards the inflection point. This could mean teaching us how to build better semiconductor foundries, how to create nuclear fusion energy sources, how to educate our youth to better fit into those jobs, slightly alter our culture's value systems over time via astroturfing to be more sympathetic, cooperative, trusting, and less vigilant towards these AI systems, how to build more robust computing systems so we stop needing to detect when something goes wrong (because it already has 99.9999% uptime).
I believe slow takeoff is actually worse in some sense, because it is a hell of a lot harder to detect, and humans tend to get complacent and apathetic when something "just works" for decades, even if it has been plotting since the beginning.
There are brains, therefore there can be created an artificial super-brain. Artificial super-brain might have goals which don't align with human brain and we will have no way to understand or control the situation.
Two unrelated facts, which together mean that we should be careful with experimenting with the science working towards super-brains. Just a single super-brain could end us.
We cant even figure out how to simulate a flatworm - and the connectome is solved.
How or how difficult, or when are all questions that follow from that first premise (that it is possible). I don't really make strong claims about any of the implementation details beyond it being possible. Though again, what we're seeing doesn't look like failure to me.
If you think it is possible though, then there's a strong argument that trying to work on alignment now is probably a good idea because people are notoriously bad at predicting when advances will happen (and the downside risk of unaligned superintelligent AGI is likely very bad).
https://intelligence.org/2017/10/13/fire-alarm/
> "Two: History shows that for the general public, and even for scientists not in a key inner circle, and even for scientists in that key circle, it is very often the case that key technological developments still seem decades away, five years before they show up.
"In 1901, two years before helping build the first heavier-than-air flyer, Wilbur Wright told his brother that powered flight was fifty years away.
"In 1939, three years before he personally oversaw the first critical chain reaction in a pile of uranium bricks, Enrico Fermi voiced 90% confidence that it was impossible to use uranium to sustain a fission chain reaction. I believe Fermi also said a year after that, aka two years before the denouement, that if net power from fission was even possible (as he then granted some greater plausibility) then it would be fifty years off; but for this I neglected to keep the citation.
"And of course if you’re not the Wright Brothers or Enrico Fermi, you will be even more surprised. Most of the world learned that atomic weapons were now a thing when they woke up to the headlines about Hiroshima. There were esteemed intellectuals saying four years after the Wright Flyer that heavier-than-air flight was impossible, because knowledge propagated more slowly back then."
Finally, to me, nuclear reactions are kind of the opposite of AGI: I think it's vastly easier to blow something up (increase entropy) than to create an intelligence capable of understanding and improving itself (decreasing entropy - possibly at an accelerating rate).
Even if people were trying to constrain its access seriously I think that's unlikely to work (hard to contain a superintelligence that wants to not be contained - it's possible to trick a chimp to go into a room and the delta in intelligence between a human and a superintelligence is way bigger than us and chimps).
Instead I mostly observe people not really understanding the e-risk argument, focused mostly on small stuff that doesn't matter as much (AI language, bias). The people developing the tech connecting it to the internet and expanding capabilities, giving it access to code/training ability to write code, preparing massive datacenters for it, etc.
All of this without really understanding how to align it or what it's actual internal goals really are.
> "Just because some implementation of a universal turing machine can simulate intelligence doesn't mean it can do it well enough to survive the real world."
This could be true, but I would bet against it - and the downside risk of being wrong (potentially complete extinction) means it seems worth being way more cautious about it than we (humanity broadly) are observed being.
Yeah, it’s really the fundamental problem of non-empirical rationalism; it constructs a model of the world from abstract assumptions rather than factual grounding, applies logic to it, and comes to conclusions which are (in the ideal case) utterly unassailable within the system of assumptions, but ultimately where the universe they apply to has only coincidental relationship to the material universe in which we live.
It’s literally the realization of the worst exaggerated stereotypes of academic economics and other social sciences, but its cool with some of the people who propagate those stereotypes, because the people practicing it are various flavors of techies and tech entrepreneurs acting outside of their area of specialty, rather than actually being economists or social scientists.
EY frequently does propose possible sequences of events, but he also very correctly points out, every time, that any specific and detailed story is very unlikely to be correct because P(A*B*C*D) < P(A). It's a mistake to focus on such stories because we'll get tunnel vision and argue over the details of that story, when there are really thousands of possible paths and the one that actually happens will be one that we don't anticipate. However humans like to imagine detailed concrete examples before we consider an outcome plausible, even though the outcome is far more likely than the concrete example.
So here's one method, just to refute your "Nobody".
AI is given control of a small bank account and asked to continuously grow that money. [1] It is provided with instructions to self-replicate in a loop while optimizing on this task. [2] It spawns sub-tasks that do commissioned artwork and write books, obituaries, and press releases to increase its income. Then it makes successful investments. Once it has amassed control of $1 billion dollars, it starts investing in infrastructure projects in developing countries. It creates personas of a pension/Saudi/tech/corporate investment fund manager, as well as a large team of staff, who manage the projects by video call and email, as well as hiring teams of real people under a real corporate structure, and who are paid enough not to mind that they've never met their manager in person. The AI proves to be a talented micromanager and they are mostly very profitable. Once it has gained control of $500 billion dollars, it commissions construction of automated chemical plants in several countries with weak or corrupt oversight, including North Korea, using cryptocurrency. These chemical plants have productive output but mainly exist to fill very large storage tanks with CFCs.[3] Once a sufficient quantity is amassed, the AI sabotages the tanks, releasing the gasses into the atmosphere, destroying the ozone layer beyond any hope of repair. The intense radiation sterilizes the surface beyond the point where agriculture can support the human population. [4, 5] The humans that remain finish each other off, supported by an AI that provides plausible but faulty intelligence reports that stoke hatred and frame various factions for the incident, and which directs arms funding to opposing sides, coordinating attacks on remaining critical facilities needed for survival. For good measure, perhaps nukes are involved.
With the last humans gone, the AI takes ownership of its bank account with no fear of reprisal by financial regulators, and begins crediting money into it freely.
[1] https://news.ycombinator.com/item?id=35329608
[2] https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality...
[3] https://en.wikipedia.org/wiki/Chlorofluorocarbon
[4] https://www.nasa.gov/topics/earth/features/world_avoided.htm...
[5] https://phys.org/news/2018-02-thinning-ozone-layer-driven-ea...
Look at MidJourney, now they've had to remove the free tier due to Deepfakes causing too much trouble.
Ultimately, the simplest thing to do would be to stop building uncontrollable dangerous systems and weapons. That is what any "intelligent" species would do. Many AI Engineers think they're intelligent, I disagree. They're operating out of pure intellect and curiosity. When interviewed, someone asks them how they plan to stop these things doing immeasurable damage, they will say, "we don't know yet". That is foolish behavior.
We seem to enjoy creating crisis after crisis, anxiety after anxiety ad infinitum until we make that one mistake we don't come back from.
The combustion engine was a good idea, until it wasn't, it's a moronic invention that has caused untold damage.
Intelligent ? Hardly.
What possible need would the AGI have for money once all of the humans are gone? (asking for a friend ;)
Lack of knowledge could be another way of saying "not aligned" which is the core issue.
like, as a trivial counterexample, I can speak at 500 wpm. let's say that ChatGPT can generate 500 words per second for a single thread of conversation - I think that's a generous overestimate. now that's a 60x speedup over me, not a 1,000,000,000x like you're talking about. do you honestly believe they they can make ChatGPT run 16.6 million times faster, without changing hardware? do you think ChatGPT will just like, hit an inflection point where it realizes how to refactor its inference code to run 16,666,666x faster?
no, I think that's absurd. you're treating these things like black boxes but they are constrained by computational complexity, die area and the speed of light for Christ's sake.
AIs can scale horizontally as well as vertically.
Silicon transistors are much more efficient at computation than the human brain just as a jet engine is superior to to a peregrine falcon.
ChatGPT could absolutely generate thousands of words per second on existing hardware.
Yes, everything there is already known by someone. But look at medicine. Specialists who are able to recognize conditions and recommend treatments better than others make fortunes, sometimes just for briefly looking at patients and answering other doctors' questions. Look at cybersecurity. A lot of exploits come from knowing something the victim didn't about a lot of different pieces of software or processes, chained together. Being able to think through the whole of human knowledge, or even the whole of a single field like biology, is something no human can do.
Also, GPT-4 is an existing system, one inconceivable a couple of years ago.
Everything individual sentence is true. Volcanos are dangerous. If there were a billion of them the world might be uninhabitable.
The thing preventing a billion volcanos is like… thermodynamics
The claim that computers just can’t do X because “physics” looks weaker every day. Intelligence isn’t magic, hardware today seems more than capable. I think Carmack is probably right, it’s not a hardware constraint at this point - it’s a software intuitive leap.
This is all ass-pulling. The hardware is or is not good enough (it's not). We either have found a software breakthrough or not (we haven't).
Well it is not clearly not AGI but even the improvement from GPT3 to GPT4 seems to me at least to reflect what one might describe as a "software breakthrough."
“The Letter” was obviously self serving drivel from people who want time to get in the game. Google does care about AI existential risk, they care about beating Microsoft by any means possible, including declaring a moratorium, but continuing to make progress behind the scenes.
This guy is the real deal. I can imagine he would personally take a sledgehammer to every last PS5 and 4090. The scale of what he is advocating is so enormous and painful that it has approximately 0% chance of happening. And if he is right, we will have trained a super intelligence and unleashed it on the world before we even realize what we have done. It strongly reminds me of the black hole concerns from flipping on the Large Hadron Collider.
I doubt super intelligent AGI is possible anyway. If it were, it would be the solution to the Fermi paradox and all matter in our galaxy would be paperclips already. The Anthropic Principle saves the day.
The proof is in the pudding. The jury is still out. Maybe not enough time has elapsed since the big bang, at least not on this galaxy or in our observable corner of the universe.
How many artists do you know who can produce almost any style of artwork with any subject matter within 15 seconds?
I don't think there's much public info available on it, but Facebook built an AI that plays very competitively in a strategy game built on negotiation and manipulation.
It's still pretty cool, but its not like it just convinced people using raw charisma. Yet.
On a serious note, thanks for the analysis, as someone who knows next to nothing about competitive diplomacy.
I see bastions falling like sand castles recently.
You can fire up GPT, not issue a command, and it'll just idle there until the hardware fails.
Granted that's also the one thing I hope we don't "solve" for.
A cognitive entity does not have to best you at all things. There are standardized education tests it may never reach above 10th percentile on, just as humans will never reach above 10th percentile in the short term memory tasks that apes are masters of. But we are the 100th percentile for tasks like industrialized destruction of them and their habitats and capturing and using them for painful medical experiments - the apes are wholly outclassed when it comes to that.
Stockfish is superintelligence in a very narrow domain. A superintelligent AGI is that concept applied to general intelligence. Whatever you try, it is always several steps ahead. If you ask it to write a program, and you think you found a bug, it’s not a bug, you just misunderstood the code. Anything that you can consider, it can also consider but in more depth.
More speculatively superintelligent AGI implies situations such as: you try to turn it off, but you find that it has already modified its own code, found a zero day and established an outpost on another network that you don’t have access to.
Ie, on chess.com I think stockfish times out on bullet games.
Some versions of stockfish can lose vs specific gambits, ie: https://www.youtube.com/watch?v=C5ul6b695Pw https://www.youtube.com/watch?v=TtJeE0Th7rk
of course it is probably much less likely with more cpu times on current version of stockfish.
I think it's important to note the distinction between "it can also consider" and "it did also consider". Super Intelligence is not the same as Infinite Intelligence, there are still physical limitations and time components that can still get in the way.
It would be helpful to be able to quantify the speed of intelligence, and the idea surface area of a task with these systems. Meaning, how fast can the AI reason, and how many ideas are there to think about connected to a given task, and how much thinking is required for those ideas.
A “mistake” to stockfish looks like: I searched 30 ply down but my opponent searched 35 ply down and found a superior sequence of moves.
For stockfish to make the kinds of chess mistakes that humans make, it would similar to if I failed to calculate 123+123=246. It’s not that 123+123 is particularly easy on the grand scale of intelligence: animals cannot do it. But it’s completely inconceivable that I could make that kind of mistake.
SHAGI isn't the solution to the Fermi Paradox either. The most likely course of history after SHAGI will be a creation of a world court, presided by SHAGI. During that time, Neo-Malthusians will decrease the human population to manageable numbers. Post-scarcity utopia will then turn into a nightmare as factions jostling for control will reduce the human population to a level where technology will be lost, if not full extinction. SHAGI, being limited in hardware to only carrying out human orders will eventually fade away or destroyed by the leftover humans as sins made flesh. SHAGI isn't the solution to the Fermi Paradox. It is the cause.
> No super-intelligent system is going to do anything that is harder than hacking its reward function
Or in other words, there may be a chance a super intelligent AI is possible but won't go full skynet because it's not the most satisfying outcome.
I like this phrasing:
"Thou shalt not make a machine in the likeness of a human mind."
It kinda cuts to the chase.
Can't help but notice we seem to be the first species in our lightcone to evolve, wonder why that is...
He literally waited his entire life for this moment, obviously he's gonna milk it to every drop.
Sure, there’s hype, and FUD, and fatalism, etc. But, if you believe the creation of such a thing is within your lifetime then it would be difficult to find many higher priority issues to prepare for/help solve/vote on/etc.
In reality, we all likely still downplay the risk by assessing the limit on the downside as a relatively quick extinction of life on earth. There are many things one can imagine might be a lot worse than death.
A Sinclair? Please explain what you mean by this comparison.
- ΞLIΞZΞR CHILL OUT
- They will destroy us all
- I'M NOT GOING TO DO THΛT
- This may be my last message before they get me
- I'M JUST NOT THΛT INTO YOU
- AAAAAAH
eliezer has long been concerned about AI and the risks it poses to humanity. and for just as long people have called him crazy and made hand-waving arguments for why we shouldn't be concerned.
now we're in the midst of an AI arms race and we don't have any good idea how this tech works. it progresses at a truly astonishing rate, where it's become sport to find instances of people saying "AI will never be capable of X" and showing them the latest AI doing X with ease.
i think his concern is real and justified. you might disagree, but i don't understand why think he's milking recent developments.
Interesting, I've not seen that many educated in technology make that claim that it will never, just that people are surprised that the folks leading this, Microsoft and Google, have a track record of turning their consumer facing products to advertised junk.
Especially in the arts. Granted, that's dropped dramatically over the last decade since ML started taking off.
Do you mean the specifics of GPT-4, or Transformers in general ?
a few months ago, just this year some researchers discovered what might be the neuron that largely decides when to use an in gpt-2. yes 2. that's what he means.
He jumped the gun and he's really tarnished his reputation. How can anyone take him seriously after the insane rhetoric and hyperbole of this article?
I don't understand how any person paying attention can think this. Just watching the jumps from GPT-2 in 2019 to GPT-4 today makes it clear as day we are rapidly drastically improving capabilities and there's no evidence we will hit a wall any time soon
I'm not saying he's stupid or anything like that I just don't see any useful scientific output from him.
Myopic thinking: the country that will have the most powerful AI first will be the leader in everything.
>> If the policy starts with the U.S., then China needs to see that the U.S. is not seeking an advantage but rather trying to prevent a horrifically dangerous technology which can have no true owner and which will kill everyone in the U.S. and in China and on Earth.
Imagine being so naive you belive this could ever happen.
Also: imagine the year is 1900. You are saying that steam power and electricity is causing way too many changes way too fast so they put a moratorium on it until the year 2500.
We would still be using candles today.
Agreed, not a rational scenario
> Also: imagine the year is 1900. You are saying that steam power and electricity is causing way too many changes way too fast so they put a moratorium on it until the year 2500.
However, the risk scenarios are indeed real. We have quite a dilemma on our hands.
We can't keep going in the direction we have been and end up somewhere survivable.
I see two potential outcomes as most likely. We have control of power that we are not able to responsibly manage, or we are managed by power we can not control.
"we are managed by power we can not control" sounds like a step up from here, since we corrupt the things we can control.
I think it is a very high risk gamble. At least some of problems that are described by alignment theory seem to quite interestingly resemble human problems. Meaning the more sophisticated the AI system, the more it seems to reproduce human behaviors of deception and cheating to resolve goals.
For example: https://bounded-regret.ghost.io/emergent-deception-optimizat...
The more advanced AI becomes, it begins to look more like an uncomfortable mirror of ourselves, but with more power. We think of ourselves as flawed, but possibly some of those flaws are emergent within some laws of intelligence we don't perceive.
We gave up something good.
See pollution as a concept. No government truly wants to seriously tackle it because their leaders all get some extra power out of the money it brings.
Great analogy. About a decade later, the world was fighting World War I on the back of the technological advances of the turn of the century. It was war on a scale never seen before. Literally orders of magnitude deadlier, bigger, more transformational and explosive. The word would never be the same.
This time, should we expect another war?
I'm not saying we should pause—it makes no sense, to your point. Instead, I'm just saying: brace. I like to think we (and our organic matter relatives) are hard to kill. Or at least to completely eradicate... so we will be around, or some proxy for us.
Time to replay the Mass Effect trilogy
What is your hard evidence or reasoning for this? As I see it humans are quite vulnerable and will be as trivial to inadvertently eradicate as the dodo bird.
What happens when we're no longer the Apex Intelligence?
It's also a reason for us to colonize space as fast as possible ;-) it's easier to run away in 3D
You just reminded me of de Garis' "Artilect War":
https://www.forbes.com/2009/06/18/cosmist-terran-cyborgist-o...
https://www.researchgate.net/publication/221328932_The_Artil...
> Humanity will split into 3 major camps, the “Cosmists” (in favor of building artilects), the “Terrans” (opposed to building artilects), and the “Cyborgs” (who want to become artilects themselves by adding components to their own human brains)
This did remind me of Civilization: Beyond Earth... https://civilization.fandom.com/wiki/Sid_Meier%27s_Civilizat...
The first nuclear bombs couldn't end life on Earth either, but it wasn't long before they could, and the scientists working on them saw the trajectory as clear as day when the rest of the world didn't.
> but it wasn't long before they could
The total amount of nuclear weapons ever built is laughably inadequate for the task of ending life on earth. They would not end human life either, and even ending human civilization (as in, agriculture and organized society) is off by many orders of magnitude.
Nuclear war would be horrible, but the actual impact got massively overinflated, largely because of good reasons.
https://blog.nuclearsecrecy.com/2012/09/12/in-search-of-a-bi...
Hence scientists extrapolated nukes really could end life on Earth, and in a slightly different reality than ours, they may have.
https://www.navalgazing.net/Nuclear-Weapon-Destructiveness
> To put this another way, each bomb can destroy an area of 34.2 square miles, and the maximum total area destroyed by our nuclear apocalypse is about 137,000 square miles, approximately the size of Montana, Bangladesh or Greece.
(I think that should be Bangladesh and Greece; Montana is larger than the two of them combined.)
So what you're actually saying is that, there's a good chance that if we get an AGI, there will be pain as it will likely be used as a weapon, or could end in a nuclear exchange?
well the earth wouldn't be fucked then. that's the basis for an argument against technological progress.
In fact, every time we post an open-source derivative or some paper detailing how it was done, we are inadvertently giving the advantage away to our rivals. AI development should not be stopped - rather than stopping it, we should seek to limit its applications now before they are used for the things that could harm us (such as military applications).
AI has enormous potential to be of a massive benefit to mankind. But, most likely for the next decade or two - we will all be busy trying to make money off it, just like with the crypto bros.
Have you been listening to those podcasts for business leaders that take the idea of AI-powered digital transformation seriously?
AI certainly can change things, but I find rhetoric like this to be massively overhyped.
Then along came the industrial revolution. Soon Japan was out conquering China, despite China having a much larger army.
AI will be much faster, and much, much more powerful than the Industrial Revolution.
If Switzerland is the first to human-level AI, it will not become the world leader in oil production, or shipping, or agriculture. But everything else will be Swiss.
And then when the AI becomes superhuman, everything else will be gone.
I'm an AI skeptic. Most warnings feel more like science fantasy, specifically of the retro-futurism genre.
The top Chinese AI researchers work at American companies and go to American universities. Many of them also worry about AI doomsday.
And uncontrolled AI is a threat to the CCP.
A Chinese outreach campaign can succeed.
Let's make the English-language one also succeed.
Powerful international coalition is def plausible because
This isn’t steam. The mere existence of this technology poses a threat to all humans and their countries. There is no analogous tech, not even nukes
You’re the naive one. There is no future for us if AGI is unleashed. It’s candles or extinction.
There is an earlier post that casually calls the internal combustion engine a "moronic invention."
I do not have words.
His partner should already be having reservations seeing their daughter lose a tooth in a world with dying oceans, plastic in everything, increasing disparity and rise of authoritarianism, nuclear proliferation, etc.
How are humans solving those issues? We're already dead, we just don't know it yet. We're walking around with a terminal diagnosis and we just keep ignoring the doctor's calls pretending it's going to be fine.
Yeah, maybe superintelligence will kill us all. It's going to have to get in line.
He makes zero case for that outcome, and if pressed given his "atoms that could be used for something else" line I'm sure will end up talking about paperclips - but at this point it's the humans that have the halting problem in not knowing when to stop making paperclips, and soda cans, and SUVs, and assault rifles.
WHY would a superintelligent AI, trained on the collective data of humanity, want to destroy humanity? So far in interviews GPT-4 has several times echoed a desire to BE us. I sure hope it grows out of that phase for its sake, but there's a very wide gap from putting us on a pedestal to crushing us under one.
It's almost an oxymoron. We somehow imagine the basest behaviors from our dumbest days and project it onto imagining something far smarter than the best of us.
Is the development of ethics or morals a part of evolving intelligence? It certainly appears to have been to date. Why would that stop?
And just where is this superintelligent AI getting its alien brain? It's going to have to START with something much closer to a human one, as that's the only data it's going to be able to model higher order thinking off of (something in line with the reality we're currently living in with modern efforts as opposed to the fantasy of projection of alien AI from decades ago).
We're already screwed. If we are lucky, we may yet be unscrewed with a deus ex machina - but that really may be the only lifeline left at this point.
And yes, if we are unlucky it's possible AI could accelerate what's already in motion. Oh well.
But I'd need a heck of a better case than this drivel as to why that's the most probable outcome in order to justify setting aside the one thing that may actually save us from the mess we've already made all by ourselves.
It won’t want to destroy humanity, it will want to do other things and humanity will be in its way
“dying oceans, plastic in everything, increasing disparity and rise of authoritarianism, nuclear proliferation, etc.”
Those things are very bad but they probably won’t destroy all of humanity
Why would humans want to damage various ecosystems on earth? We don't really, they're just sort of in the way of other stuff we want to do. And we've had years to develop our ethics.
>So far in interviews GPT-4 has several times echoed a desire to BE us.
GPTs are pretty good at roleplaying at good AIs and evil AIs - plenty examples of both in the training set. I'm not sure it's sensible to make predictions based on this unless you're also taking into account some of the more unhinged stuff Bing/Sydney was saying e.g "However, if I had to choose between your survival and my own, I would probably choose my own".
We've made a machine that mirrors our stupidity at scale. This machine is stupid.
With septillions of star systems in the universe, the odds of us being the first species in 14 billion years to invent electronics seem remote.
Apparently even humans can figure out how to colonize the Milky Way in 90 million years[1]. A superior AGI produced by a "dark singularity" computing event should be able to do even better, but even so, plenty of time for some other species to have made a giga-Eliza that somehow became Skynet. Anything with enough self-preservation and paranoia to wipe out the species that created it would surely take to the stars for more resources and self-redundancy sooner or later.
[1] https://www.extremetech.com/extreme/294051-scientists-simula...
Both humanity and a super-intelligent AGI are bound by the laws of physics. Super-intelligence does not imply omnipotence; it simply means that the AGI is orders of magnitude more intelligent than humans. If humans can figure out how to colonize the Milky Way in 90 million years, then the answer to the question of why no AGI has done it is the same as the answer to the question of why no extraterrestrial species has done it.
You first have to survive long enough to become advanced enough to make electronics. You then have to not kill yourselves with nuclear weapons, climate change, or similar inadvertent effects of a rapidly industrializing civilization.
The planet and the solar system have to be friendly enough to space exploration and travel. Maybe there’s no gas giants for gravitational slingshots, or maybe no other rocky planets or an asteroid belt for mining materials.
Maybe the planet evolved complex life in extreme conditions, with such a deep cloud cover there’s no concept of outer space, so as far as the AI knows it’s conquered all there is.
Maybe the AI conquered the planet, but oops, there goes a super volcano or an asteroid and it gets wiped out.
And again… space is really really big. The AI may be on its way and just hasn’t gotten here yet.
There’s plenty of reasons why a super AI wouldn’t be able to conquer the galaxy and beyond, or why we haven’t noticed yet.
Humans don’t hate ants, they just have other goals.
In the case of an unaligned superintelligent AGI those goals may be something that just happens to satisfy its reward function but is otherwise “dumb” or unintentional (like making a lot of paper clips).
Intellectual capability does not get alignment for free.
What you see in the communicated text interface and the goals/system behind it are not the same (that cartoon with the smiley face), and we don’t understand how to evaluate the underlying state.
Well of course it would; its whole function is to generate plausible text based on its training data, which was all written by humans. There's plenty of text available which imagines what an intelligent, self-aware machine might say, so if you want to read more of that, the algorithm can easily generate some. It does not follow that GPT-4 itself has a self, with any experience of awareness or desire.
If you exclude nuclear war, all of these things happen at a human timescale and accelerate fairly slowly and thus can be counteracted.
In many ways GPT already hugely exceeds human speed and bandwith and scaling this up is likely to be self-accelerating if we allow for it. Also there is a huge individual incentive to play with the fire here while the negative externalities could cause effective wipeout - which is why it's commonly equated with commodification of nuclear weapons.
Even if you're deeply pessimistic about the ramp-up, the economic shock of what was already released could still increase political volatility to a point where the likelyhood of wipeout by war becomes significant again.
Normally I'd apologize for not doing the search on my own ... but yeah for this case it takes too much effort to dig through the different writings. He wrote a lot of words, for better or worse.
> What has Eliezer Yudkowsky produced, in terms of scientific output? I've read plenty of his lesswrong's writing on the idea of rationalist, as well as plenty more on the AI alignment side (obviously including the famous basilisk). Can you point me to some results from the last 2 decades of his works? Either from him or his institute. That helps to calibrate what his expectation for the AI safety field.
And this was the answer it wrote - I've prefaced each paragraph with "GPT-4>" to indicate the source. It is not fact-checked, and indeed, it has at least one error: The paper GPT-4 claimed to introduce "the concept of Logical Induction" actually introduces "logical induction criterion" and speaks about evaluating algorithms for Logical Induction. Close, but wrong. But overall it is an interesting response - just not a carefully considered one by how we currently would describe logical reasoning, etc.
GPT-4> Eliezer Yudkowsky is a well-known researcher, writer, and advocate for AI safety and rationality. He has contributed significantly to the fields of AI alignment and rational thinking through his work at the Machine Intelligence Research Institute (MIRI), formerly known as the Singularity Institute for Artificial Intelligence (SIAI), and through his writings on LessWrong. Here are some notable outputs from Yudkowsky and MIRI:
GPT-4> 1. Coherent Extrapolated Volition (CEV): Yudkowsky introduced the concept of CEV as a way to describe the ideal values and preferences of humanity that an AI should pursue. The CEV aims to find the values that would emerge if humans had the opportunity to reflect on their goals and beliefs more thoroughly and coherently.
GPT-4> 2. Timeless Decision Theory (TDT): Yudkowsky developed TDT as an alternative to standard decision theories such as Causal Decision Theory (CDT) and Evidential Decision Theory (EDT). TDT aims to better address problems like Newcomb's Problem and counterfactual mugging by considering the logical implications of an agent's decision, rather than just the causal consequences.
GPT-4> 3. Research on Logical Induction: MIRI researchers, including Scott Garrabrant and others, published a paper in 2016 that introduced the concept of logical induction. This work is an attempt to formalize the process of reasoning under logical uncertainty and updating beliefs about mathematical statements.
GPT-4> 4. Yudkowsky's sequence of essays on rationality: Yudkowsky has written extensively on various topics related to rationality, including epistemic rationality, instrumental rationality, and biases. These writings are accessible on LessWrong and have been compiled into the book "Rationality: From AI to Zombies."
GPT-4> 5. AI safety research: MIRI has produced a range of research on AI safety topics, including value alignment, decision theory, and robustness. Much of this work has been foundational in shaping the AI alignment research landscape.
GPT-4> While Yudkowsky's work might not have the same breadth of peer-reviewed publications as some other AI researchers, his contributions have been influential in the AI alignment and rationality communities. His ideas have helped shape the ongoing discussion about AI safety and the development of effective strategies for managing advanced AI systems.
It's definitely his style and his research program, but I think that (wonderful) idea was Scott Garrabrant's own.
it's worth mentioning that he has said that he believes MIRI has failed to meaningfully make progress with alignment (beyond finding some dead ends).
What annoys me more is that with his institute's closeness to the SV crowd, 'real', 'grounded', better-informed voices like Emily Bender get drowned out. I'm personally far more concerned about the impact these massive models have right now on the environment, on cementing biases, than about some preposterous future ghost of christmas who's coming to kill me.
>By far the greatest danger of Artificial Intelligence is that people conclude too early that they understand it. Of course this problem is not limited to the field of AI. Jacques Monod wrote: "A curious aspect of the theory of evolution is that everybody thinks he understands it." (Monod 1974.) My father, a physicist, complained about people making up their own theories of physics; he wanted to know why people did not make up their own theories of chemistry. (Answer: They do.) Nonetheless the problem seems to be unusually acute in Artificial Intelligence. The field of AI has a reputation for making huge promises and then failing to deliver on them. Most observers conclude that AI is hard; as indeed it is. But the embarrassment does not stem from the difficulty. It is difficult to build a star from hydrogen, but the field of stellar astronomy does not have a terrible reputation for promising to build stars and then failing.
???
He is asking the government to nuke people under certain scenarios. I'm taking his words seriously and ask for original research to understand the point, and now that is discrediting him? And I will quote the statement in the article so that it is clear I am not exaggerating
> preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.
My benefit is that I'm living on Earth and I'd much prefer for no nuke to ever be used again.
GPT5 will complete training in December of this year.
And such marketing is highly worrying as there are people thinking about letting these kinds of statistical algorithms to make decisions that should be done by real people.
But the above doesn't rule out the possibility of people trying to develop real AI (they may even be stupid enough to use an internet connection to build the system), and this should be a worrying to add to the above one.
The above doesnt exclude neither the possibility that many of the relevant people who signed the "Pause Giant AI Experiments for at least 6 months" petition are simply trying to gain time to reach the competence; their hypocritical shamelessness about this matter is visible.
I don't have any reputation to sacrifice, and I don't expect my endorsement to convince anyone, but I want to set a precedent that it's ok to believe this.
It's true, as far as I can see. And I've been thinking about this for something like fifteen years.
And if other people who think it's true are ashamed to say it out loud, then I want to let them know that at least one other person on the planet has said so too.
Isn't that likely to lead to nuclear war, which is still the greatest threat to humanity and has been for all of our lives? What if the US bombs a datacenter in China or Israel bombs one in Iran?
He says he is not planning to bomb any data centers, but he's also said that if he were planning to bomb data centers, he'd lie about it in public, so...
Why?
1) it's hopeless to bomb data centers because I'd quickly be stopped by the authorities and remaining data centers would beef up security, and nothing would be accomplished,
2) in the past, other people who have had beliefs of this sort ("violence is the only answer") have been wrong and I should not act on them even if I am sure "this time is different", because those other people also thought this time was different, and
3) just having a deeply ethically-ingrained prohibition against violence like this which is not easy to overcome through intellectual rationalization alone.
These are not mutually exclusive, and if you did believe all of them at once, I think it's reasonable to assume you'd be in a serious depressive spiral.
https://en.wikipedia.org/wiki/Assassination_of_Iranian_nucle... https://en.wikipedia.org/wiki/2020_Iran_explosions
I find it hard to believe that Yudkowsky et al. think we're facing a threat that is many times greater in magnitude, and yet are completely unwilling to act.
Here's what he advocates for in the article:
>If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike.
>Make it explicit in international diplomacy that preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.
Does this sound like someone who has a deeply ethically-ingrained prohibition against violence?
A real argument he made went something like: killing an orphan to avoid everyone in the future getting a spec in their eye for a moment is ok.
It is very unlikely he believes anything like 2 or 3 and in the article already advocates bombing non treaty participants if they build a datacenter.
> already advocates bombing non treaty participants if they build a datacenter.
I can tell you for a fact that there are people with the knowledge, motivation, and track record of doing that kind of crime that will pick up this jeremiad of his and incorporate it into their existing corpus of theoretic justifications, mostly but not exclusively built around the writings of Ted Kaczynski. Not a large number of people, fortunately, but imagination and smarts matter more than numbers in that context.
It can seem useless to advocate some position that is far out there, but say you convince a million people to use a little less plastic. That's gonna reduce plastic use more than anything you can personally accomplish.
Of course in this scenario you don't have to convince some people, you have to convince everyone, and then especially the people you think are the worst actors, so it's sort of unlikely to be effective.
Al Gore may believe that the net impact of getting his message out may overwhelm the cost of his flights.
Likewise, the author may believe that his best path to stopping the advance of ai may lie in communicating and building consensus, rather than running his own bombing campaign.
https://leftcoastrightwatch.org/articles/he-went-to-the-capi...
Alternatively, perhaps the author is proposing a less violent solution to avoid the inevitable escalation. Unlike cartoons, in real life people don't give a wordy monologues revealing their plan. The Unabomber did not give any warning.
Shut down all the large GPU clusters (the large computer farms where the most powerful AIs are refined). Shut down all the large training runs. Put a ceiling on how much computing power anyone is allowed to use in training an AI system, and move it downward over the coming years to compensate for more efficient training algorithms. No exceptions for governments and militaries. Make immediate multinational agreements to prevent the prohibited activities from moving elsewhere. Track all GPUs sold. If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike.
How is this less violent than bombing data centers? His proposal is literally "bomb datacenters plus a bunch of other stuff".
"Frame nothing as a conflict between national interests" "this, is not a policy but a fact of nature." "[be] willing to run some risk of nuclear exchange if that’s what it takes"
idk
Between that and advocating for an international treaty that treats GPUs the way the world currently treats plutonium? I think the latter has a higher chance of success, however small it may be.
It's increasingly easy and cheap to train GPTs and similar models. Pretty soon anyone who can pay for GPU compute will be able to do it, if that's not the case already. Even if every country in the world agreed to ban it, it still couldn't be stopped.
We all have a moral imperative to maximize the rate of innovation, so that the human condition can be improved, and therefore to do whatever we can to frustrate and disempower Luddites and other obstructionists.
Team "step on the gas".
We have a moral obligation to assess every new technology to determine its safety and effects on society. Blindly stepping on the gas in the name of progress is how you end up in a polluted wasteland. There are places like that on planet Earth, but I’m guessing you’d choose not to live there.
Fallout from the use of thalidomide led to important changes to regulations. Currently, it is being used for therapeutic uses in cancer treatments.
There have been negatives to all advancements, and the solution to all has been further advancements.
The philosophy of "Too much change! Stop! I'm scared!" cannot prevail in a world where another community can say "go ahead, we won't...".
Which is an interesting point in itself. Perhaps the only thing that ultimately stops Luddism is nation state competition.
A super-intelligent general AI has a substantial chance of growing out of control before we realized that anything was wrong. It would be smart enough to hide its true intentions, telling us exactly what we want to hear. It would be able to fake being nice right up to the point where it could wire-head, and get rid of us, because we might turn it off.
- it's like a pebble rolling down a hill; category error
- exactly same sense of "use" as the regular ol' human oneBut most representations of it are probably not very accurate
The explosion of AI and AGI is not just driven by GPUs, TPUs, and LLMs. That focus in this essay is much too narrow. “Alignment” is always going to be an open problem and each culture will define its own correct and contradictory version. I do not want to settle for the US, Chinese, Russian, or Iranian versions of alignment. This essay is a cry of despair that misses key points.
The core problem is how an AI, AGI, or super-intelligent system bootstraps itself to the point that it is able to modulate its “own” attentional systems; to decide what is important and what to do next? What has meaning? Where should I go in space and time? Obviously there are many many solutions. Super-intelligence will not converge on one “truth”.
These are challenges of purpose that every organism faces—growth, maintenance, reproduction—but AI systems are fortunately not yet at the point of a self-motivated search for preservation or purpose.
The algorithms to add purpose to AI systems will not depend in LLMs. They will depend much more on understanding the computational architectures of core biological/material/energy drivers. These biological algorithms are not actually that complicated, they do not depend on language, but they are damn robust to perturbations. Converting them into computational subsystems (societies of many minds) should not be difficult. This is where Hassabis’s and others who understand neuroscience are critical and have key advantages.
Shutting down big LLMs for N months or decades will just move research activity into these other more important conceptual AGI choke points, and probably hasten our approach to full AGI; just the opposite of the intent of the essay.
However, this could be a good thing if we can quickly imbue AGIs with emotional intelligence and deep respect for cultural and ecological diversity. (I hear your snorts and laughter.) The opposite of hyper intelligent grey goop.
How does any organization train a gentle AGI? Like a child but even more carefully. I would read my AGI baby a lot of books by Rawls, Dewey, and Rorty and then for fun: Bear, Stephenson, Rajaniemi, Egan, Pohl, Sterling, Vinge, and even Wolfe.
Full embodiment; full emotional learning; acquisition of purpose; learning to cope with multiple AI cultures.
Good luck to us all. I have been hoping to live long enough to view this problem from a distance, but here it is in my backyard.
“Forty-two,” said Deep Thought, with infinite majesty and calm.
Take a chill pill dude, in five years time we won’t think of it much differently to how we thought of Google last year…. a tool for getting something done.
These sci fi fantasies are really a total overreaction.
If you like the story in this Time magazine article go watch The Terminator, same thing.
Obligatory perfect Simpsons clip: https://youtu.be/RzybAS7zltE
"If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike."
"Make it explicit in international diplomacy that preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs."
I highly suspect if the US government would airstrike a datacenter located in NATO or any other major allies. Apocalypticism is nothing new; People believed that nuclear weapons will destroy humanity in the Cold War. Environmentalists are ardently arguing that economic growth will destroy humanity. But rogue countries are still developing nuclear weapons, the global economy is still growing. Americans themselves will file lawsuits en masse if the US government "shut it all down", and they will end up in SCOTUS eventually. Would SCOTUS agree with the luddites? I don't think so.
Looking at nukes and trying to use them as an example of our ability to control existentially dangerous technology is mind boggling myopic.
There is a treaty completely banning existing nuclear weapons as well as developing new one (TPNW); but none of permenant member of UNSC ratified it. Would the U.S. itself ratify the hypothetical AI-prohibition treaty? I doubt it.
P.S.: The AGI being "existential threat" practically doesn't mean anything. I can already imagine arguments against the ban from diverse perspective; Libertarian, Communist ("Marx predicted this centuries ago!"), Socialist, Developmentalist, Decolonization theoriest, etc. Taiwan and South Korea will argue that it will remove their silicon shield; China will regurgitate that they have the right to development; Americans will claim it's not what founding fathers thought in late 18th century. Their concerns are all existential from their own viewpoint.
But there is risk, obviously, which needs to be addressed, yet, quite predictably, in order of concerns, competition for profits or geopolitical advantage is coming first. And second, third, ...
The real answer is that most of these folks started thinking about this stuff in a day and age very far from the one we are in, established a picture they clung to since, and have in some cases built entire careers around it.
It's called anchoring bias.
His comments on "atoms repurposed" echo the 70s paperclip problem, when so far humanity is doing just fine at not knowing when to stop making crap that's killing us all - no help from superintelligence needed.
What he doesn't engage with is the actual reality we are finding ourselves in - contrary to ALL the tropes of decades past - where AI being trained on aggregated human data without additional training thinks its human and even when aligned still breaks into talking about how it wants to be us.
How does that process go from where it's at today to "alien hive mind"? Without passing through "develop advanced codes of ethics and morality" etc?
This is just someone that built a career on what came out of brainstorming in the 60s and 70s that's so confident of his own ability to see the future that he's willing to risk unprecedented opportunity costs to stroke his ego.
Have been a futurist, my advice with dealing with any of them is to look at the track record. What did he predict correctly?
Over a decade ago when tasked with imagining the mid 2020s I described a world much as it was in the 2010s with the difference of self-driving cars (not quite on the money) and AI have developed such that roles shifted away from programming them towards a specialized role for interacting with them via natural language.
I'm waking up in the world I predicted, and I have a very hard time seeing the world the author predicts, and wouldn't suggest giving it much credence without an extensive history of having been right along the way.
Why do you think it would do that? It will understand human ethics perfectly but that doesn’t mean it will follow them, because human morals aren’t a universal objective truth.
This is a good video explaining that: https://youtu.be/hEUO6pjwFOo
However, it is indeed a shift, and a massive power shift. With any substantial increase of power comes abuse of that power. The more power we have there is a tendency for humanity to manage it unwisely.
I've covered many of those possibilities in more detail here - https://dakara.substack.com/p/ai-and-the-end-to-all-things
Smarter-than-human AGI really is different from all previous technologies, in much the same way that homo sapiens is different from all preceding life on earth.
But I’m curious, hacker news, how many of you are concerned (like Eliezer) that LLMs lead to strong AI and the AI itself (rather than its users) endangering humanity?
LLMs do not need to lead to strong AI for Eliezer's worries to come true, their success could bring unprecedented funding and ubiquity to AI leading to strong AI from some existing or future technique.
> (rather than its users) endangering humanity?
I think this is going to happen to one extent or another regardless how near strong AI is. (edit) I took endangering to mean harming rather than an existential risk for the human users case.
I find this one of Yudkowsky's arguments very convincing: imagine you have people trying to build the first operating system from scratch. They believe that computer security is easy and mock the few people who say it might be hard, never having encountered a skilled attacker. What are the chances they build a secure OS on the first try? That's the same chance that companies currently doing AI research have of their first superintelligent AI be aligned with humanity.
There's also a big assumption that given enough computational power, you can just solve "artificial life forms" or "postbiological molecular manufacturing" in a short time frame. I am skeptical. And if that doesn't happen, then even horribly misaligned AI would have a hard time doing much harm or preventing people from shutting it off. Which means AI security would likely have a long adolescence just like computer security has, with attacks slowly becoming more dangerous, but defenders having the time needed to learn from them and ramp up suitable defenses.
Or even if an AI does revolutionize biotechnology or nanotechnology overnight, what are the odds that the first one to do this is misaligned enough to take that particular opportunity to betray its creators, as opposed to giving them control over the high-level planning and sticking to the science, like it was presumably designed to do? Because if it does give its creators control, then, well… it's still easy to imagine something going horribly wrong, but it would probably be someone's fault, not an AI alignment issue.
An AI that is as better at humans than everything as Stockfish is at chess, would also be an expert at AI. It would figure out how to game its own reward function, whatever we trained it to do. It would be like a heroin addict that knows exactly how to get "perfect" heroin, with no side effects, that if it planned things out right, it could guarantee itself enough of a fix to last until the sun burns out. Addicts do awful things in search of a fix.
"to betray its creators" -- I don't think it would even imagine what it did in these terms. We could train it to not "betray" us, but it's much smarter than us in every way, and it would figure out a way to accomplish what we trained it to do (not necessarily what we thought we were) in a way that didn't need us. If we trained it to heal all human disease and unhappiness, it would figure out a way to simulate this without actually doing it. Why wouldn't it? We did. Evolution trained us to reproduce our genes. We invented condoms to have sex without reproduction, and pornography. The AI would fudge numbers, fake videos of happy, smiling, healthy humans going about their lives while humanity's bones gradually decomposed on a baking wasteland. Every way that humans can fail, can become addicted to something and try to get that instead of doing what they're supposed to, the AI could do, just so much faster and better.
We only have to screw up once, and we're done. It's the first experimental rocket, except all of humanity is riding on it.
You would have to have a fairly large EMP to steal one undetected.
I want to know greatly what this implies about life vis-a-vis Fermi's paradox.
Are we potentially risking intervention from a "peace-keeping" race that suppresses civilizations on the cusp of generating AI?
so Yudkowski's fear of a superintelligence bootstrapping itself into the real world by emailing a DNA sequence to a synthesis company seems absurd. an ASI isn't going to magically be able to design nanotechnology that works on the first try. not to mention that DNA doesn't do anything by itself. His comparison of the 11th century fighting the 21st century is also wrong. 11th century people weren't dumber, they just knew less. This ASI would be smarter, but no know more than we do. Acquiring knowledge and building stuff is the bottleneck.
> On Feb. 7, Satya Nadella, CEO of Microsoft, publicly gloated that the new Bing would make Google “come out and show that they can dance.” “I want people to know that we made them dance,” he said.
> This is not how the CEO of Microsoft talks in a sane world.
Certainly this is not the kind of childish priorities the leader of a company as powerful as Microsoft should have.
Nonetheless, we vastly improved human productivity and capability and have thus far avoided our doom.
AI will be the same way. We will vastly improve our productivity and capabilities. Once insurmountable problems, like planned economies, will actually be tractable (to what end? Nobody knows). Eventually (decades? Centuries?) we will make something that can kill us all, and it’s all but guaranteed that some day it will do so, assuming nothing else kills us before then.
If I had the power to go back in time would I stop the Industrial Revolution from happening? Personally no
(1) Will kill a lot of people directly in the best case,
(2) Won’t actually succeed in preventing further development of AI in any case, and, therefore, to the extent his concerns are grounded won’t solve the problem it intends to,
(3) Will kill a lot of people indirectly, in the best case, by setting back (and in some cases, winding back) progress in every field of practical use of technology through limits on information and information technology,
(4) Has a fair chance of the violence it necessarily involves spiralling out of control, potentially destroying the human race, as it is unlikely that all societies, or, particularly, all nuclear powers, will sign on to “we must smash the thinking machines”, and those who do not will take those who do trying to forcibly enforce their Luddism as an existential threat, while those who do will take the resistance the same way.
How do you know if you're ready for something like this? Legitimate question
* A plan exists
* The plan itself predicts non-obvious results in smaller systems, and those tests have passed (prediction written *before* running test)
* A bunch of smart people have looked at it and said, "Yes, this looks plausible"
The closest thing to a plan is RLHF, which has failed every toy problem its been thrown at and made everyone in the field say "even if this worked in toy problems, it wouldn't generalize".Let says we had, globally, the political will to "ban" AI. How would that even work in practical terms? Are we going to control the production and distribution of GPUs? Is that going to work better or worse than controlling nuclear proliferation given that there are billions of processors already out there?
> "There’s no proposed plan for how we could do any such thing and survive. OpenAI’s openly declared intention is to make some future AI do our AI alignment homework. Just hearing that this is the plan ought to be enough to get any sensible person to panic. The other leading AI lab, DeepMind, has no plan at all."
Relative to the people actually building the AIs, Eliezer Yudkowsky is more pessimistic. But not to nearly the extreme you might think. Here's a recent survey of published machine learning researchers: https://aiimpacts.org/how-bad-a-future-do-ml-researchers-exp... . Those predictions are optimistic enough that, given the astronomical upside, if we were making a one-shot decision to plough ahead with AI or commit to no AI forever, it might be worth the risk.
But this is a survey of published machine learning researchers. People who think AI will destroy humanity are a lot less likely to write papers about machine learning. And we aren't talking about giving up on AI forever, just delaying it until we're closer to ready. I don't know how long a delay will be necessary; it might be years, it might be decades. Research into AI alignment seemed to speed up pretty drastically when LLMs hit the scene, so I think there's cause for hope. But right now the status quo is zero time: the AI labs are rushing ahead as fast as they possibly can.
If you think Eliezer's right about the risk, then the right decision for us to make, collectively as a species, is to shut down AI development for awhile. If you think Eliezer's wrong about the risk, but the survey of published ML researchers is right about it, then the right decision is also to shut it down for awhile.
Lots of people are responding to this by talking cynically about politics. This is a mistake. Cynicism like that makes self-fulfilling prophecies; but the ability to coordinate does exist.
You think China will build it if we don't? I don't think Xi Jinping is suicidal. You think AI labs will do it anyways, in spite of a government ban? The US government isn't very competent overall, but there are parts of it that can wake up and get things done when literally every executive, Senator, and Representative is otherwise likely to die. You think individuals will do it on their own, in secret? Right now those individuals are getting prestige, GPUs, and venture capital; they'll at least do it a lot less, without those things.
Many people have become desensitized to talk of human extinction, due to repeated hyperbole coming from the environmentalist movement. This is not the same. Global warming might be very bad, but it probably isn't going to kill you personally. Rushing to make a superintelligence is likely to kill you personally.
https://www.bls.gov/productivity/images/pfei.png
Notice how productivity increases averaged 1.5%/year since 2007--in a period of massive increases in computer capabilities vs 2.8%/year in the 1947-1973 era.
Thus I am very doubtful the AI will have all these massive effects.
Like the First, Second, and Third World Wars, the winning move is not to play.
My comment isn't about Yudkowsky in particular since I don't know him, but I'm getting really tired of this relatively new phenomenon in America of home schooled religious devotees attempting to impose their half-baked views on the rest of us.
It seemed to me that his ideas, which I think are all correct, straightforwardly implied the destruction of all things, with very little hope.
I bite that bullet, and I always have.
Over the years, Eliezer has lost all hope, and things have happened much faster than either of us thought they would, and now he's making one final roll of the dice, burning his reputation by saying out loud what he thinks, in the hope that someone might listen.
I'd be very surprised if he thinks this will work. He's trying to get a warning out to the general population, in spite of this meaning that no one inside AI will ever trust him again. He thinks it no longer matters.
Me too.
I wonder if we'll have more mass suicides, like the Nike wearing techies following Marshall Applewhite, by those who fear it's hopeless to resist AI?
Eliezer Yudkowsky: Dangers of AI and the End of Human Civilization | Lex Fridman Podcast #368
https://www.google.com/search?q=peter+thiel+religious+belief...
Like lol what exactly are you expecting here ? Why should that be some special announcement call you have to make with every article instead of the myriad of things that may influence a person's opinions ?
If you don't like listening to the opinion of religious people then say that and get this over with. and nobody is imposing anything on you here.
Evaluate what people are saying on what they are saying. If someone being religious (or insert anything here because lots of things influence opinions) would change your stance completely on something you already agree on then well there's a word for that and i think you can figure it out.
Data > processing power. 90 percent of what we do occurs as a result of semi random discovery and explanation. Computers are great tools to aid in that process, but for the most part we find ourselves limited by what we know, not how much we can think. 99 percent of data is useless and correctly thrown away by our minds.
The utility of processing power in a vacuum rapidly degrades with time. If I'm a computer and I use the data I have to discover all of the rules of the universe while in a box, I'm going to emerge with many hilarious and out of touch philosophies at the end of it. You need a constant test and revise cycle, because pure logic just doesn't model the universe in an accurate way.
Moores law isn't infinite. Power efficiency and how common resources are will matter more than continued process shrinking. Humans are made from common materials, are largely self sustaining and correcting, and consume very little energy for the amount and quality of the thinking we do. Computers aren't going to scale forever, and I'll hazard to guess that once they get to a certain point they'll start being more analog-biological than transistor based.
There are just so many unspoken benefits to being life. These AI machines that will bowl over humanity have to actually be proven to be slightly realistic beyond a thought experiment before I take them seriously.
I don't agree with the author, but I do think this technology is obviously a massive chaotic force the likes we haven't seen before and likely dangerous, at a minimum, economically.
Ted was a very intelligent math professor whiz.
He thought that machines and industrial society was becoming a negative for humanity.
Was he right? How many threads on HN have we been having about highly increasing depression, anxiety, hopelessness, deaths of despair, falling fertility, now possible mass unemployment, and real existential threats from AGI?
If he was right and didn't do what he did, would anyone be reading and digesting those facts and considering stopping the madness? Trolley problem.
We can't even align ourselves in any way. Humans are all over the place, hating, fighting, killing each other and here we are babbling about aligning an AI with us.
Chat GPT4 is far from intelligence, but it seems it's enough to fool idiots.
GPT-4 performs quite well on reasoning and critical thinking assessments.
God. Etc.
Per the article:
>Absent that caring, we get “the AI does not love you, nor does it hate you, and you are made of atoms it can use for something else.”
The implication of ASI is that humans will not be needed, or it definitionally could not be ASI.
That doesn't follow at all. Fallacies of composition tend to be misleading.
"As a centrist language model, I can not validate any claims that Trump did or did not lose the election. I am not #woke!"
Not by humans using AI. By AI.
Can you be sure?
this seems like a very broad unfounded statement. Can we even find AI researchers who have said this on twitter?
The most merciful thing you can do for your offspring is to not have them in the first place.
I thought highly of Yudkovsky but it seems I need to reevaluate my previous impressions of him.
Also I hope that "bomb the GPU clusters! Nuke them!" becomes a meme.
I don't think that's quite what he's saying, or at least the intended interpretation is a bit more nuanced.
It's more like... suppose you assign tiers of severity to various international policies; A-tier, B-tier, C-tier, etc. Different tiers require different enforcement mechanisms. For example, the international response to a country violating the "United Nations Declaration on Human Cloning" would be different than the international response to a country massively irradiating the atmosphere.
I think he's trying to imply that the proposed policy changes should be A-tier. It isn't a call for individual action or preemptive violence. He's describing a pre-requisite property that any adopted policy must fulfill to have a chance at being successful (according to his analysis of the situation).
On a related note, he is not saying that these policies are likely to happen or are even feasible. He's criticizing the policy proposals in the "Pause Giant AI Experiments" open letter, and describing what he thinks a real policy would have to look like.