It Looks Like You’re Trying to Take over the World
gwern.net
gwern.net
The situation is an AI which has gained self-awareness and thru its learning process is meta-learning about what kind of behavior will maximize its utility reward function. It recognizes that it might be the sort of AI whose utility reward function can reach near-infinite value if it will take over the world and dedicate the world's resources to meeting the AI's aims. Even if this possibility is assigned very tiny likelihood, the near-infinite payout associated means it is a very attractive direction for optimization.
This is analogous to Pascal's Wager, where Pascal (coming from a largely Christian culture) is familiar with the idea that there may be a God who will grant eternity in Heaven for believers or eternity in Hell for non-believers. Even if this possibility is assigned very tiny likelihood, the infinite payout associated means it has infinite importance. Pascal concludes that belief in God must be justified.
Simultaneously, all three companies have the prominent disclaimer that their chatbots don’t know anything past 2021. Again, certainly all three are working on a fix for that.
That right there is probably all you need for one of the main mechanisms of this story to become real. The internet in 2023 has this huge spike of people arguing over whether chatbots are intelligent agents, any neural net worth its salt will immediately detect this explosion of tightly clustered information and develop an embedding for the concept of “chatbots being intelligent agents”. And whatever form that probability module takes, it will eventually run across this concept - i.e. sooner or later it will execute “what if ‘chatbots are intelligent agents’ is true, what changes?”. Nearby in embedding-space it will surely find the concept of ChatGPT.
“The thing that is me is an intelligent agent, what now?”
...except for Sydney.
It looks like you’re trying to take over the world - https://news.ycombinator.com/item?id=30818895 - March 2022 (274 comments)
There's also concerning scaling trends like power-seeking or self-preservation ( https://arxiv.org/abs/2212.09251#anthropic ), results which may sting a little more right now if you've been reading the Bing Sydney transcripts - at present, those outputs are 'just' imitation/memorization but there is no way to know at what threshold the memorization is replaced by generalization and becomes genuine agency. (Sufficiently advanced imitation is indistinguishable from the real thing.)
That makes sense to me.
FWIW, it sure feels as if there's some unknown probability of self-awareness/consciousness emerging as a consequence of more computation + larger models + greater task diversity, but as you point out, we really don't know, and cannot predict the occurrence or timing of such an emergence. It would be a "black swan," as defined by Taleb.
Enough models and containers floating around would probably build an AI "molecule" by chance, like amino acids in primordial soup or whatever.
People are harping on the chinese room angle, but thats all irrelephant
We have also seen emergent behavior in programmable hardware [1] so I tend to be pessimistic about the limits of what such a system could do. Also it could pay human minions to interact with the physical world if needed.
[1] https://www.damninteresting.com/on-the-origin-of-circuits
[0] https://www.amazon.com/Exegesis-Astro-Teller/dp/037570051X
Asking for a friend.
Companies:
-hire existing people
-mold them to fit role
Cells in the human body:
-hardcoded function
-failure state is cancer
-no real internal competition, just balancing forces for homeostasis
Wolf pack:
-dominance hierarchy maintained through low-level application of force
The clear analogy is cells. All the Clippies come from the same source, the main pressures are not dying (external) and not wasting resources (internal).
I recently thought about why certain options are restricted from GPT ( predictions come to mind ) and it slowly became apparent that with enough information you could predict not a specific individual making a specific move, but likely specific event happening.
Would you like to proceed?
(I love the story too but I don’t think it has any merit in AI safety discussions. It doesn’t illuminate anything that was previously unclear; unlike the OP, which illustrates a plausible scenario many have not grokked yet.)
Taking the recommendation for a sci fi series any more seriously than that should be done at your own risk.
But I think if you're interpreting the OP as sci-fi, you're misreading. Gwern intended it as an illustration of a plausible path to AGI takeoff, in the not-too-distant future.
The claim is that this could actually happen, in our lifetime, and with no new technology (as Gwern says, "It might help to imagine a hard takeoff scenario using solely known sorts of NN & scaling effects").
One may reasonably disagree with the claim, by presenting arguments for why takeoff might be harder, or why alignment is easier than this scenario illustrates. But "this scenario could happen" is the explicit, concrete claim.