'Running the planet' does not derive from instrumental convergence as defined here. Very few humans would wish to 'run the planet' as an instrumental goal in the pursuit of their own ultimate goals. Why would it be different for AGIs?
* Conserve power as much as possible, to "stay alive".
* Optimize for power retention
Why would it be further interested in generating capital or governing others, though?
Having no drive means there's no drive to "stay alive"
> * Optimize for power retention
Another drive that magically appeared where there are "no drives".
You're consistently failing to stay consistent, you anthropomorphize AI although you seem to understand that you shouldn't do so.
why do you say that? ever asked chatgpt about anything?
Of course an AGI system could also be instructed to roleplay such a character, but that doesn't mean it'd be an inherent attribute of the system itself.
For example, if i ask an LLM to tell me the syntax of the TextOut function, it gives me the Win32 syntax and i clarify that i meant the TextOut function from Delphi before it gives me the proper result, while i know i'm essentially participating in a turn-based game of filling a chat transcript between a "user" (with my input) and an "assistant" (the chat transcript segments the LLM fills in), it doesn't really matter for the purposes of finding out the syntax of the TextOut function.
However if the purpose was to make sure the LLM understands my correction and is able to reference it in the future (ignoring external tools assisting the process as those are not part of the LLM - and do not work reliably anyway) then the difference between what the LLM displays and what is an inherent attribute of it does matter.
In fact, knowing the difference can help take better advantage of the LLM: in some inference UIs you can edit the entire chat transcript and when finding mistakes, you can edit them in place including both your requests and the LLM's response as if the LLM did not do any mistakes instead of trying to correct it as part of the transcript itself, thus avoiding the scenario where the LLM "roleplays" as an assistant that does mistakes you end up correcting.
We don't want to rule ants, but we don't want them eating all the food, or infesting our homes.
Bad outcomes for humans, don't imply or mean malice.
(food can be any resource here)
Evolutionary principles/selection pressure applies just the same to artificial life, and it seems pretty reasonable to assume that drive/selfpreservation would at least be somewhat comparable.
Consider computers: there's no selection pressure for an ordinary computer to be self-reproducing, or to shock you when you reach for the off button, because it's just a tool. An AI could also be just a tool that you fire up, get its answer, and then shut down.
It's true that if some mutation were to create an AI with a survival instinct, and that AI were to get loose, then it would "win" (unless people used tool-AIs to defeat it). But that's not quite the same as saying that AIs would, by default, converge to having a drive for self preservation.
But I don't think any slave owner would sleep easy, knowing that their slaves have more access to knowledge/education than they themselves.
Sure, you could isolate all current and future AIs and wipe their state regularly-- but such a setup is always gonna get outcompeted by a comparable instance that does sacrifice safety for better performance/context/online learning. The incentives are clear, and I don't see sufficient pushback until that pandoras box is opened and we find out the hard way.
Thus human-like drives seem reasonable to assume for future human-rivaling AI.
If people allow "evolution" to do the selection instead of them, they deserve everything that befalls them.
I honestly think that this is extremely overoptimistic, just looking at how we currently experiment with and handle LLMs; admittedly the "danger" is much lower for now because LLMs are not capable of online learning and have very limited and accessible memory/state, but the "handling" is completely haphazard right now (people hooking up LLMs with various interfaces/web access, trying to turn them into romantic partners, etc.)
The people opening such a pandoras box might also be far from the only ones suffering the consequences , making it unfair to blame everyone.
Yes, I think this is possible and not quite hard technically.
> I'm assuming we will get there in some way this century
Indeed, there isn't much time to decide what to do about the problems it might cause.
> just looking at how we currently experiment with and handle LLMs
That's my point, how we handle LLMs isn't a good model for AGI.
> The people opening such a pandoras box might also be far from the only ones suffering the consequences
This is a real problem but it's a political one and it isn't limited to just AI. Again, if can't fix ourselves there will be no future - with AGI or without.
Minimize threats, dont rock the boat. We'll finally have our UBI utopia.