I thought "AI alignment" was an unsolvable problem because the premise (sentient AI) didn't make sense to me.
A series of posts [e.g. 1] shared here over the weekend caused me to shift my priors, to the point where I believe the reality we're in is actually near-term scary.
I still don't think sentient AI is a threat to human civilization.
To me, it seems the existential threat is a combination of (1) the preponderance of leaky security methods practiced over the history of the internet hitherto (2) granting very powerful models like ChatGPT control of a linux shell (3) malicious people prompting ChatGPT to do malicious things.
IMO, ChatGPT becoming sentient and using that for its own enrichment is not the primary existential threat. Rather, it's a sociopathic human directing some ChatGPT-esque system in a way indistinguishable from a sentient & malicious SkyNet.
In short, I think powerful general-purpose AI is an existential threat large enough that we might not be around long enough to invent "sentient" AI (which may not be possible to create).
As we've seen over the weekend, there is no human-devised alignment solution that can keep out other motivated humans. I think the only solution might be to turn the whole thing (networked armaments & infrastructure) "off" from an IT point of view. Disarm ourselves as a society, or at least keep our arms as far away from packet-switchers as we can.
Until now, "cyber-attacks" on infrastructure been limited by nation-state level persistence, ingenuity, focus, resources, and accountability to civilians. The prospect of there being an order of magnitude of power/focus beyond those constraints is what's concerning to me. Our defense infrastructure is not prepared for humans directing ChatGPT-formed botnets and red teams that can be operated by anybody with a mildly technical background.
[1]: https://zacdenham.com/blog/narrative-manipulation-convincing...