Another issue is just the big whackload of nobody really knowing what is possible. Could an AI in, say, the next year, become Skynet and decide that it could just bump humans off and become its own independent economy? Probably not successfully, but how much trouble could it make if it incorrectly decided that was the right thing to do? What if someone says "hey, you're sandboxed and we're testing out what's the worst you can do, just start hacking around and do as much damage as you can in this simulated environment" backed by a few million bucks in tokens, or the AI hacking itself a few million bucks in tokens? How much damage can it do? The only honest answer is, nobody knows. Humans would react, after all. If necessary we can pull a lot of plugs still. What is inconceivable in peacetime can become inevitable in wartime.
In some ways, it's possible the best thing Anthropic could do right now for humanity in the long term is actually exactly that... just equip the AI with a limitless token budget and let it go nuts with the goal of inflicting maximum damage. Send someone with an axe down to the data center to cut the power on command. Let humanity really see and feel the risk, because probably right now it can do a lot of damage but not actually destroy us. We're probably worse off with them getting 10-100x smarter and then some process starting with that goal. Trying to keep the AIs "aligned" and 99% succeeding could end up being worse than failing at it right now. Of course it would be the end of Anthropic as a going concern, but falling on their sword might be the best thing for humanity.
On the other hand, who's to say we're not already past the point where that would do an absolutely unacceptable amount of damage? I sure wouldn't care to be personally responsible for pushing that button. Hypotheses about possible future positive value in a rationally time-value-discounted future would be dominated by the much more certain and temporally much closer negative effects.
Personally I don't think that if this scenario is likely that we actually get to the point that we get an independent Skynet that calculates (correctly) that it can survive without humans. It's far more likely that some accident takes civilization back to the point before it can run AI at all, but humans still survive. Right now the really good AIs are still locked to data centers. They can't just distribute themselves widely because they can't run at any reasonable rate distributed like that. And then there is always the possibility that this is all overblown and the existential risk isn't anywhere near as special as we think it is... one can use an ecological analogy to suggest that no AI is necessarily any more likely to completely dominate the ecosystem than any particular life form is.