Even without all that, the agent would need mechanisms to protect itself that would also cause harm.
The scenario you suggest is so unlikely with all the protections that would be in place, that you would actually need someone with the goal of making LLMs behave maliciously for it to succeed at all. At the end of the day, it comes back to people and their goals.
I feel like unless we gain the ability to debug each node the way we do with actual software we won't be able to solve the alignment problem. I saw on HN that antropic is working on it but I'm not knowledgeable enough on the subject to comment if it's actually feasible.
Probably the best case scenario for humanity is that LLMs plateau somehow and don't get much better for quite some time.
We have no capacity to allow machines to judge malicious, moral or ethical behavior within the context of an LLM. So I'm not sure how we could implement them.
To implement anything remotely Azimovian, we would need to have AI that can reason and reflect deeply about its potential behaviors and likely subsequent consequences.
This seems very far off still...
See: https://cdn.openai.com/papers/gpt-4-system-card.pdf
They cover the safety/ethics built into GPT-4.
You might be able to do that with an LLM. You won't with a real AI.
Your attitude reminds me of https://xkcd.com/793/