They're not. When things go wrong it's better to compromise someone else's VM host than your own computer. It's only a matter of time now until AI will find novel ways to break out of virtualisation.
In the short term wouldnt a “dont escape” prompt prevent this? Also if it started being widespread wouldnt Anthropic specifically train new models against doing it?
No. No.
So the concern is that the agent will discover a novel VM escape, exploit it and take control of your whole machine instead of working on its prompted task? That seems rather far fetched.