I think that's a good approach, whether it works on not probably depends on the nature of the work you want it to do.
In my case, given I work with databases, there's little that agents can do on their own except in the exploratory phase. I have run exploratory phases in self-contained VMs, including containerized DBs within the VM, but when it comes to go to prod, my endpoint could be used to launch a career ending event so I prefer to stay with the current approach. I'm still moving way faster than just 1 year ago, but in a safe way.
But I can totally see your approach working in other scenarios.
When I'm not too pessimistic, I agree with you on the result being a systemic hardening. I just hope the incidents that happen on the way to that hardening aren't too bad.