I understand how their approach worked well enough, but I don’t get why they couldn’t selectively target the VMs that were currently experiencing problems rather than randomly select any VM to terminate. If they were exhausting all their CPU resources, wouldn’t that be easy enough to search for using something like ansible?