Great question, I might well write another post about this because it's a really interesting problem.
The software and services that we looked at to address this problem all left us pretty lukewarm. We needed distributed configuration and resource management, logging and alerting at a minimum.
To cut a long story short, we've written our own system, which draws on ideas from things like rush, capistrano, munin and nagios and tailors it for our needs.
At the moment, we have a system which gives us the information we need to make decisions on commissioning new servers etc, and wakes us up in the middle of the night when something goes wrong. I expect the capabilities to grow to include auto-scaling, auto-healing and some other bits and pieces as I get the time and our needs mature.