I can imagine it would be a tough sell to the CEO. :)
I can imagine it would be a tough sell to the CEO. :)
I remember reading years and years ago about bandit algorithms... this kind of ops work is at a level that's found only in a few different companies.
The "sell" can be tricky for some people, until your first production issue. Never let a crisis go to waste.
Machines go away all the time in the cloud. This tool increases the frequency so you can ensure your system handles it gracefully.
Some people believe their system can tolerate this class of failures, but without continuous validation, that is more of a hope than a certainty.
Why? Your customers use your production environment, not your test environment. Something will cause loss of an instance for you: * Mistaken termination * AWS retirement (and you missed the email) * Cable trip in the data center * <Something else we can come up with> * <This list goes on>
So, vaccinate against the loss of an instance cratering your service. Give your prod environment a booster shot (with Chaos Monkey or something like it) every hour of every day. Then, when anything from the above list happens you're infrastructure handles it gracefully and without intervention. Continued booster shots ensure that this stability continues through config changes, software version changes, OS changes, tooling changes, etc.
I think the better question is "Why wouldn't you do this?"
See https://medium.com/production-ready/chaos-monkey-for-fun-and...
As for chaos being a hard sell: https://medium.com/production-ready/chaos-engineering-a-shif...
You catch bugs, and no one says you can't run Chaos Monkey in staging or a similar environment if it really is a tough sell.
The biggest issue IMO is explaining need to make things more resilient. Actually the technical people (mainly developers) might be the biggest obstacle, because it adds more work for them (with no visible benefit to them, because when application fails it's ops who get woken up).