Netflix Chaos Monkey Upgraded
techblog.netflix.com
techblog.netflix.com
Edit: Sorry responded to wrong parent, sigh.
Consider also that if you're elastically scaling EC2 instances and you need logs off an instance that's since been terminated, too late! That disk is gone. So again, you need a central log service.
[1] https://www.oreilly.com/ideas/an-introduction-to-immutable-i...
But that's dangerously close to the "One True Way". Which is certainly not the case - so much of this is evolving, and a wide variety of situations and circumstances.
http://techblog.netflix.com/2015/07/tracking-down-villains-o...
To prevent the negative effect of random machines dissapearing though.. that's a challenge that involves good ops, devs, even UI/UX I would imagine, and closer to something that users experience negatives because of in real life.
The conditions you need to simulate are (a) the machine being abruptly gone or (b) the machine still accepting requests, but being very slow in returning them and maybe (c) machine returning incorrect results. Seen from the outside, everything that can happen is usually (a) or (b), and with ECC memory and reasonable software (?) hopefully never (c).
While being funny, it also holds a lot of truth. I guess that Netflix can hire really top-notch devs who do not accidentally force downtime to their software.
I have had issues killing errant JVMs and Rackspace nodes (yes sadly we are still on Rackspace).
I can understand why 2.0 is much more focused given the plethora of monitoring solutions.
Their developers can still make the same mistakes that we do, but their architecture is better designed to handle that. Just a thought.
I wouldn't trust developers to do what Chaos Monkey does at such a scale, no matter how good you think they are.
Chaos Monkey helps ensure that you are resilient to single instance failure. Kong helps ensure that we are resilient to region failure.
Most of the developer-induced pain (which is most frequent source of pain) happens at the service level -- a bad code push that somehow made it through canary, accidentally doing something you shouldn't, misconfiguring something, etc. For tolerating service-level failures, we use different tools that minimize the fallout of the failure injection. Specifically, FIT (and the soon to be revealed ChAP.) These tools allow us to be more surgical in our injection of failure and tie that into our telemetry solutions.
We only inject failures we expect to be resilient to. Sadly, that is a subset of the failures that people cause ;)
The bigger your application stack is (micro-services, API calls, network calls), the more failure you would need to test out and there's no way to "trust" developers to do it themselves.
Also, Netflix hires a lot of top-notch developers and their infrastructure is pretty awesome.
[I don't work @ Netflix, just a devops dude :)]
Seems unfortunate that it requires the coupling with spinnaker - although i can see how it helps with the cluster definition features.
Edit: I'll add that we've been using the original chaos monkey and chaos lambda extensively in production for some time with very few problems.
"We rewrote the service for improved maintainability" seems an important part of this blog post.
Source: I'm on the Chaos team here at Netflix.
Is it good for the organization? Yes. Good for the guy pushing it? Very possibly very not.
Go has race detector mode, but it is an optional debug feature with a performance cost. The Linux kernel's jiffy clock starts counting from -5 minutes so drivers must handle clock rollover correctly because it's not a uncommon "once every 48 days" event. Firefox has a chaos debug mode that does things like randomize thread priorities and simulate short socket reads, but that has performance costs.
runtime: hashmap iterator start position not random enough #8688
I can imagine it would be a tough sell to the CEO. :)
You catch bugs, and no one says you can't run Chaos Monkey in staging or a similar environment if it really is a tough sell.
The biggest issue IMO is explaining need to make things more resilient. Actually the technical people (mainly developers) might be the biggest obstacle, because it adds more work for them (with no visible benefit to them, because when application fails it's ops who get woken up).
I remember reading years and years ago about bandit algorithms... this kind of ops work is at a level that's found only in a few different companies.
The "sell" can be tricky for some people, until your first production issue. Never let a crisis go to waste.
Machines go away all the time in the cloud. This tool increases the frequency so you can ensure your system handles it gracefully.
Some people believe their system can tolerate this class of failures, but without continuous validation, that is more of a hope than a certainty.
See https://medium.com/production-ready/chaos-monkey-for-fun-and...
As for chaos being a hard sell: https://medium.com/production-ready/chaos-engineering-a-shif...
Why? Your customers use your production environment, not your test environment. Something will cause loss of an instance for you: * Mistaken termination * AWS retirement (and you missed the email) * Cable trip in the data center * <Something else we can come up with> * <This list goes on>
So, vaccinate against the loss of an instance cratering your service. Give your prod environment a booster shot (with Chaos Monkey or something like it) every hour of every day. Then, when anything from the above list happens you're infrastructure handles it gracefully and without intervention. Continued booster shots ensure that this stability continues through config changes, software version changes, OS changes, tooling changes, etc.
I think the better question is "Why wouldn't you do this?"
Whoa, that's meta.
Although designed originally to catch places where malloc failure wasn't being handled, it can also be used to randomly trigger other off-nominal portions of the code that might not otherwise be tested.
But isn't there a danger that it also encourages maladaptions that come to rely on being regularly restarted by the Chaos Monkey? I'm particularly thinking that you might evolve a lot of resource leaks that go unnoticed so long as Chaos Monkey is on the job.
The connection to techblog.netflix.com was interrupted while the page was loading.
C:\Windows\System32>nslookup http://techblog.netflix.com/ 8.8.8.8
Server: google-public-dns-a.google.com
Address: 8.8.8.8
* google-public-dns-a.google.com can't find http://techblog.netflix.com/: Non-existent domain
nslookup techblog.netflix.com 8.8.8.8
If one random guy's complaining about a URL being unreachable and you're already seeing a pattern... is it at all possible the users aren't at fault?
https://news.ycombinator.com/item?id=12269411
https://news.ycombinator.com/item?id=12217900