Principles of Chaos Engineering
principlesofchaos.org
principlesofchaos.org
https://github.com/Netflix/chaosmonkey
https://medium.com/netflix-techblog/chaos-engineering-upgrad...
http://www.oreilly.com/webops-perf/free/chaos-engineering.cs...
It reminds me of the Amazon "everything must be a network service" architecture. Painful to institute but with great payoffs later in terms of robustness and scalability.
I see this as a version of Monte Carlo testing - I'm not totally sold on doing it in production, but certainly if you have a load that's difficult to replicate in scale, then this is the type of thing you should be doing.
But I don't think anyone should be doing this from the start. This is something that you evolve into. Sure, it's easier to bring it in if you planned it from the start - but as an engineer, you cannot proceed on that basis. You know and understand stuff has to change as you grow, and the art of it is making (more often than not) good calls on architecture to allow that.
Because we had already been dealing with chaos monkey, everything recovered by itself within a few minutes of the script stopping, and we barely had any alerts from any services going down.
Also, when aws was pushing our meltdown patches, we just ignored all the notifications because we didn’t care about instances being shut down, because they go down randomly all the time.
It’s a pain in the ass at first to deal with, but it definitely forces you to fix a lot of potential problems unless you want to spend all your time manually fixing stuff every day.
Even just aside from chaos monkey, we build new base amis and push out rolling updates daily. That’ll also quickly suss out problems with not having version numbers locked, etc.
For instance, you may want to check that the authorization service can function without access to the A/B testing service, so you cut the network connection between them. If authoritazion errors start rising, you stop the experiment and investigate why this happened. You may also find that clients did not retry properly on unexpected errors from the authorization service (e.g. http 500 error codes)
If you have otherwise upheld your targets, you can use the remaining slack at the end of the period to run larger experiments such as these, for a long term benefit.
Basically you are creating a fault at a time of your choosing to prevent the same fault from occurring at the worest possible time.
improper fallback settings when a service is unavailable;
retry storms from improperly tuned timeouts; outages when a
downstream dependency receives too much traffic; cascading
failures when a single point of failure crashes;
I don't see how one couldn't replicate those in an environment other than the production. They all involve bringing down complete services and sending some unexpected high load to other services.Kyle Kingsbury == Aphyr
Also, I take issue with it being “destructive”. Compute resources are effectively free on the margin (at least at the scale that concerns us here) and ephemeral. Killing such a process doesn’t meaningfully destroy anything of value.
Are they using the formal definition of chaos, or just a colloquialism?
The engineering comes in when we try to simulate such situations and create processes for hardening applications against those types of failure.
> Are they using the formal definition of chaos, or just a colloquialism?
Is it a colloquialism to use a word in its normal English sense? Chaos as a word precedes Chaos Theory by some thousands of years. If I write a theory and take over an existing word in everyday use, it seems a bit much to accuse every one else of colloquiallism when they use an existing but less strict definition?
It does, but engineering is a technical profession and it's practitioners are likely familiar with the mathematical concept.
I've read about applications of chaos theory in system design, and I expected 'Principals of Chaos Engineering' to be about that topic.
They’re just trying to co-opt “chaos” as a buzzword for fault-tolerant systems.