The Discipline of Chaos Engineering
blog.gremlininc.com
blog.gremlininc.com
Right - this makes sense. I thought this philosophy sounded Netflix-ish. Interesting to see it "spun out" into a product.
I don't like how they define "Chaos Engineering" as being strictly related to distributed systems.
I would consider applying invalid inputs, etc. to non-distributed systems to be more along the lines of traditional testing. Perhaps you could implement chaos engineering principles in a non-distributed system by simulating the failure of a CPU core or a region of memory? It seems less useful though, as those things seem very difficult to effectively recover from.
How would you define "chaos engineering" to apply to non-distributed systems?
Could be. In a general sense of chaos engineering.
For me, I think chaos engineering would be to keep generating "chaos" in a system or module or any unit. Chaos in this sense would be to break it, or any kind of maltreatment to it.
As a matter of fact, Netflix is running a big distributed system, so that's where they focus their testing efforts. In general, I think it's fair to talk about Chaos Engineering and systems in the general sense, distributed or not.
About the database : Cache. Offer a different flow that does not need the database, like with hardcoded stuff. Etc etc
About doubling the costs : It highly depends. On a small app, sure it may double the cost. On a big thing with thousands of servers, you can play a bit more with redistributing roles and all.
For your specific question about databases, you generally have clustering to reduce the impact of any one database instance going down, caching of data to guard against temporary db outages/network issues, and sharding of data across multiple databases to reduce the blast radius of any one logical database going away entirely.
2. Netflix has done a great job at publicizing their efforts and open sourcing software that helps you do this kind of testing in a continuous, automated fashion.
So I think the "basically what google has been doing" comment is reductive.
It is true that Google's Disaster Recovery Testing events are also about breaking things on purpose as a means of preparation. However, those events are typically large-scale, company-wide drills targeting not only critical systems but also business processes involving people.
(They even prevent experts from participating to make sure knowledge is spread across the organization. I recommend reading http://queue.acm.org/detail.cfm?id=2371516 for more.)
As dastbe has pointed out, Chaos Engineering is more about experimenting in a continuous, automated (and hopefully safe) way. Compared to DiRT, experiments are typically smaller in scope, involving fewer people, if any.