Principles of Chaos Engineering (2018)
principlesofchaos.org
principlesofchaos.org
Funnily enough it also improved our latencies a lot (which I guess is mostly due to memory leaks et al.)
Of course this was a decade ago, but I think the fundamentals are still sound, as far as being skeptical about the quality and longevity of your nodes in a virtual environment.
https://www.pingcap.com/blog/chaos-practice-in-tidb/
Regarding the Fault Injection tools we are using:
- Kernel Fault Injection, the Fault Injection Framework included in Linux kernel, you can use to implement simple fault injections to test device drivers.
- SystemTap, a scripting language and tool diagnose of a performance or functional problem.
- Fail, gofail for go and fail-rs for Rust
- Namazu: a programmable fuzzy scheduler to test a distributed system.
We also built our own Automatic Chaos platform, Schrodinger, to automate all these tests to improve both efficiency and coverage
- there's a service that needs config data from a DB on another node to initialize itself to become useful - should the service die if it doesn't have connection to the DB on startup (so that the error propagates), or should it start and perform retry indefinitely until DB connection is set? Until that happens it sends back error code to its consumers.
identifying that it needs to do something, and whether it does it or not is part of chaos engineering. eg by turning off the DB for a bit and seeing what it does
- Chaos Monkey Guide for Engineers https://www.gremlin.com/chaos-monkey/
- Recent HN discussion on Resilience Engineering: Where do I start? https://news.ycombinator.com/item?id=19898645
It seems like this setup works great if built from the get-go but incredibly painful and possibly dangerous if starting with existing applications.
It's basically the strangler pattern[1]. It is painful, but can be made arbitrarily safe.
Also the term 'antifragile' (lightly controversial) comes to mind.