Mistakes teams new to Chaos Engineering make
dadontherunblog.com
dadontherunblog.com
What this is a form of testing, and applying boring old requirements to operations. The whole idea of "chaos" is just a really sad tell about the state of planning for some places.
You should be able to articulate your reliability requirements in certain situations and the verify your stack meets them before a release. And if it's cost prohibitive to do that in a non-prod environment then planning around testing in production in your pre-release planning.
It's silly, it's like classifying a test engineer as an "exploratory tester", which would clearly be a mistake. That is just a type of activity of an engineer, not an engineer role. This is just exploratory load and scalability testing, and falls under the responsibility of a test or devops team.
This is in case you didn't quite articulate or verify your real requirements. The more confident you are that your test environment matches your prod environment, the less you need this.
> The more confident you are that your test environment matches your prod environment, the less you need this
I'd also add "workload" alongside "environment", which is often challenging to accurately simulate.
Call me old fashioned, but I don't like messing around with production. I think you can reasonably scale down production environments and simulated wan networks packet loss/latency. I can't imagine a platform moderate importance not having some sort of benchmarking and scalability as part of their pre-prod pipeline automated acceptance criteria. It feels like this work would fall under that engineering role (again, performance test engineer or devops/sre).
Anyway, I think I chafe at the name more than anything. I know nonlinear dymanics, stress testing, queuing theory, et al and it just feels overly glib for what should be the most serious sort of activity a company considers.
"What would happen if..." experimentation doesn't negate the need for requirements. It seeks to actively test whether or not those requirements were adequate in the first place. It also recognizes that no system is static - almost all distributed systems are evolved over long periods of time.
No matter how careful your planning, shit sometimes breaks in the real world in ways you didn't/could not have reasonably anticipated.
"A list of the mistakes that teams make if they are new to the concept of Chaos Engineering"