AWS Fault Injection Simulator
aws.amazon.com
aws.amazon.com
I always considered the simplest form of chaos engineering to be a process that randomly kills processes on a server, and see what happens.
If, on the other hand, you mean that it’s important that the “chaos” aspects needs and on/off button and that someone needs to be managing that, then I agree. :)
The chaos itself should be random, but the time and place should not be.
New features are generally ported quite fast to the other regions, so this doesn't seem like a valid reason in 2020.
The proper way to handle this is using multi-region HA/DR.
What I don’t understand, however, is what’s preventing Amazon from upgrading the infrastructure in us-east-1 and making it more stable. Is it about mitigating risk of unknown side effects and/or incompatibilities?
[1] https://github.com/amzn/awsssmchaosrunner
[2] https://aws.amazon.com/blogs/opensource/building-resilient-s...
If you run the agent/daemon on your production stack, then it's a potential vector for misconfiguration or attack. But if you don't run the agent/daemon in production, then it's another way in which your test stack diverges from production!
I saw various PR/FAQs related to Chaos engineering while I worked in both EC2 and the AWS developer tools org. I've been gone over a year now, but I would bet that FIS does something at the EC2 Network level so that you don't have to install stuff on your instances or containers.
Hire me to do a half-assed job so your infrastructure can randomly fail to make it more resilient.
Many AWS outages result in APIs failing or returning errors, which should be possible to simulate. I'd like to create experiments where EC2 instances won't spin up in a particular AZ, EBS disks can't be reattached, Route53 zones can't be updated, IAM changes don't take effect, etc.
But AWS made it more configurable, and not in a single region.
A truly cunning pricing scheme would be to strike a Coasian Bargain: "we inject faults for free! If you want us to stop, you'll need to pay extra". That would mean the cost would fall mostly on those who are unable to avoid paying it.
The charitable name for this lucrative segment of the market is "Enterprise".
Our workaround was to put the IP address in the hosts files - not an ideal setup, but it got the job done.
https://en.wikipedia.org/wiki/Chaos_engineering#Chaos_Monkey
See: https://en.wikipedia.org/wiki/Jesse_Robbins#Contributions_to... https://medium.com/the-cloud-architect/chaos-engineering-ab0...
Edit: Funny how my original comment is downvoted for contradictory reasons, both because "obviously that was irrelevant to the state of the art in chaos engineering", and because "oh don't worry, that's relevant, they just discuss it in other contexts".
> but maybe Wikipedia is leaving off a lot of relevant history on the Chaos Engineering article then.
This is certainly true.
Case in point, I've been yanking out active drives, cables, PSUs and CPUs since the 1980s, and not always by accident.
What's more, I was taught to do by someone who'd been babysitting VAXes for years; they had a ceramic axe that, at the age of fourteen, I'd seen them swing right through a three-phase minicomputer power feed at my local university's data center "for test purposes".
Made it easy to insert comms errors and failures.
are their SDRs trying to boost their numbers for the year?