I assume so but would be nice to know its working as expected!
I assume so but would be nice to know its working as expected!
We have mostly found one or two classes of jobs in the orchestration for which nomad stopped retrying deployments before the ec2 instances running the allocations were fully failed and removed from the cluster - and our on-call was unsure how to handle that situation in nomad right. Network and routing were really weird at some point. Additionally, we ended up with a couple of container instances orphaned from the container management, which was strange for a moment.
This was made a bit more hectic over here because a second hoster apparently fried their own network at the same time so we needed some time to realize we have two issues.
Overall, 5/5 Outage, would fail again once we've updated our jobs. We're happily close to not caring about such an incident.
Amazon had released a version of their AWS Linux edition that rebooted randomly due to a kernel bug, and I was working on our staging cluster, but I didn't even notice that I had EC2 instances that randomly rebooted and dropped because Nomad just kept the workload up.
"To ensure that resources are distributed across the Availability Zones for a Region, we independently map Availability Zones to names for each account.".
my eu-central-1a is not necessarily yours.
[0] https://docs.aws.amazon.com/ram/latest/userguide/working-wit...
[1]: https://docs.aws.amazon.com/ram/latest/userguide/working-wit...
I wonder whether its because a large load of capacity needs suddenly in unaffected AZs put abnormal stress on things like networking to the whole unaffected AZs.
My only experience there is in Azure, when they deployed patches for the Heartbleed etc issues, for a few weeks things were much slower in CPU power (our response times just shot up 20% for no reason then recovered a few weeks later) and there were network related timeouts that were abnormal and it all settled down eventually.
For reference https://heartbleed.com/
We also do have some EC2s running with nodejs applications and there the aws-sdk just errored out with "UnknownError: 503" and simply stopped logging until we restarted the machines. The machines itself were not stopped at all.
Other than that, I can't see any effects across our accounts. Also not RDS or anything else. Fascinating. Glad its under control and seemingly nobody died or so.