Amazon EC2 Issues in Frankfurt AZ
status.aws.amazon.com
status.aws.amazon.com
We have the ability to do everything and anything in Javascript EXCEPT for fucking UTC or localized time zones.
Is PST the new UTC? Is there an RFC for this? Did I miss the memo?
Un-fucking-real.
Doable, of course, but it is not just a matter of putting it into the issue description.
"We are experiencing elevated API error rates and network connectivity errors in a single Availability Zone."
Key fact "A SINGLE AZ". Availability Zone's are isolated from each other with redundant power supplies and internet connectivity and most often physically different datacenter locations. Well architected applications are designed to allow for a single AZ to become unavailable. This is precisely why the cloud is useful: you can bring up new capacity in the other Availability Zones behind the same load balancer with zero effort - the autoscaling does that for you automatically.
https://pastebin.com/WC1hkh0c (see ts on uptime command, ignore timezone differences)
>The beautiful thing here is that the failovers work to spec
Who's failover? AWS' certainly did not for us.
Also, I don't know how AWS implements failovers internally, but I can see several potential issues that could arise from connectivity loss between availability zones, especially if both EC2 and RDS are involved – few of which are trivial to handle if your application cannot tolerate loss of consistency.
Are you referring to this kind of scenario? https://github.blog/2018-10-30-oct21-post-incident-analysis/
Yes, after skimming your referenced article, this is exactly the type of issue that prevents a "simple" failover when databases are involved.
Depending on your application, even returning stale data on one AZ in a split brain scenario might not be acceptable (this is a stronger requirement than not allowing non-reconcilable conflicting writes).
RDS lets you choose a DB Engine (MySQL, Postgres, etc). the replication depends on the engine you choose, e.g. Postgres native streaming. Heres a good writup on how it works and best practices (https://aws.amazon.com/blogs/database/best-practices-for-ama...)
This is creating a huge number of issues as you lose capacity with the AZ going down and can't recover by creating new instances.
How were you launching? Using EC2 directly (e.g. `aws ec2` cli or RunInstances API call - https://docs.aws.amazon.com/AWSEC2/latest/APIReference/API_R... ), or using e.g. ASGs?
If you were launching directly with EC2, were you doing targeted launches (e.g. specifying Placement.AvailabilityZone for classic or SubnetId for VPC)?
In the case of an AZ failure, launches which don't specify an AvailabilityZone or SubnetId may be impacted for a few minutes, whereas launches which do specify them should still succeed.
Not a big deal to remediate as all the other AZs are working.
https://twitter.com/TransferWise/status/1194168200210124800?...
12:08 AM PST We are investigating increased network connectivity errors to instances in a single Availability Zone in the EU-CENTRAL-1 Region.[edit] I know it can be done with combined metrics etc, but it would make it a lot more complicated ;-)