The EC2 has responded, and informed me that they do not provide analysis of individual issues in AWS infrastructure.
They referred to the AWS SLA that states that "AWS will use commercially reasonable efforts to make Amazon EC2 and Amazon EBS each available with a Monthly Uptime Percentage (defined below) of at least 99.95%".
The AWS SLA is available here:
https://aws.amazon.com/ec2/sla/
(The above is a part of the response I received when I attempted to ask AWS about sporadic reboots and outages on one of my instances.) Thanks for that information. It really help us understand the issue faced. As you know changing a A record for a domain can take time to replicate through all the Root DNS servers and then onto the none authoritative servers from there on. When you are editing your DNS Zone file it is highly related to the TTL settings for quicker updates on changes to them.
However it really isn't uncommon to see full DNS replication when making changes to an A record.
..
Also for even more control and perhaps faster DNS resolving times within AWS look into our Route53 service.
https://aws.amazon.com/route53/
I hope this has addressed your questions and please feel free to ask if there is anything else.
Tools like http://mxtoolbox.com/dnspropagation.aspx and http://www.viewdns.info/propagation/ indicated that the change had been picked up by every listed server but for AWS.Solution proposed by AWS: Use Route53.
And don't forget the CPU throttling, noisy neighbors, IOPS and what not.
When things don't work, AWS doesn't tell you a reason why it didn't work. Instead they teach you lessons about how to architect your applications for failure by saying:
The EC2 team also recommend that you architect for failure using the white paper linked below:
https://d0.awsstatic.com/whitepapers/AWS_Cloud_Best_Practices.pdf
See my earlier comment as well: https://news.ycombinator.com/item?id=11822298For people who use it as a VPS provider, they are going to be hard hit by other tenants who are applying the intended purpose.
There are a ton of these teams. AFT (several hundred devs), for example, is practically on their 10th year of a 3 year mandate to get off of their monolithic Oracle database backends, a constraint which forces them to run in legacy datacenters. You'll find similar issues with plenty of Supply Chain, Operations, and Transportation teams as well. The primary data warehouse clusters (non-redshift) and management interfaces are also in legacy data centers, last I recall.