The Amazon S3 team recently completed some maintenance
changes to Amazon S3’s DNS configuration for the US STANDARD region on
July 30th, 2015.
You are receiving this email because we noticed that your bucket
is still receiving requests on the IP addresses which were removed
from DNS rotation. These IP addresses will be disabled on August
10th at 11:00 am PDT, at which time any requests still using
those addresses will receive an HTTP 503 response status code.
Applications should use the published Amazon S3 DNS names for
US STANDARD: either s3.amazonaws.com or s3-external-2.amazonaws.com
with their associated time to live (TTL) values. Please refer to
our documentation at:
http://docs.aws.amazon.com/general/latest/gr/rande.html#s3_region
for more information on Amazon S3 DNS names.
Something to do with that perhaps? AWS sent us that last thursdayhttp://javaeesupportpatterns.blogspot.ie/2011/03/java-dns-ca... has more detail.
Companies 1/1000th the size of Amazon can manage it.
But the last thing you want to do is put inaccurate information onto a status page; so mere administrative personnel isn't enough -- you'd need people who understand enough about the system to be able to write about it without introducing errors.
I'm guessing that the intersection of "administrative personnel", "willing to carry pagers" and "understand the internals of AWS services" is a very small set.
They also claim to have "customer obsession" as a leadership principle, this whole thread is an excellent example of that being failed in a big way.
Not every job can be full of self-directed aspirational spiritual awakenings. If that were the case, nobody would deliver my dinner on a bike when it's -20ºF outside.
Being a non-engineer doesn't mean they don't know anything about the technology. And they don't need to know the internals, just enough to convey information from the engineers managers to the public.
Plenty of other organizations manage resolving issues while transmitting information about the issue to other stakeholders.
Also, most administrative personnel have far less job opportunities than engineers. If they can get the engineers to carry pagers they can get a PR minion to carry one.
1:52 AM PDT We are actively working on the recovery process, focusing on multiple steps in parallel. While we are in recovery, customers will continue to see elevated error rate and latencies.
or would the inter-machine chatter continue ad infinitum? would they run out of IPs or successfully transition to IPv6?
So many questions.
http://craphound.com/overclocked/Cory_Doctorow_-_Overclocked...
2:38 AM PDT We continue to execute on our recovery plan
and have taken multiple steps to reduce latencies and error
rates for Amazon S3 in US-STANDARD. Customers may continue
to experience elevated latencies and error rates as we
proceed through our recovery plan.Can't pull any images either.
Unable to fetch source from: https://s3-external-1.amazonaws.com/heroku-sources-production/heroku.com/<some-uuid>?AWSAccessKeyId=<some-access-key>&Signature=<some-signature>&Expires=1439198046looks like the cavalry are coming
Feels like in future... If cloud provider goes down... All internet will stop working :)
[3:25] AM PDT We are still working through our recovery plan.
Man, I'd love to see that plan. $ vagrant up
Bringing machine '...' up with 'virtualbox' provider... ==> ...: Box 'debian/jessie64' could not be found.
...
...: Downloading: https://atlas.hashicorp.com/debian/boxes/jessie64/versions/8.1.0/providers/virtualbox.box
An error occurred while downloading the remote file. The error message, if any, is reproduced below. Please fix this error and try again.
The requested URL returned error: 500 Internal Server Error
EDIT: not Markdown.3:46 AM PDT Between 12:08 AM and 3:40 AM PDT, Amazon S3 experienced elevated error rates and latencies. We identified the root cause and pursued multiple paths to recovery. The error has been corrected and the service is operating normally.
There are many use-case when paying 2x for storage is a reasonable tradeoff for higher availability and also be provider independent.
Importantly, make sure you CNAME your bucket under your own domain so that you can switch services.
edit: Much easier than AWS Lambda, actually: http://aws.amazon.com/about-aws/whats-new/2015/03/amazon-s3-...
Apache jclouds® is an open source multi-cloud toolkit for the Java platform that gives you the freedom to create applications that are portable across clouds while giving you full control to use cloud-specific features.
So far S3 seems to be reliable enough...
12:36 AM PDT We are investigating elevated errors for requests made to Amazon S3 in the US-STANDARD Region.
S3 is in yellow, which means "performance issues". But not being able to download files from many buckets it's clearly a "service disruption" (red).
> 12:51 AM PDT We are investigating increased error rates for the EC2 APIs and launch failures for new EC2 instances in the US-EAST-1 Region.
edit; does not seem to affect all the buckets though. Only one of ours is experiencing this, others are fine.
3:36 AM PDT Customers should start to see declines in elevated errors and
latencies in the Amazon S3 service.
Fixed ?3am "Dry Run" Staging on Heroku...fail 7am Deployment to Production....fail
Now: "We have confirmed elevated latencies affecting our SendEmail, SendRawEmail and SendSMTPEmail APIs in the US-EAST-1 Region and are working to address the problem."
Which is perfect since most of our PO orders are being placed between 6am-10am.
Tonight, adult beverages will be needed after everything is resolved.
12:51 AM PDT We are investigating increased error rates for the EC2 APIs and launch failures for new EC2 instances in the US-EAST-1 Region.
Elastic Compute Cloud
Elastic Load Balancing
Elastic MapReduce
Relational Database Service
CloudTrail
Config
Lambda
OpsWorks
good luck to the on-call engineers at amazon!
Amazon Elastic Compute Cloud (N. Virginia) Increased API Error Rates 12:51 AM PDT We are investigating increased error rates for the EC2 APIs and launch failures for new EC2 instances in the US-EAST-1 Region.
Whomever being a DevOps or SysAdmin probably cannot sleep tonight :(.
In our case, we put Fastly on top of our assets/images so that only a partil of request get errors. The cached object on Fastly is still fine.
However, if it doesn't come back before 5AM I think we are screw :(