Amazon DNS error
amazon.com
amazon.com
I love AWS but they really need to improve their procedures for communicating during outages. If your parent companies' billion dollar site can be affected in any way on the night before black friday, and even your own site is down, and you are not acknowledging that the service is FUBAR - you have a problem.
We've built our systems to recover from failure states once we know there is a problem. AWS's inability to do that reliably is forcing us to own the problem ourself, and as a result, we will probably migrate away to use cheaper boxes in the cloud.
On the positive side, RDS has been solid.
--
Cloudfront DNS is hosed. Doing a DNS lookup on my cloudfront distribution fails. Many folks on Twitter also see the issue too [1]. Maybe a failed upgrade or something, their whois info was updated today @ 2014-11-26T16:24:49-0800.
[~]$ host d1cg27r99kkbpq.cloudfront.net
Host d1cg27r99kkbpq.cloudfront.net not found: 2(SERVFAIL)
[1] https://twitter.com/search?q=cloudfrontNot that I've noticed they tend to lie through their teeth on the status page or anything....
<title type="text">Informational message: DNS Resolution errors </title>
<link>http://status.aws.amazon.com</link>
<pubDate>Wed, 26 Nov 2014 17:00:39 PST</pubDate>
<guid>http://status.aws.amazon.com/#cloudfront_1417050039</guid>
<description>We are currently investigating increased error rates for DNS queries for CloudFront distributions. </description>However, digging the Cloudfront name servers times out intermittently:
$ dig +short @ns-666.awsdns-19.net cloudfront.net
;; connection timed out; no servers could be reached- All of Vox Media's properties (The Verge, Polygon, Vox.com, SBNation, etc)
- All of Atlasssian's services (Bitbucket, Jira OnDemand, etc)
- Flowdock
- aws.amazon.com has no assets
Edit: console.aws.amazon.com has no assets, either, so it's also currently worthless.
It's probably worth having a DNS failover strategy for Route53 (if that's what you're using) that doesn't involve the UI on console.aws.amazon.com.
Which is one of the reasons I setup https://dns-api.com/ - A way of updating Route53 DNS via git hooks.
Columbus:
[11-26-2014 16:46:24] SERVICE ALERT: public-www;CDN - Logo;OK;SOFT;2;HTTP OK: HTTP/1.1 200 OK - 2960 bytes in 0.165 second response time
[11-26-2014 16:45:34] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
[11-26-2014 16:39:24] SERVICE ALERT: public-www;CDN - Logo;OK;SOFT;3;HTTP OK: HTTP/1.1 200 OK - 2960 bytes in 0.030 second response time
[11-26-2014 16:38:34] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;2;Name or service not known
[11-26-2014 16:37:34] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
[11-26-2014 16:25:24] SERVICE ALERT: public-www;CDN - Logo;OK;SOFT;2;HTTP OK: HTTP/1.1 200 OK - 2960 bytes in 0.030 second response time
[11-26-2014 16:24:34] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
[11-26-2014 16:21:24] SERVICE ALERT: public-www;CDN - Logo;OK;SOFT;2;HTTP OK: HTTP/1.1 200 OK - 2960 bytes in 0.066 second response time
[11-26-2014 16:20:24] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
Portland:
[11-26-2014 16:49:40] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
[11-26-2014 16:43:40] SERVICE ALERT: public-www;CDN - Logo;OK;SOFT;2;HTTP OK: HTTP/1.1 200 OK - 2960 bytes in 0.148 second response time
[11-26-2014 16:42:40] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
[11-26-2014 16:39:40] SERVICE ALERT: public-www;CDN - Logo;OK;HARD;3;HTTP OK: HTTP/1.1 200 OK - 2960 bytes in 0.186 second response time
[11-26-2014 16:21:40] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;HARD;3;Name or service not known
[11-26-2014 16:20:40] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;2;Name or service not known
[11-26-2014 16:19:41] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known
Santa Clara:
[11-26-2014 16:24:26] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;HARD;3;Name or service not known
[11-26-2014 16:23:25] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;2;Name or service not known
[11-26-2014 16:22:25] SERVICE ALERT: public-www;CDN - Logo;CRITICAL;SOFT;1;Name or service not known<script> window.jQuery || document.write("<script src='js/jquery-1.10.2.min.js'>\x3C/script>")</script>
config.asset_host = -> {
cdn_up? ? "http://mycdn.com" : "http://mydomain.com"
}What would be really nice is if you could specify a fallback host in your DNS prefetch, and the browser would make it "just work."
I would hope they honor this dns issue under the same guidelines although its technically not the route53 service we are paying for.
What's everyone using to monitor external asset hosts? Is anyone dynamically switching between them, or failing back to local assets?
$ dig @ns-666.awsdns-19.net cloudfront.net
Times out for that server as well as for all of the nameservers listed for cloudfront.net.If you have an alias to a cloudfront distribution the answers aren't being provided by the Route53 servers.
You can verify this with a dig:
dig mx [your domain like domain.com] - will probably still work.
dig a [name of alias like www.domain.com] - isn't working.
The opposite might also be true: Amazon might now have reached a point in size where they can't scale further upwards without losing visibility and control of part of their hardware.
It's all speculation anyway.
Noticed: - trello.com - import.io - intercom.io
as some sites as well as ours who we've noticed issues with.
Finally, chat came through--but only after a wait far longer than what their SLA guarantees.
Point being that support probably isn't going to get you much either, especially given that they aren't holding to the 1hr SLA.
Deleted comment