In addition you can reach out to our customer support team at support@pagerduty.com or +1 (844) 700-3889.
Tim Armandpour, SVP of Product Development, PagerDuty
In addition you can reach out to our customer support team at support@pagerduty.com or +1 (844) 700-3889.
Tim Armandpour, SVP of Product Development, PagerDuty
Specifically, they shouldn't have all of their DNS hosted with one company. That is a major design flaw for a disaster-handling tool.
I ask because my secondary question, as a network noob, is was anybody prepared / preparing for a DDOS on a DNS like this? Were people talking about this before? I live in Mountain View so I've been thinking today about the steps I and my company could take in case something horrifying happens - I remember reading on reddit years ago about local internets, wifi nets, etc, and would love to start building some fail safes with this in mind.
Two pronged comment, sorry.
Re question #1, check out PagerDuty's reliability page here: https://www.pagerduty.com/features/always-on-reliability/
Namely "Uninterrupted Service at Scale - Our service is distributed across multiple data centers and hosting providers, so that if one goes down, we stay available."
It seems fair to expect them to have a backup dns too, but I am not an expert.
Yes.
I have, personally, been under attack with as-large or larger than todays attacks at my DNS infrastructure and survived.
Knocking half of the web off the grid because their DNS provider is under attack? It happened recently to DNSimple.
https://blog.dnsimple.com/2014/12/incident-report-ddos/
The irony is that I noticed it when dotnetrocks.com went offline, at that time dotnetrocks was sponsored by dnsimple...
Even a single quad-core server with 4GB RAM running TinyDNS could serve 10K queries per second, based on extrapolation and assumed improvements since this 2001 test, which showed nearly 4K/second performance on 700Mhz PIII CPUs: https://lists.isc.org/pipermail/bind-users/2001-June/029457....
EDIT to add: and lengthening TTLs temporarily would mean that those 10K queries would quickly lessen the outage, since each query might last for 12 hours; and large ISPs like Comcast would cache the queries for all their customers, so a single successful query delivered to Comcast would have (some amount) of multiplier effect.
See MaxCDN for example who uses a mix of dns providers (AWS Route53 and NS1):
ns-5.awsdns-00.com. ['205.251.192.5'] [TTL=172800]
ns-926.awsdns-51.net. ['205.251.195.158'] [TTL=172800]
ns-1762.awsdns-28.co.uk. ['205.251.198.226'] (NO GLUE) [TTL=172800]
ns-1295.awsdns-33.org. ['205.251.197.15'] (NO GLUE) [TTL=172800]
dns1.p03.nsone.net. ['198.51.44.3'] [TTL=172800]
dns2.p03.nsone.net. ['198.51.45.3'] [TTL=172800]
dns3.p03.nsone.net. ['198.51.44.67'] [TTL=172800]
dns4.p03.nsone.net. ['198.51.45.67'] [TTL=172800]
Curious, are you the kind of person that runs their own smtp email server and complains about GitHub pricing being too expensive?If both Dyn and R53 go down, it's exactly when you want a service like PagerDuty work without a hitch.
Spoiler: I'd bet my complete net worth against your assertion and give you incredible odds.
Golden rule: Fixing a DNS outage with actions that require DNS propagation = game over. You'd might as well hop in the car and start driving your content to people's homes.
I was giving a bare-minimum example of how this or (some other backup solution) should have already been setup and ready to be switched over.
DNS is bog-simple to serve and secure (provided you don't try to do the fancier stuff and just serve DNS records): it is basically like serving static HTML in terms of difficulty.
That a company would have a backup of all important sites/IP addresses locally available and ready to deploy on some other service, or even be built by hand via some quickly rented servers, is I think quite a reasonable thing to have. I guess it would also be simple to run on GCE and Azure as well, if you don't like the idea of dedicated servers.
I just wish it scaled to multiple cores :(
Calling it a "challenge" implies that there is some difficult, but possible, action that the customer could take to resolve the issue. Since that is not the case, this means either you don't understand what's going on, or you're subtly mocking your customers inadvertently.
Try less to make things sound nice and MBAish, and try more to just communicate honestly and directly using simple language.