DNSSEC Outage on www.cloudflare.com – 2019-03-21
ianix.com
ianix.com
Specifically, we were breaking out our "www" subdomain as a separately managed zone, and some of the configuration wasn't properly migrated.
This resolution failure was regrettable, but there was no impact to the service we provide for our customers.
• You get the default certificate for the web server rather than the correct one for the vhost.
• You have a redirect from www->bare or bare->www and the redirect vhost don't have a certificate.
• The certificate don't include the www subdomain name.
• A wild card certificate doesn't include the bare name.
• There are multiple web servers listed in DNS but only a subset has a correct HTTPS configuration.
I for myself pay for a smaller DNS provider like CloudDNS with a bigger redundancy in NS servers (4 instead of 2 at Cloudflare) than going for free at Cloudflare.
I don't think this is really all that important, as typically all these servers are running anycast. If you really wanted redundancy, you would have NS records from at least two different providers.
Someone could probably build a better DNS provider than Cloudflare, at least for their specific site, but it would be difficult. I don't know of any other DNS provider I'd consider more robust, particularly against high load, DoS, etc.
(Disclaimer: worked for Cloudflare 2014-2016)
Are your saying you have inside knowledge and that's the architecture? (I'm guessing no)
If you use juniper routers, for example, a juniper specific software, hardware problem or a mistaken assumption about juniper specifics could take down all routers and an entire network.
You may be onto something that nobody else spotted but so far the arguments are too generic (“failure”) or incompatible with the premise (rules of natural evolution applied to very artificial systems).
Don’t get me wrong, I’m happy to learn something new and understand a different point of view but so far the arguments don’t help me.
Often the only thing preventing mass upgrades from becoming mass outages is some part of the network just happens to work differently.
I was actually hoping for an interesting human independent failure. A cosmic radiation caused bit flip that would occur on a particular device. :) Human error can overcome any setup.
Where ever you host something, the provider may make mistakes which causes you to be off the air for some time. It doesn't matter if a big provider makes a mistake and takes down lots of domains or if lots of smaller companies make their own mistakes and over time (together) also take down a lot of domains.
If, as a user of the internet, you only care if there are cat video somewhere, then yes, having lots of small providers makes a difference. If you just need to visit one website, it doesn't matter how many other sites are down.
The Dyn incident is a different story. Because there a (ddos) attack on one site, took down lots of other sites. In that case using smaller providers would have helped. But one of the features cloudflare is selling is ddos resistance. Which is something only very big players can offer.
Now I'm curious as to why www has NS records...
(DS records are used to tie DNSKEYs to a child zone from the parent zone).