Together with a short TTL we were able to recover without relying on their dashboards.
I'm running the web servers, official wiki, and game external resource portal for the most active open source video game on github, through cloudflare, and maybe we might not want our 60 million requests a month website to go down when cloudflare does.
Because I can tell you right now our 300 a month budget (that mind you, is capable of covering 7 game servers that can handle 100 connected players (each)) can't take the 80 dollar hit just to make cloudflare not a single point of failure.
Edit: Wait, do they use Cloudflare?
$ dig namecheap.com +short
198.54.117.250
whois: CIDR: 198.54.112.0/20
NetName: NAMEC-4
Organization: Namecheap, Inc. (NAMEC-4)
Updated: 2015-11-13
https://whois.arin.net/rest/net/NET-198-54-112-0-1.html $ kdig +short www.namecheap.com
www.namecheap.com.cdn.cloudflare.net.
104.16.99.56
104.16.100.56If this is something you want to be able to mitigate, you really need to be running a seperate DNS infra from your hosting/CDN and use short TTL cnames to delegate hostnames to the CDN. This becomes a big challenge if you host on an apex domain (eg example.org instead of www.example.org), so don't do that.
We wrote about a strategy to circumvent this sort of thing a little while back https://www.ably.io/blog/routing-around-single-point-of-fail.... Given two incidents in a matter of weeks, I think a revisit of that article in light of most businesses who operate on a single domain would be useful :)