Let's Make DNS Outage Suck Less
kvz.io
kvz.io
echo "nameserver 127.0.0.1" > /etc/resolv.conf
http://www.thekelleys.org.uk/dnsmasq/docs/dnsmasq-man.html
--all-servers : By default, when dnsmasq has more than one upstream server available, it will send queries to just one server. Setting this flag forces dnsmasq to send all queries to all available servers. The reply from the server which answers first will be returned to the original requester.
Dnsmasq also has avoidance of unresponsive servers built in, this is a bit more of a blunt instrument.
Really wanted to solve this using as transparent & archaic tech as possible in the hope this still stands when big player's resolving servers fall.
Granted, it is blunt; archaic tools often are ; )
echo "nameserver 127.0.0.1" > /etc/resolv.conf
If you use dhcp3-client (chances are you do), add this line to /etc/dhcp/dhcplient.conf to try your local DNS cache first: prepend domain-name-servers 127.0.0.1;
Note that this will slow down stuff like captive portals at airports that want to push you to their "accept terms" page first.I use djb's dnscache instead of dnsmasq, but dnsmasq works fine too.
I'm sitting here contrasting this against my normal approach* and the approach here is:
1) easier to explain to others
2) self documenting
3) just as effective as running a caching nameserver
* i have vm hosts configured in CM to have a bunch of promises applied, one of them is to run a caching nameserver, these hosts are the only ones allowed to do zone transfers. The vm instances running on top of these have a promise applied which has them use their underlying dom0 for dns queries.Do you have data you can share regarding "Amazon EC2 resolving nameserver (172.16.0.23) is unreachable too often" ?
In case of Amazon, you actually prefer their nameserver as it resolves names to LAN IPs where possible. With your strategy, ~66% of my instance hostnames would be resolved to the less efficient WAN IPs, that's with all 3 nameservers being reachable.
As for data, there's papertrail and Premium Support tickets with Amazon admitting to their downtimes. In most cases (e.g. the one on Jan 14, 2013 11:27 PM PST) "there was a problem with the underlying host".
Remember it doesn't even need to be the service itself going down, there's weaker links in this chain like misconfiguration of the virtual host. Cables can get tripped over. Anycast is not possible here.
I'm not blaming them. This article was about mitigating suckyness in case of outages, not proving that such a thing even exists : )
Agreed on aws I would use their resolver to request the internal ip. Another option is setting up your environment in VPC. With VPC you'll be able to assign a static internal ip to your instance.
btw I did find many reports on the 172.16.0.23 lookup issues.
Good suggestions otherwise. I'm not sure what will happen with an RDS Multi-AZ failover if you address it by it's static VPC IP?
Also about your worst case scenario. It's true that I rely on cron to be working. While other proposed solutions rely on more obscure daemons to be up & running, there still is a risk, and it could be mitigated by just writing all the nameservers to the resolv.conf. I've created an issue for that: https://github.com/kvz/nsfailover/issues/1