Stop using low DNS TTLs (2019)
blog.apnic.net
blog.apnic.net
When I worked at reddit, it took a month for traffic to shed off of the old load balancer, despite me lowering the TTL to 5 seconds. And even then there was still some traffic, so I just had to cut it off at that point. I've had similar problems when using DNS based traffic management. Sometimes when we say 50% to one IP and 50% to another, it can end up at 70/30 because of ISP and client caching.
And lastly, my friend who ran an ISP in Alaska told me that they just set all TTLs to 7 days minimum to avoid traffic back to the lower 48.
It doesn't matter what you set your DNS TTL to, so you might as well make it low for your own sake. It's a catch-22. No one plays nice with TTLs anymore because everyone sets them low, and everyone sets them low because no one plays nice.
Traffic signs shouldn't be taken down just because some people speed regardless of the posted speed limit.
Fundamentally however, giving clients control over when failover occurs means giving up control over failover.
Literally just a couple of hours ago, I was patiently explaining this to some Azure techs from Microsoft.
The conversation went like this:
> "Just use Traffic Manager for active-passive fail over!" they piped up, helpfully trying to solve my problem.
> "Client software often ignores DNS changes."
> "We designed Azure so that Traffic Manager does failover and load balancers are always active-active."
> "That's nice, but client software ignores DNS changes."
> "Umm..."
Fundamentally, the DNS protocol is just too old and crusty for the modern Internet, much like SMTP and FTP. It's best avoided to the extent possible.
The big providers seem to be converging on Anycast IP routing, where a single address can be routed to multiple data centres dynamically. You're all probably familiar with 8.8.8.8 and 1.1.1.1, both of which are Anycast IPs, but anyone can create a global static IP like that in minutes via a service such as Azure's Cross-region load balancer: https://learn.microsoft.com/en-us/azure/load-balancer/cross-...
however. Inside ISP;s routes are usually distributed by an IBGP with a couple of route reflectors, and route reflectors will drop routes if no router within the IBGP domain is advertising the route.
Routes between ISP's are usually fairly static thanks to mechanisms like graceful restart, which makes sure routes do not get dropped, even if BGP has a restart between routers. THere are various other mechanism in which to achieve this. (look up route flap dampening if you want to learn more).
So really it's processing time of every node in the chain. Few seconds to few minutes, not really that different from DNS.
The advantage is that there is no updates if there is no changes, instead of update every time TTL expires but that's about it.
Fundamentally the problem is that you want to direct a client to particular datacenter and it is not always shortest part (shortest != lowest latency) so anycast DNS isn't helping here.
There is reason biggest CDNs use it to direct to particular DC instead of " just anycasting everything" like you're proposing, it's suboptimal.
> You're all probably familiar with 8.8.8.8 and 1.1.1.1,
And both companies behind that still direct you to particular datacenter, not anycasted IP, for the actual content.
Anycast IP are also "expensive", you need to use entire /24 route to get one, and it takes routing table space in every router's memory that is limited, and requires costly replacement if it runs out everywhere, not just where you distribute it.
If you're using a DNS service that charges per query, set the TTLs high! Otherwise, 60s is probably good enough for most failover. I wouldn't go too far under 30s, unless you have a good reason; some recursives do stupid things with low values, and why tempt fate. Otoh, a lot of high profile sites and services do have < 10s ttl, so it probably works?
That doesn't really matter. Once a server becomes saturated you just remove it from being returned. Having multiple records being returned just slows down how fast a server can become saturated. You don't need the multiple records to be saturated at the same rate.
Why would you think that it doesn't matter that the traffic doesn't match the ratios I picked? I picked them for a reason.
People without cached records will always get the returned servers with the lowest load.
I was referring to load balancing and not splitting traffic for an experiment.
Then, if you return both IPs, you're very likely to eventually come across some people who sort multiple returned A/AAAA records and will prefer one that's "closer" to their current IP. Because if you're in 1.0.0.0/8, it's certainly going to be better to connect to 1.2.3.4 instead of 5.6.7.8. Same thing happens in v6 land, of course.
Lumpy results just means the times for an IP to be rotated out will be lumpy, and not that the server load will be lumpy.
>That's possible, but not always available.
If you build it, then it's always available.
There are many of DNS services which offer an API for querying and updating records.
It is very easy to write a service / script that just fines the X least loaded servers and then call an API to set those as the available records. In practice you will also want to implement some monitoring that it is actually working.
From this I have chosen a minimum of 5 minutes of TTL. Anything that might have caused a large amount of dns traffic in the early day of the internet to go off every minute, or worse ever few seconds, might killed their server and taught the young programmer to sanitize inputs.
DigitalOcean does this, or at least some of the collocation/hosting providers they are using for their infra. They appear to redirect :53 connections to their own resolvers that cache NXDOMAINs.
IMO that is a little misguided. 60-second and higher TTL will be honored most of the time. And I don't really get what that ISP in Alaska was doing: DNS is about 0.1% of Internet traffic (according to some quick googling), so decreasing that will save you only a tiny tiny bit of bandwidth at disproportionate inconvenience for your customers.
0 - https://docs.oracle.com/javase/7/docs/technotes/guides/net/p...
By setting the DNS TTL, we had established the terms by which we offered the service. Failure to accept those terms is not something we can control.
What we generally found was that clients disconnected would reconnect almost immediately to one of the other IP addresses. Any disruption to the client appeared to be almost negligible.
Likely because the clients just weren't doing DNS lookups. I would often run into clients written to resolve a hostname and then query the resolved IP for the duration of the process (often this is done as part of a library or runtime). Our customers wouldn't restart their apps for weeks or months so there was always a long tail of clients which never picked up dns changes.
DNS wasn't designed for failovers. We can use it for that, but it's never going to be perfect by design.
Because of this, IP level failovers are the way to go wherever possible. Routing is designed to handle this sort of thing in the best possible way.
But you can't just suddenly make one IP address refers to another datacenter with another ISP in another country, can you?
DNS is very simple, easy, and mostly good enough. And DNS failover is now well-enough known that we can mostly expect clients to accept fault if they don't engineer to expect it.
What you can't do with the above is more fine grain load balancing. If your LAX node is getting overloaded, you can use a combination of geo and DNS to start splitting some amount of traffic away.
That's a weird conclusion. My conclusion would be that you can't force low TTLs for everyone or even most users to you don't can't use them to allow for quick changes and may as well let everyone else benefit from caching too.
As the poster says, when it comes to the public internet DNS TTLs are more like guidelines.
Luckily most client DNS caches are cleared on a reboot.
FYI, if I remember correctly Windows (and MacOS?) used to cache negative responses for "a while".
I've seen this happen on internal networks where IT support tries to guide end users through flushing their local DNS cache so the users can get back to work before TTL expires on a broken entry.
The model I've been trying recently is to make DNS changes in three parts. First: change existing DNS entry TTL from 1hr to 1s. Wait 1hr as the caches clear. Update entry to the new address, keeping 1s TTL. Wait, test, and monitor for a while. Finally, increase the TTL back to 1hr. It takes more planning, but I avoid a 1hr partial outage if I mess up the address.
Most sites I've worked with change DNS <1 time per year, so outside infrequent, planned maintenance a long cache time is best for performance.
This is my preferred way of doing it, especially for MX records.
Beware that cache times below 300 seconds will be treated as 300 by some clients, and NS cache times below 2 seconds can cause lookup failures.
I know there was a DNS resolver that did worse than impose a minimum like that, it treat short TTLs (<300s) as errors and "corrected" them by using its default (often 24 hours). I actually encountered it and its effect in the wild, but this was back in the very late 90s or very early 00s. It was a DNS cache daemon on an old Unix variant (I forget which one, my university had a few floating around at the time), so old that if someone is still using it they have far greater issues than my low TTL value!
My current attitude is that if I have the need for short TTLs then I'm going to use them (though 300s is more than short enough for anything I envisage doing ATM & 300s seems to be reliable) and if someone's DNS cache is broken that is their problem, much like I don't pay any attention to issues Internet Explorer users might have looking at some HTML+CSS+JS I've inexpertly strung together.
Some heavily loaded ISP's have modified daemons that set a lower and upper TTL threshold. Anyone running Unbound DNS can also do this. Some versions and distributions of Java do all manor of odd things with DNS including ignoring TTL's though I do not have a current table of those misbehaving. The same goes for some IoT's but I have no idea what resolver libraries they are using.
Yep. Just checked the docs for that and “cache-min-ttl” is a thing. For those unfamiliar, unbound is a pretty common resolver: it is the default for most BSDs and things based on them (such as pfSense) and used by the popular PiHole and its forks. How common it is to use this setting I can't comment on.
> Some versions and distributions of Java do all manor of odd things with DNS including ignoring TTL's
This at least is less of an issue if all you are futzing with the DNS for is services intended to be consumed by a web browser: if common browsers and DNS resolvers behave OK then you are good, if someone consuming your stuff through their own code hits a problem they can work around it themselves :)
It goes without saying that only a fool would play trick like this for email as that is already a mountain of things that can be easily upset by anything off the beaten track.
EDIT:
Also, yep, people are using the setting in their normal environments: https://stackoverflow.com/questions/21799834/how-to-determin... (that person currently having it set to 30 minutes)
Java/JVM historically ignored TTL and just did a single DNS lookup and cached the result unless configured otherwise.
PGBouncer ignores DNS TTL and has its own configuration.
I've seen lots of software that will create a "connection" object and re-use that with TCP reconnect handling. In practice, that means the DNS resolution happens when the connection is recreated but not necessarily when the TCP connection reconnects due to timeout/failure. I've also seen some issues with long lived connections (where the connection greatly outlives the TTL). Long lived connections aren't necessarily a DNS TTL problem but they mean DNS load balancing/redirection don't always work how you expect.
That said, the solution is usually to just fix the issue in the software and not ignore DNS TTLs.
As counter examples, AWS ALBs and Kubernetes service discovery both rely on short TTLs with fairly good success.
Slow onset of impact is probably more useful in minimizing incidents than fast onset of fix (edit: or rather this is a trade-off and one might consider both directions).
Disclaimer I'm not OP
I've handled DNS changes where the underlying host was changing IP address. In that instance the old DNS entry immediately became invalid, so a quick DNS cut-over was needed for all clients.
From the tinydns-data documentation:
--
You may include a timestamp on each line. If ttl is nonzero (or omitted), the timestamp is a starting time for the information in the line; the line will be ignored before that time. If ttl is zero, the timestamp is an ending time (``time to die'') for the information in the line; tinydns dynamically adjusts ttl so that the line's DNS records are not cached for more than a few seconds past the ending time. A timestamp is an external TAI64 timestamp, printed as 16 lowercase hexadecimal characters. For example, the lines
+www.heaven.af.mil:1.2.3.4:0:4000000038af1379
+www.heaven.af.mil:1.2.3.7::4000000038af1379
specify that www.heaven.af.mil will have address 1.2.3.4 until time 4000000038af1379 (2000-02-19 22:04:31 UTC) and will then switch to IP address 1.2.3.7.--
Do other authorative resolvers support this, too?
raw.githubusercontent.com has 1 hour TTL. github.map.fastly.net is 5 minutes TTL. detectportal.firefox.com is 60/120 seconds TTL.
Another thing missing from the post is exactly _how_ bad is it -- because latency matters. If there is a choice of 30mS extra latency on initial site visit vs a chance of 1 hour of downtime in case of hardware failure, it is not clear what the best choice is. Especially things like detectportal.firefox.com which are asynchronous and normally do not introduce extra user-visible latency at all.
- NS Records at 86400
- A and CNAME at 3600
- 1-2 days before moving stuff to other infra (ISP, DNS reseller, etc), switch to 600 or lower, then 1-2 days later switch back.
Can't even say why, it's cargo culting but it has always worked out fine, but ofc I'm not talking about production APIs.And every single time there was some problem, setting the TTL to < 1minute would not have helped, so I guess I'll stick to it.
I don't know, but this issue probably has a name ... selection bias?
Also no clue how actual RR DNS load balancing (without anycast) is still widely practiced, because I guess that also influences it a lot, you may be able to fix a DC going down on a completely different level than DNS. And my experience with cheap home routers is basically that they might cache hostnames for a month (we had stuff with tracking pixels in 2013-2017, EOLing some stuff was wild, we switched DNS and days and weeksl later still got valid requests to the old IPs).
The statement here is misguided.
Use whatever TTL serves your purpose.
DNS was designed to be distributed and scalable. And in my opinion it is THE MOST scalable protocol as it basically allowed the entire internet to work that way. Any user in the world can get any public DNS record for a site.
My uninformed model is that it's all pull based, and for changes to propagate you basically have to wait for cached entries to expire. It seems to me that if there were a cache invalidation mechanism you could have fast updates, but long cache TTLs.
(I know that if this doesn't exist, adding it would be basically impossible, since some servers won't update to new protocols)
There's also DNS Push (RFC 8765), but it's a very recent mechanism. I'm not familiar, but I doubt it's widely supported.
It does. You need to get loaded servers out of what's being advertised. The only way to do this is with a TTL.
Another example of the Internet routing around damage. The DNS caching design dates back to the 1980s when it made perfect sense. It no longer does.
Since a single zone can have multiple labels ({abc.def}.ghi.example - abc.def is a NAME in ghi.example), the server needs to receive the entire requested name in order to properly respond (yeah, root queries usualy optimize this, because all NAMEs there are one label long).
Of course, you can force a flush of the cache using `pihole restartdns` in case you need to.
So you set them low the first time you run into a DNS issue where you had to wait an hour or a day to resolve some bad configuration and your "fuck this" meter went from 0 to 100.
I have a commit of ~15bn queries/month with my DNS provider ATM, and if just one busy record has its TTL set erroneously low, it can cost an extra few thousand $ a month in overages. It’s not a great conversation to have with the CFO when that happens. Caching matters.
Unless your traffic is non-revenue generating or somehow poorly monetizable...
In any business doing proper budgeting and trying to make a profit, a few thousands in UNEXPECTED costs, is huge. It can be the difference between your department having the money for new equipment, or begging for a budget increase because of unexpected costs to get new equipment.
It seems like the problem is expecting a fixed infrastructure budget while trying to create ever larger customer engagement. Is it just modern companies that don't realize you can't count your profits until you've got your AWS invoices for that month?
It's a fixed _overhead_ issue, not a fixed _budget_ issue.
Your only way out is the built in tools to limit service in these cases or to build your own circuit breakers and implement them, or make your usage so small as to not truly require the cloud in the first place.
I'll admit I don't have the stats in front of me, but I'd bet over half of them either don't attempt caching at all or do a horrible job at it.
Of course, your upstream resolver (hopefully) does do some caching, but as most of the latency is getting to this resolver in the first place that doesn't really matter.
They mention latency but I wonder what the real world impact is. On a home network, most users are using ISP supplied DNS which /usually/ has pretty low latency and will have a huge cache already built up. Corporate users will have DNS cached to all the common host they're visiting.
Anecdotally, AWS ALBs (HTTP load balancers) have a CNAME with a 10 minute TTL (maybe it is 60 seconds) and can have their IPs change 100s of times a month and this works fine as a default configuration for many use cases.
Now most of my certs come from Let’s Encrypt via DV, which checks DNS. So I have to repoint DNS first, and risk users seeing a cert error before certbot finishes getting the new cert. So I keep my DNS TTLs a lot lower than I did before.
Also, DNS service is a lot cheaper than it was years ago, so it doesn’t hurt my budget to send more requests back to the name servers.
It would be great if there were some reverse notification to let revolvers know to refresh their caches.