Why are DNS records set with such low TTLs?
- Legacy load balancers left with default settings
- The urban legend that DNS-based load balancing depends on TTLs (it doesn’t - since Netscape Navigator, clients pick a random IP from a RR set, and transparently try another one if they can’t connect)
- Administrators wanting their changes to be applied immediately, because it may require less planning work.
- As a DNS or load balancer administator, your duty is to efficiently deploy the configuration people ask, not to make websites and services fast.
- Low TTLs give peace of mind.
- People initially use low TTLs for testing, and forget to crank them up later.
Why is this not valid?
If your DNS hosting provider charged you per query (some do, especially when adding features like health checks & load balancing), then it might make a big difference.
...
Please tell me you're not testing in production.
If a load balancer or DC fails we need to ensure traffic moves away fast. Similarly if you want to take a system out for maintenance or perform migrations.
Well, if that's the case, you better have your redundant systems on your normal DNS entries, because there is no chance you will distribute new entries over the internet in 5 minutes, whatever value you specify at the TTL.
Easily 90% of new connections will move following the TTL. Of course, some traffic got a DNS result once in 2003 and is going to use that forever. If it's important traffic, you can trace it and follow up with them. If not, you do the best you can and let the rest go.
One would probably be OK if it was 5 or 10 mins, but it depends on what's behind that dns entry and how often ISP can change the IP.
That’s what is missing from this discussion.
Indeed the root DNS servers are not a single server but pools of geographically distributed servers via anycast.
Every request that comes from the same network is going to be routed the same way. Anycast works great for regional load balancing in general, but it doesn't work for subdividing individual networks.
[1] Google Cloud networking in depth: Cloud Load Balancing desconstructed - https://cloud.google.com/blog/products/networking/google-clo...
[2] What is AWS Global Accelerator: https://docs.aws.amazon.com/global-accelerator/latest/dg/wha...
[3] Tumblr: Hashing Your Way To Handling 23,000 Blog Requests Per Second: http://highscalability.com/blog/2014/8/4/tumblr-hashing-your...
[4] Load Balancing without Load Balancers: https://blog.cloudflare.com/cloudflares-architecture-elimina...
The issue is that you can't do percentage based routing with anycast... in fact, you can ONLY do shortest hop routing with anycast (at least for WAN anycast). That means that, while different edge networks can go to a different datacenter, every individual edge network will hit only a single datacenter.
The key issue is that anycast is a very blunt tool. You are relying on your BGP announcements to route traffic, but you aren't actually in control of where a particular request goes.
Determinism isn't necessarily required either. Probabilistic shifting works fine mostly.
Then you were supposed to update it to a longer TTL when your change had propagated.
So, I guess, after understanding things better... there is no use case really since if it's a new record your change will always propagate, and if it's an old record, lowering the TTL on update doesn't really matter since the old TTL will still be in effect.
It's still no guarantee all changes will propagate within 5 minutes. But it gives some ease of mind to know the bulk of change won't take a day.
Also a lot of people forget the negative caching of NXDOMAIN records which is set by the TTL of the SOA record. Which means that it will take a while for your new record to be resolved if you started querying before you set the record.
Especially it gives you peace of mind that should stuff go badly wrong you can easily revert the change.
Does it actually make a difference? I dunno but it just feels right.
Eg. Five days from now, I'm going to make an infrastructure change that will affect public DNS. Today, I lower my TTL on affected records. Five days later I make the public changes, still using the lower TTL. Once I am satisfied my change will stay in production, I modify TTLs to be more sane.
In our case, primary is AWS with a protection service in front of it, and secondary is our own servers at a data center. So something like VRRP wouldn't work.