DNS can fail when someone who is managing it have no idea what the hell they are doing and abuse it to do things it wasn't meant to do.
1. Clients query a list of servers (IPs) and handle failover when you don't quickly get a reply.
2. Most of those servers at the root and TLD level are actually anycasted from multiple locations globally, so you connect to the closest instance.
3. Those instances are often clusters of physical servers. The big ones have fully redundant networking, so any router or switch failing doesn't take it down. Some run different DNS server software on each physical server, so even software bugs won't take down the whole system.
And then the problem is confounded by the fact that, ironically, DNS works so well that we don’t think of it as a primary point of failure. So inevitably, when there is a DNS problem, it’s the last thing we check. This just reinforces the idea that the problem is always DNS… because in those long, hard to troubleshoot instances… the problem was DNS.
All variants of ‘DNS is working perfectly, just your expectations of how it will work in your situation are not completely correct’.
// wrote what might have been first commercial-use dynamic DNS server for a regional ISP in early 90s, invented an unreasonably effective geo+latency balanced anycast-like DNS for global video delivery network in 00s
Except when Windows is the architecture.
This may be a place where Garnet is a good alternative.