Google Public DNS's approach to fight against cache poisoning attacks
security.googleblog.com
security.googleblog.com
As a side note, whatever group in Google are running DNS are doing a great job. I do not see malformed garbage coming from them. Some people watch birds when they retire, I watch packets ... and critters.
Why are your records high? Is the load on your servers too intense otherwise? (not trying to debate, am curious because I've really appreciated your comments/insights in the past)
if (SYN->dest port && TTL > X) { ACK };Are you maybe confusing IP packet TTL with DNS record TTL?
For my personal hobby sites I change things so rarely that I can keep my TTL's very high, especially for NS records and most TXT records. There are some records I could theoretically set to 68 years though obviously nobody would keep a record that long. Going against the grain I also keep neg-ttl negative cache really high to help spot the bots that ignore neg-ttl. It's a hobby of mine to study bots and what they are enumerating. Sometimes it gives me a jump start on zero-day vulnerabilities. For example a number of bots suddenly start looking for a specific A record such as cpanel to use an old and silly example.
Non bot clients will respect TTL's within reason. Most ISP recursive DNS servers will cap the TTL to 24 hours for NS records and sometimes higher for A, CNAME, PTR, etc... Most corporate DNS servers will be close to the defaults of whatever recursive daemon they are using, usually Active Directory, sometimes Bind. The remaining limiting factor is memory but most recursive DNS servers these days have obscene amounts of RAM and CPU time that the DNS admin may allocate. There are ways to even further optimize recursive servers such as periodically flushing junk zones via cron and many other things that do not need to be there as well as tuning slabs and threads based on core count. This will vary by organization and requires getting detailed zone and client statistics. An example of a junk zone would be a zone used by a corporate spy to exfiltrate data over DNS that will result in hundreds of thousands of unique A records data-flow-outbound customer or intellectual property data or TXT records inbound that are actually encrypted data data-flow-inbound encrypted malware, instrucitons.
Bot scripts will bypass recursive DNS servers and will typically talk directly to authoritative servers. Most of these scripts do not even look at zone or resource record TTL times which is unfortunate for them as it makes spotting them trivial. I then dig deeper into the networks they are originating from, find their IPv4/IPv6 CIDR blocks, who they peer with, what business they claim to be. Sometimes they are squatters that are announcing routes from businesses that went under and had laid off the people that would have released the IP allocations. I help get those clawed back when I can, taking the IP allocations away from the squatters.
Side project idea for Google / Cloudflare / OpenDNS / etc... These companies have a unique view into DNS traffic and could also help spot the squatters. They could free up a significant amount of IP space if they allocated a small team of interns to analyze their traffic, use automation to build graphs and reports, then once confidence is high enough submit this data to all the IP registries. Registries are also much more likely to take communication from these companies far more seriously than some retired hobbyist.
This will probably come up but some may say having high TTL's is risky. It can be, but not for me. If all of the internet shared one massive /etc/hosts and DNS ceased to exist, it would rarely have updates from me and if a record is stale there would be no harm, no foul as it pertains to me. There are a myriad of other reasons I have unusually high zone/record TTL's but it would turn into a blog post and I have been too lazy as of late to make any as interest is usually very low. I don't even have my blog VM's spun up.
TTL is still in seconds, but every hop has to decrease it by at least one.
I'm not sure if there is any implementation out there that cares, but if e.g. wifi retries or a massive buffer queue lead to a packet spending more than a second on a hop, it should decrease by two.
If an intermediate server handed out the original TTL for all requests then the TTL would be effectively multiplied by each caching layer³, which is why you will only get the true TTL at the client end if a fresh request was actually made to the authoritative server(s).
--
[1] the fraction of a second difference caused by latency isn't likely to be significant for either long or short TTLs
[2] and on a really high latency link (like back when GPRS or POTS landlines were common, or some links even now in deep rural areas) 1s might not be enough anyway for the benefit they think they are giving
[3] so at my home where lookups go PiHole->8.8.8.8->authoritative it would make TTLs potentially up to 2x their intended length. Example: for a 1000s TTL if PiHole gets the lookup request with 1s remaining but hands out the true TTL, my client will not check for 1000s, potentially 999s late. Depending on the sequence of related requests to each DNS server from other clients, that could become 1998s in my case, longer if there were more caching layers.
It would be interesting to hear how often the google dns servers see attempts to poison their cache. The mitigations probably prevent folks from even trying, but he numbers would be interesting.
The OARC 40 presentation PDF mentions cookies deployment is low for large operators but open source software has compliant implementations. Are large operators writing their own dns servers, but badly? I would think there wouldn't be many custom implementations, and that you would be able to detect which software nameservers are running, each with known capabilities. But from the way the numbers are presented it seems they only look at behaviour without considering software (versions).
Yes. It is trivial to build a DNS server, and near impossible to write a correct DNS server. Eventually your organization gets large enough that someone thinks it is a good idea without understanding the implications.
> you would be able to detect which software nameservers are running, each with known capabilities
There are pseudo-standards for asking an authoritative server what software it is running, but everyone turns that off because somehow it makes you "more secure." What you end up having to do is probe auth servers by replaying user queries on the side, measuring if the responses you get are correct, and then keeping a database somewhere of which servers support which flags.
Today I let my curiosity dive deeper and quickly found the ietf publication on 0x20 encoding and this article.
Just odd to see others post it to hn on the same day.. coincidences are weird.
That's why they don't mention DNSSEC: because it isn't a significant security mechanism for Internet DNS. It's also why they do mention ADoT: because it is.
I think this is also why DNSSEC advocates are so fixated on DANE, which is (necessarily) an even harder lift than getting DNSSEC deployed: because the attacks DNSSEC were ostensibly designed to address are now solved problems.
Note also that if ADoT rolls out all the way --- it's already significantly more available than DNSSEC! --- there won't even be a real architectural argument for DNSSEC anymore, because we'll have end-to-end cryptographic security for the DNS.
Thanks for calling this out! That Google feels case randomization is more important than DNSSEC is indeed telling.
Maybe you can help figure out where you went wrong here by explaining what - in your understanding - were the problems that DNSSEC was "ostensibly designed to address" ?
> because we'll have end-to-end cryptographic security for the DNS.
In 1995 a researcher who was annoyed about people snooping his passwords over telnet invented a protocol (and gave away a free Unix program) which I guess you'd say delivers "end-to-end cryptographic security" for the remote shell. Now, when you go into a startup and you find they've set up a bunch of ad hoc SSH servers and their people are just agreeing to all the "Are you sure you want to continue ..?" messages, do you think "That's fine, it's end-to-end cryptographic security" ? Or do you immediately put that on the Must Do list for basic security because it's an obvious vulnerability ?
ADoT relies on NS records to be DNSSEC signed.
The TLS certificates that ADoT relies on need to be hashed into TLSA records (DANE, DNSSEC).
https://datatracker.ietf.org/doc/draft-dickson-dprive-adot-a...
While I agree transport confidentiality is important, that is not what DNSSEC solves, nor should you see people saying that it does solve confidentiality.
DNSSEC protects the transport integrity of DNS responses. DNSSEC enabled zones 100% defeat on-path cache poisoning attacks to recursive resolvers that are DNSSEC enabled. Full stop.
ADo"X" protects the transport confidentiality of DNS responses. I suppose this "weakly" protects the transport integrity of DNS responses, but, again, the primary purpose is confidentiality.
After reading RFC 9539, it's clearly stated that it's opportunistic encryption. An on-path attacker will find it trivial to disable encryption and start poisoning caches. The RFC states in multiple places that if TLS setup fails, fall back to plaintext DNS.
If DNSSEC fails, an on-path attacker has no similar recourse. A properly configured recursive resolver will SERVFAIL and _never_ send back a potentially poisoned response to clients for DNSSEC signed zones.
DNSSEC is the ultimate defense against cache poisoning attacks, no matter the adoption percentage.
(I also note that you didn’t answer my question, and instead opted for a rhetorical cheap shot reply to my second paragraph only.)
Unfortunately, most people do not DNSSEC sign their zones, so Google have to resort to also enabling 0x20, which is helpful, but also (to an extent) security theater.
[0] - https://www.verisign.com/en_US/company-information/verisign-...
[1] - https://www.statdns.com/
https://news.ycombinator.com/item?id=36171696 - Calling time on DNSSEC: The costs exceed the benefits (2023)
And also many news regarding validation failures:
https://www.potaroo.net/ispcol/2023-02/dnssec.html
(Geoff Huston is an Internet infrastructure giant.)
But really it all just boils down to the fact that the DNS zones that matter --- the ones at the busy end of the fat tail of lookups --- just aren't signed, despite 25 years of work on the standard. IPv6 is gradually mainstreaming; in countries where registrars auto-sign zones, DNSSEC is growing too, but very notably in countries where people have a choice, DNSSEC deployment is stubbornly stuck in the low single digit percentages, and the zones that are getting signed are disproportionately not in the top 10,000 of the Tranco list.
Some other counterpoints to general DNSSEC doomsayers:
• <https://blog.technitium.com/2023/05/for-dnssec-and-why-dane-...>
• <https://www.redpill-linpro.com/techblog/2019/05/06/sshfp-and...>
Edit: Is this authoritative enough for you? <https://www.icann.org/resources/pages/dnssec-what-is-it-why-...>
This is what I mean when I say we're not really talking to each other. I don't think you understand or care about the argument I'm making, and so you're not engaging with it. That's fine! But then: let's just stop engaging.
If you can't securely authenticate your server (as HTTPS/TLS does) you have other problems too.
Still, good to see that monitoring the CT logs would have caught this problem much sooner.
The reasoning is that DNS is not important enough to go through the trouble of deploying DNSSEC. These days TLS is often cited as the reason DNSSEC is not needed.
At the same time we see a lot of interest in techniques to prevent cache poisoning and other spoofing attacks. Suddenly in those cases DNS is important.
If all DNS client software would drop UDP source port randomization and randomized IDs, then lots of people would be very upset. Because DNS security was more important than claimed.
DNS cookies are also an interesting case. They can stop most cache poisoning attacks. But from the google article, big DNS authoritatives do not deploy them.
The interesting thing is what happens when BGP is used to redirect traffic to DNS servers: https://www.thousandeyes.com/blog/amazon-route-53-dns-and-bg...
The weird irony is it's the old "worse is better" winning again. HTTP and TLS are fairly bad protocols, in their own ways. But put them together and they're better than whatever else exists. It's just too bad we didn't keep them and ditch the browser.
The problem is that applications typically use TCP connections, but IPSEC works at the IP level. Early on, the (BSD socket) kernel API was basically fixed at the IP level instead of associating it with a TCP socket.
So the whole thing became too complex (also for other reasons). So SSL and SSH were created to have simple things that worked.
SSL took many iterations to get any kind of security, so IPSEC had plenty of time to get it right and take over. But as far as I know, there just never happened. It also doesn't help that TLS is trivial to combine with NAT, and for IPSEC that is quite tricky.
If it becomes popular enough you will certainly face future security challenges you failed to even imagine. Leave some room for that.
Otherwise, this is great work.