Chromium's impact on root DNS traffic (2020)
blog.apnic.net
blog.apnic.net
> In a previous blog post, we quantified that upwards of 45.80% of total DNS traffic to the root servers was, at the time, the result of Chromium intranet redirection detection tests. Since then, the Chromium team has redesigned its code to disable the redirection test on Android systems and introduced a multi-state DNS interception policy that supports disabling the redirection test for desktop browsers. This functionality was released mid-November of 2020 for Android systems in Chromium 87 and, quickly thereafter, the root servers experienced a rapid decline of DNS queries.
https://blog.verisign.com/domain-names/chromiums-reduction-o...
“I’m the original author of this code, though I no longer maintain it.
”Just want to give folks a heads-up that we’ve been in discussion with various parties about this for some time now, as we agree the negative effects on the root servers are undesirable. So in terms of “do the Chromium folks know/care”; yes, and yes.
”This is a challenging problem to solve, as we’ve repeatedly seen in the past that NXDOMAIN hijacking is pervasive (I can’t speak quantitatively or claim that “exception rather than the norm” is false, but it’s certainly at least very common), and network operators often actively worked around previous attempts to detect it in order to ensure they could hijack without Chromium detecting it.
”We do have some proposed designs to fix; at this point much of the issue is engineering bandwidth. Though I’m sure that if ISPs wanted to cooperate with us on a reliable way to detect NXDOMAIN hijacking, we wouldn’t object. Ideally, we wouldn’t have to frustrate them, nor they us.”
A way for DNS root servers and Chromium devs to cooperate that couldn’t be hijacked by domain redirecting ISPs would be nice.
I doubt Google would be happy if Microsoft or Apple were to build a network availability tool into their systems that executes random Google queries that won't get you a result to see if searches are being redirected.
This isn't Google using a competitor's resources. It's Google's users using core Internet infrastructure. It's core infra for a reason. We're not required to use it parsimoniously.
I have the impression that most of the concerned comments on this thread believe that Chrome's NXDOMAIN detection is somehow a burden on the Internet roots. It is not. It's prominent on a graph, and that's all. The Internet roots are, by necessity, designed to handle immense amounts of traffic.
It would, on the other hand, create enormous new load on the DNS infrastructure.
DNS was designed with trusted resolvers in mind (was it?), but we've shifted towards using resolvers of our ISPs and random companies.
I can see the appeal of caching across a WAN, but then again, we don't use HTTP proxies anymore, which provide a similar caching feature. Integrity protection should be placed higher in our priority list than caching effectiveness.
Maybe the impact on root servers created with selfhosted resolvers would decrease as chromium wouldn't have to query random TLDs -- root could just be cached for days.
Actually, thinking about caching root, why don't resolvers already cache root? Chrome wouldn't have such an impact if NSEC from root was remembered by resolvers. DNSSEC solves this problem even nowadays.
DNSSEC is hardly deployed at all and doesn't really enter into this story. It doesn't change anything about what Chrome is trying to do, and, were it actually deployed, would dwarf all other causes of increased root server loads by itself.
Yes, but see, it's not a distributed denial of service attack. It's a productive use of a capability the DNS roots were designed to serve. It's a huge amount of traffic because Chrome is one of the 3 most important and popular applications on the Internet, and the most popular of those, at that. The cart doesn't drag the horse.
Everybody can look at this DNS probe code and kibitz about whether it could have been done more efficiently or whether it should be done at all, but this is a straightforward use of the DNS to do something important for end-users (detect whether their ISP's nameserver is lying to them about NXDOMAIN). Whatever else it is, it isn't a DDOS attack, and whatever else these servers are, they're not quiet out-of-the-way private systems; they're core Internet infrastructure.
Like any computer system, Root DNS was designed around expectations of typical utilization and expected response time, with margins for burst traffic. Capacity was left unused for a reason.
Google's action here is an example of cost shifting. It doesn't cost Google or the individual user anything to test for a hijacked DNS resolver, because the costs are shifted to the Root DNS. As the article puts it, this cost shifting is indistinguishable from a DDoS because it consumes network capacity and CPU resources and electrical power in an unplanned fashion.
Root DNS was designed around resolving names, not testing for hijacked DNS. And just because you can use DNS to test for hijacked DNS doesn't mean you should.
Says who? You? Have you double checked that belief with the RFCs? Last I looked, there's a lot more than just name->IP in the DNS. "Testing for hijacked DNS" is a DNS function. The root servers are there to make applications work, not the other way around. They're fine, and people watching them don't dictate terms to the browsers.
Nothing that I can find there about checking for hijacked DNS. Did you have a specific RFC/section in mind that you were thinking of?
(Damn, there are a lot of RFCs devoted to DNS!)
There is an obvious flaw which is that this behavior is detectable and avoidable by ISPs that become aware of what that domain is being used for... But that argument applies to the current detection method as well (as this post demonstrates).
No, because the post is a post-mortem analysis. If ISPs are going to block this they will have to do so in live, and as presented I don't think they can do that without a disruption.
The original point stands in that this feature abuses a common good at the expense of an operator (requiring 2x the capacity without it by the metrics given in this post) for a niche feature that they could run with equal effectiveness against a domain they control instead.
Attempting to look up domains with the intention that they don't exist and are expecting the general case to be a failed query is not using the DNS system as intended. As I mentioned this could be solved by the application including a domain they control so they are not excessively consuming a public good at the expense of others (someone does have to pay for these servers).
The roots should not be expected to double their capacity because one application implements a feature in bad faith.
That's apparently 60 billion queries per day. Doing the math, a query is some 92 bytes on the wire (50 bytes UDP payload, but the root servers are probably also on ethernet or something similar so I'll include headers), so 487 megabits per second per root server is taken up by this, assuming there are no spikes and everything is perfectly averaged. Edit: and that's the downlink. I forgot the uplink traffic that will be larger due to DNS' nature of echoing the query back with the response(s) attached, plus DNSSEC... Woa. /edit.
So what's the next step, Google generously stepping up to running (some of) those root servers and gaining further control and observability over what people do online?
Reading further, Firefox solves this by "namespace probe queries, directing them away from the root servers towards the browser’s infrastructure".
For infrastructure hosted in a legit data center, this is nothing.
On top of that come captive hotspot/corporate portals - and here especially, I do wonder why messing around with DPI is necessary, when every public network has a DHCP server running that could be used to distribute DHCP options for gateway URLs or information if the connection is to be considered metered (mobile phone hotspot) or bandwidth constrained (train or bus hotspots).
DHCP isn't reliable as many clients don't do anything with advanced settings you provide. Good luck getting a phone to accept the proxy server you've configured over DHCP.
This legitimate-ish DPI usually runs at the network itself, it doesn't traverse the uplink to cause any load.
The other thing that people would be very suprised about is how old the software is on those root server, forget modern libraries with Rust/c++ and the like, it's pretty old tech that is very inefficient.
Edit: when talking about old tech I'm talking about the architecture of the DNS server used, IO libraries and models, caching, data structure and the like, for example a lot of stuff has been done arround web servers to serve things very efficiently, the same could be done on DNS servers.
There is nothing about Rust or C++ that make them faster than C.
In what way are the root servers inefficient?
I can guarantee you that this "old tech" was coded with more thought invested into it than at least half of these "modern libraries".
Well written C code can easily blow a C++/Rust application out of the water.
Of course, for the sake of completeness it should be noted that:
> With the run-time checks disabled and the restrictions loosened, Rust presents a performance indistinguishable from C. [1]
Though I believe my original statement to hold none the less, as disabling these restrictions disables (amongst other things) bound and overflow checking, which is one of if not the major selling point of rust.
As for C++ depends on the features that one uses. If one writes "just what one could do in C" then the machine code produced by the compiler will be exactly (almost) the same. This is due to the fact that many c++ features are only compiler relevant but compile to (almost) the same instructions as code.
However, I would once again raise the question I did above with rust: If we use little to no c++ features then can we distinguish that codebasse from a c codebase in any meaningful way? But assuming we write idomatic code we will have the c++ code behaving somewhat slower due to factors such as:
- automatic collections/object allocation. Datastructures growing "on demand" do in general perform slower than a comparable "none automatic" datastructure increased in larger chunks by hand (using malloc/etc. in C). While this is an implementation detail admittedly, I believe the libstdc++ does not use chunking, though I would not swear on that. - strings. While no doubt a big upgrade from \0 terminated char sequences idomatic strings in C are less efficient. Especially when it comes to concatenating or manipulation of said strings. In addition it may lead to memory fragmentation, though this should be an afterthought most of the time.
In general the performance difference of C++/C comes down to "hidden" code. While by no means large, assuming software such as the dns root servers which are running essentially 24/7 and will most likely continue doing so for quite a while even small differences in performance will add up.
Admittedly however my original statement of
> Well written C code can easily blow a C++/Rust application out of the water.
May not have been well formulated. It would have been better to split the statement and be more specific about the individual performance differences in regards to rust/c++ instead of bunching them together.
Then again, maybe it's revenge for everyone pinging 8.8.8.8
I can see some purposes for detecting middleboxes. I've done it. It usually doesn't involve DNS. It does involve certificate pinning though.
> detecting if the user has a captive portal between them and the Internet
That's easy. Try to browse to something. If succeeds but the certificate isn't valid then the user probably has a captive portal. That, or your pinned certificate has been revoked.
> detecting ... if the user's provider messes around with DNS
Certificate pinning, again, comes to the rescue. Pin a certificate to your own DoH server and then use DoH to look up whatever you need.
If you can't connect to your DoH server then you effectively aren't (or shouldn't be) connected to the internet.
That the browsers are reacting in this way says we have a failure at the DNS software level. Those projects do seem to be giving the security that is wanted. So we are starting to get some fragmentation. Which can not be good. Perhaps we need new record types to support this?
The solution would have been DNSSEC, the problem is that authenticating NXDOMAIN responses comes with a ton of challenges on its own and so there, in the end, was just workarounds and messy hacks [1] that IIRC no one ended up utilizing.
[1] https://en.wikipedia.org/wiki/Domain_Name_System_Security_Ex...
If Karen takes my sandwich from the company fridge, so I take some of Jon's lunch, so I don't starve, I'm not innocent because Karen started it, I've created a situation where there are two arseholes instead of one. This isn't quite what is happening here as the root servers are effectively a public resource and stuff in the fridge is all private resources, but close enough to make the point.
> The networks that hijack DNS request should share some of the blame
They should have all the blame for deliberately breaking part of agreed protocols for their own gain.
But that doesn't make anything we do in response to that right by virtue of us doing it because we have been wronged.
Google manages entire TLDs, surely they can use their own DNS servers for this purpose.
Reminds me of the recent "Go module mirror fiasco" where Google found it fair to clone repositories at a rate of ~2,500 per hour in order to essentially proxy Go modules.
- "Sourcehut will blacklist the Go module mirror" - https://news.ycombinator.com/item?id=34310674
After the drama become very much public, they finally decided to address the issue in a good way.
This isn't some tiny authoritative DNS server being flooded unexpectedly with queries. These are the Internet DNS root services. They have to keep up with this kind of traffic. It's their literal job description.
But I dislike the whole GOPROXY design, like it breaking private repos by default and having to set some env variables to make this stupid tool download stuff from server I told it to download.
Well, taking the context into consideration, I'd still say it's too much. Context being they were full git clones and the traffic ended up representing "70% of all outgoing network traffic from git.sr.ht".
So calling it abuse, is wrong. It is a dirty hack for a non-existing feature. It is technical debt of the DNS platform and the root server suffer for it because the ISPs and in-house DNS resolvers create the problem.
Being angry at Google we can anyway be. They have enough money, enough people and enough power to either fix this financially or as a feature within the DNS platform.
G wants it
Honestly, this is why I prefer using FQDN's, and bookmarks. I ask for what I want, and I get it (barring some captive portal, etc).
I don't know what a typical user of Chromium expects to happen when they type arbitrary strings into the Omnibox, but I can tell you that they probably don't expect the full text to be forwarded to several random, centralized, and highly popular DNS servers that are neither under the control of Google nor their ISPs nor their employers/institutions.
This may have the effect of logging all sorts of stuff that normally Google would be hoovering up, but instead it's in the hands of third parties. Extremely trustworthy parties indeed, but they are also high-value high-stakes targets for anyone who'd want to steal valuable data in the form of query logs.
Obviously there's no good solution to this. I believe Firefox led the charge with the "Awesome Bar" or whatever it was first called, which habituated me and millions of users to type search queries directly into the URL bar.
It's one of those features that's so amazingly useful that it really forces the giants to carefully calculate the tradeoff costs, such as the one discussed in TFA.