The case of the recursive resolvers: What happened during Slack’s DNSSEC rollout
slack.engineering
slack.engineering
1. "This strict DNS spec enforcement will reject a CNAME record at the apex of a zone (as per RFC-2181), including the APEX of a sub-delegated subdomain. This was the reason that customers using VPN providers were disproportionately" - This is non intuitive and maay people are surprised by that. You cannot create any subdomain (even www.domain.tld) if you created "domain.tld CNAME something...". Looks like not every server/resolver enforces that restriction.
2. "based on expert advice, our understanding at the time was that DS records at the .com zone were never cached, so pulling it from the registrar would cause resolvers to immediately stop performing DNSSEC validation." - like any other record, they can be cached. DNS has also negative caching (caching of "not found responses". Moreover there are resolvers that allow configuring minimum TTL that can be higher that what your NS servers returns (like unbound - "cache-min-ttl" option) or can be configured to serve stale responses in case of resolution failures after the cached data expires [2]. That means returning TTL of "1s" will not work as you expect.
[1] https://blog.powerdns.com/2020/11/27/goodbye-dns-goodbye-pow... [2] https://www.isc.org/blogs/2020-serve-stale/
Like networking there can also be existing protocol errors and plain broken things that has for one reason or an other been seemingly working for decades without causing a problem. Internet flag day is one of those things that pokes at those problems, and maybe one day we will see a test for CNAME at the apex.
The Saltzer and Reed paper, if I'm remembering right, even calls out security as specifically one of those things you don't want to be doing in the middle of the network.
See also: Zero Trust / BeyondCorp.
There has been many that has suggested that we should just scrap the whole thing called The Internet and start from scratch. It would be safer, but I don't think it is a serious alternative. DNS, BGP, IP, UDP, TCP, and HTTP to name a few are seeing incremental changes, and the cost is preferable over the alternative of doing nothing. Ambitious security things would be much less costly if we had working redundancy in place, which is one of those things that flag day tend to illustrate. Good redundancy and people won't notice when HTTP becomes HTTP/2 that later becomes HTTP/3. It also helped development at google that when they added QUIC, they controlled both ends of the connection.
See second-system effect:
If it takes a designer of DNSSEC to implement it, then how should I, a peasant implement DNSSEC for my infra?
1. A bug in Route 53 which caused wildcard record not to work with DNSSEC signing. Anyone not using Route 53 would not have had any problems with DNSSEC.
2. Slack decided to revert the DNSSEC rollout, but botched the process badly, effectively locking themselves in the trunk and throwing away the key. If they hadn’t tried to revert the DNSSEC rollout, or if they had been a bit more deliberate and careful while doing it, this would not have happened.
https://news.ycombinator.com/item?id=29381778
That thread, which is big, is probably the right place to take general discussion of DNSSEC itself, though I'll snipe DNSSEC here too. :)
Google Workspace is a good point though. I know there are many users of it in government... maybe some AOs are fine signing off on it even without the needed security controls, which is an option they have in their discretion with and without FedRAMP.
> While we are aware of the debate around the utility of DNSSEC among the DNS community, we are still committed to securing Slack for our customers.
The argument is specifically that it doesn't provide that security. At least it's neat to see actual begging the question in the wild, I guess.
Not everyone agrees with the linked argument. For example, I disagree that browsers can't take advantage of DNSSEC, since many are using DoH, and the rest of the article reads like someone complaining that we need to wait for the perfect protocol or nothing at all.
That's the thing about a debate... it's got arguments on both sides.
But I'd also say that DoH (1) largely obviates any need for DNSSEC (the last-mile DNS problem is the only on-the-wire DNS security problem that needs solving) and (2) doesn't enable DANE in browsers, which is what people are talking about when they talk about DNSSEC intersecting with browsers in any way other than randomly making sites fall off the Internet.
Those security controls come from a document NIST SP 800-53, 2 of which (that Slack linked to in the linked post-mortem), SC-20 and SC-21, effectively seem to me to conspire to require DNSSEC. Both of these are included as part of the "Low" baseline of security controls, so they are effectively required for all Federal IT systems unless your Agency Authorizing Official wants to walk on the wild side.
So even if you get a FedRAMP certification, if you do it without fully implementing SC-20 and SC-21, that just means your customer needs to either convince their Agency Authorizing Official to sign off on an ATO despite the missing SC-20 and SC-21 security control, convince them to sign off on some sort of Plan of Action and Milestones where Slack will commit to fix this in the future (which is just kicking the can down the road), or somehow manage to implement the same effect completely within the customer end without help from Slack. All you would have done is to spend a lot of money on FedRAMP paperwork without making it appreciably easier for potential customers who have to deal with compliance regimes to buy your product.
Cloud.gov's argument is valid but all they posted is that they don't implement SC-20 or SC-21 for their government customers, and that the OMB M-08-23 mandate for DNSSEC is no longer operative (not that no other DNSSEC mandate applies). Indeed they even give explanation for how their customers should work to enable it (presumably by refusing to use the non-DNSSEC compliant .app.cloud.gov services and instead using only their DNSSEC-compliant custom domains).
FWIW I fully agree with tptacek's arguments against DNSSEC, and will note that I recently stopped being able to navigate to literally the entire .mil on my Linux host until I disabled DNSSEC in systemd, for reasons that are still unclear to me even now.
Complex systems can and will fail. Try to do better, of course, but let’s acknowledge that perfection will always exceed our grasp. The world will continue to turn regardless.
One day it might just be your turn to break production.
Perhaps I did not read the room appropriately. Mea culpa.
I think another interesting question here is why Slack bothered in the first place. As was pointed out on the other DNSSEC thread today: practically nobody in the technology industry uses DNSSEC in the first place. Presumably, Slack did DNSSEC (they don't anymore!) in service of FedRAMP compliance. Why? Slack has one of the most popular products in all of computing. What bad thing was going to happen if they said "nah, we're going to go with Cloud.gov's recommendation and not this FedRAMP document"?
(Kenn White points out on Twitter that some of this may be due to grandfathering --- though, the FedRAMP DNSSEC requirement is pretty old.)
When the DOD tried to mandate Ada, lots of projects were bid as Ada, then switched to C++ at the very first sign of any trouble whatsoever. I would 100% believe it if someone told me that this horrible rollout could be leveraged into an exemption from needing DNSSEC
Was it a hard requirement? No, but the fat fingered audit companies really like to tick that "should" box green and would be more lenient with other debatable findings, so it was suddenly "in our best interests" to comply.
As just one example, it's tremendously difficult, if not impossible, to sell your cloud-based SaaS to Navy customers if you have open FedRAMP compliance issues that you aren't at least working to address.
I say "compliance" instead of "security" for a reason as well, as "compliance" truly runs the show in Navy cybersecurity. And if you want to sell to that market (and it's hardly just Navy who runs this way), it's easier to check the checkboxes than it is to argue about whether NIST is right or cloud.gov is right.
I'm pretty surprised that slack doesn't have a more robust testing network. Is it really that hard to set up another DNS on Route53 for staging these changes? Idk, but that type of thing is the least you can do if you want some FBI agents to discuss active investigations on your chat platform...
(There's a whole thread here, and more on Twitter, getting into the actual details of what FedRAMP and NIST require here, and engaging with the fact that Slack is the only large tech company in the past several years to have attempted to flip the DNSSEC switch on.)
Your blog post makes the supposition that DNSSEC is only being pushed as an alternative means of security to CA for TLS. While it makes a compelling case that this isn't realistic, there are other security concerns that occur from the compromise of DNS records. If the government is going to use a DNS record, it should be signed by a zone owner.
Slack is actually a good use case for this security enforcement, because they maintain a handful of domains that are extremely authoritative for their messaging service[1]. If you can't maintain a security protocol on four domains that are crucial to the operation of your service, you maybe aren't cut out to supply software for the government.
1: https://slack.com/help/articles/360001603387-Manage-Slack-co...
Unfortunately, the GSA product market is its own bubble, as is people who work in IT for the USG in any capacity, and so it's easy to see how people with limited exposure to modern industry practice --- experiences almost wholly gated through vendors that snake through the GSA acquisition process --- might believe themselves to be operating several levels above where they actually are.
I would take Slack's security practice --- their infrasec, their corpsec, their software security, the whole shebang --- over anything done in any USG agency. Slack is better at this than their USG clients are, full stop. And Slack, while strong, is far from the S tier of industry security teams.
I just want to hammer home the point that requiring service providers to get their DNS records signed by DNS zone owners is a reasonable ask for USG software service vendors. Even if DNSSEC isn't capable of securing the whole internet.
Either way, this argument is starting to become political. Is Facebook a role model for cybersecurity, and keeping data out of the wrong hands? Or do NIST researchers know better? Neither - the government outlines its security requirements, and private companies play ball to compete for their business. And if a federal agency wants to be able to prove a DNS record's authenticity, even if it is maintained by a vendor, even if that isn't sufficient to secure their infrastructure, that's their prerogative.
Slack's second attempt wasn't a DNSSEC problem. Slack depended on a permissive fallback of revolvers when encountering a plain DNS protocol error. It is similar to how some websites in the past relied on permissive browsers implementation when facing broken HTML/JS/CSS. Slack fixed their broken DNS as a result of this.
Slack's third attempt was not the fault of Slack but rather a software bug at Amazon. I would make the argument that Amazon's primary product isn't DNS services, but they did fixed their bug after this.
The general conclusion I get from the article is not that DNSSEC is broken, nor that is too complicated. It is that when doing changes with your core infrastructure to make it more secure, bugs that may have been laying dormant might pop up and bite. I am sure some people has had that experience in domains outside of DNS.
What one can't ignore is the underlying chicken-and-egg problem that DNSSEC must overcome: Not many DNSSEC deployments and hence not much of it has been tested in the real-world, which results in colossal outages despite the attention of some of the most qualified engs, including the ones running one of the largest nameserver deployments in the world.
TLS and WebPKI has had a similar, perhaps even more painful route to ubiquity. So, this problem isn't unique to DNSSEC. What isn't working in DNSSEC's favour is, the world has not just moved on, but it has built solutions atop DNS' weaknesses, like it once did with IPv4 and NAT. Internet's strong network-effects coupled with its heterogeneity, make battling "the System" an even harder proposition.
See also: System design explains the world: Vol 1, https://apenwarr.ca/log/20201227
Sometimes it's BGP.