1.1.1.1 lookup failures on October 4th, 2023
blog.cloudflare.com
blog.cloudflare.com
- There's a clear uptick in SERVFAIL responses at 7:00 UTC but they don't start their response until an hour later after receiving external reports. This uptick should have automatically triggered an alert. It can't have been within the normal range because they got customer reports about it.
- The resolver failed to load the root zone data on startup and resorted a fallback path. Even if this isn't an error for the resolver it should still be an alert for the static_zone service, because its only client is failing to consume its data.
- The static_zone service should also alert when some percentage of instances fail to parse the root zone data, to get ahead of potential problems before the existing data becomes stale.
I have my own unrelated issues with Cloudflare as a company
They might just want a prober; ask every server for cloudflare.com every minute. If that errors, there are big problems. (I remember Google adding itself as malware many many years ago. Nice to have a continuous check for that sort of thing, which I am sure they do now.)
Yeah this was my first thought too. Why don't they have such a system in place for as many permutations of their public services as they can think of? We're a small company and we've had this for critical stuff for several years.
Add issues like network-visibility, and you wind up talking about a cross-team, cross-org effort to stick an HTTP poller and get the traffic back to some server - and so going into production without it winds up being the easier path (because it'll work fine - provided nothing goes wrong).
7:57 UTC: first reports coming in
I noticed this issue quite quickly ("reported" at 7:54 UTC [1]), and I noticed I wasn't alone thanks to Twitter / X. I tried to get in touch with Cloudflare to report this issue - but I haven't found any meaningful contact other than Twitter.
For such an important service, I'm impressed there is no contact email / form where you can get in touch with the engineers responsible for keeping the service up and running.
Other than that, kudos for the well written blog post - as always!
[1]: https://nitter.net/DenysVitali/status/1709476961523835246
Even when I just had 2 paid domains with Cloudflare, there was real time chat support that seemed to always have an available engineer when something came up.
Maybe they don't provide support like that to free users.
For something as core as DNS, I'm sure engineers were aware about the issue within minutes. There's a lot of politics and processes between that and public acknowledgment of an incident.
What I want as an engineer is an engineering contact (similar to a NOC) that "normal customers" aren't aware of.
Something as critical as DNS should have that - but maybe I expect too much from a free service.
cloud: you are the product, not the customer
I get the "you're the product, not the customer" for social networks or ad-based search engines, but certainly not for cloud products.
I got on particular fix one time in just a few minutes, and scalated pretty quickly. I got a recomendation fix on my side (a code change on my deployment that cover the problem) and a permanent fix on CF side 4h latter (time that took to changes propagate to all colos)
They're usually busy doing other things.
but yea, repeatedly running into a wall of ignorance encourages shifting to "busy doing other things"
Ding ding ding.
At any rate: it's very funny that DNSSEC took 1.1.1.1 down, but this bug can't honestly be pinned on DNSSEC itself.
If the DNSSEC didn't add new unnecessary complexity to an otherwise working system, there would be no bug, and no stale data.
If you noticed your brakes had failed because you ended up in a ditch I wouldn't really say that's a positive outcome.
Frankly I can't believe they don't have better monitoring for a system as critical as that.
They used 8.8.8.8 in their example, which I of course changed to 1.1.1.1!
(Un)surprisingly the issue did not resolve after this, and I went on to do other things, hoping the gremlins would find something better to do in the meantime.
webassembly? what is that word even doing in a post mortem about DNSSEC failures?
All joking aside, it occurred to me that the vast majority of internet and even tech users know very little about DNS. For the longest time, I was in the same boat. After having been in a role where it was necessary to understand the record types and deploy DNS configs, I’m quite thankful I learned. Just remember…it’s always DNS.
edit: typo
I've had sales teams pitch me on their authoritative DNS service running heterogenously, although I guessing both partitions gathered config from a single place, but we didn't get that far in the potential customer pipeline.
Of course, just because it's possible doesn't mean they're likely to do it.
That's my way of saying: it is seperate infrastructure
Remember that Cloudflare is probably the second largest DNS Resolver in the world. They aren't just going to tack two IP's onto the same system. The entire system would be completely independent for reasons exactly like what happened today.
It's incredible how far the SW industry gone in the last decades, but the way monitoring is done is the same as in the 90's.
Thinking not correct?
Ironic.
4(0b10) 7:00 ends at 11:02 (4 hr 2 min) on a 4 sum 2x2. And refs to 1.1.1.1 vs 1.0.0.1