Cloudflare outage on June 21, 2022
blog.cloudflare.com
blog.cloudflare.com
It was very pleasant to find the cloudflare status page pointing me to the issue right away (minutes after our alerts triggered), even though I couldn't replicate the issue myself yet.
I wish more companies would take note of the transparency and sense of urgency on updating their status page. (Looking at you Azure)
Looking at you Twilio...
Me and a bunch of my colleagues were all able to access it (different ISPs) on the east coast.
Like the post-mortem says, they will put mitigations in place, but this is something every network admin has to implement bespoke after learning the hard way that the default management approach is dangerous.
I’ve personally watched admins make routing changes where any error would cut them off from the device they are managing and prevent them from rolling it back — pretty much what happened here.
What should be the default on every networking device is a two-stage commit where the second stage requires a new TCP connection.
Many devices still rely on “not saving” the configuration, with a power cycle as the rollback to the previous saved state. This is a great way to turn a small outage into a big one.
This style of device management may have been okay for small office routers where you can just walk into the “server closet” to flip the switch. It was okay in the era when device firmware was measured in kilobytes and boot times in single digit seconds.
Globally distributed backbone routers are an entirely different scenario but the manufacturers use the same outdated management concepts!
(I have seen some small improvements in this space, such as devices now keeping a history of config files by default instead of a single current-state file only.)
Alternatively some sort of watchdog timer would be a great addition (e.g. rollback within X minutes if the changes are not confirmed).
I’m saying this is the problem — the device configuration approach should be safe by default.
Always good. The system built into display settings ("click yes if you can read this, or the change will be reverted in 15 seconds") has saved me a number of times. No reason not to apply that to other settings where the data channel is the same as the control channel.
OpenWRT router distribution has had this for years, it's amazing! (As is OpenWRT)
OpenWRT also has SQM CAKE which saved my sanity on parents DSL connection for years. As far as congestion control and bandwidth sharing goes, nothing else compares
(Which is an issue with the /automation/ defaults. I've learned enough to do commit confirm, but by default ansible does a hard commit.)
Looking at the hourly chart in Google Analytics (compared to the previous day) there isn't even a blip during this outage.
So for all the advantages we get from Cloudflare (caching, WAF, security [our WP admin is secured with Cloudflare Teams], redirects, page rules, etc) I'll take these minor outages that make HN go apeshit.
Of course it helped that most our traffic is from the US and this happened when it did but in the past week alone we served over 180 countries which Cloudflare helps make sure is nice and fast :D
Many thanks
(Or DM the puppet email in my profile)
I did get an alert from Uptime Robot but when I checked everything was fine and so I thought it was a false positive.
Ouch
Now I'm remembering the story of how, when a certain blue website fell off the Internet for a day ~a decade and a half ago (due to some slightly broken database migration logic), out-of-band access boiled down to who was still logged in (!): https://rachelbythebay.com/w/2019/01/20/quiet/
Amusingly when things went wrong again last year, it was BGP's fault (is this the hyperscale equivalent of "it's always DNS" or something?). Engineers (with adequate credentials) had to actually drive to the datacenter haha.
I would be very interested to hear more about how the break-glass process worked.
Imagine the first eng team did a forward movement revert that corrected the issue and had a new head bits that gets deployed, where shortly after another eng fires off the second process type and tells the system to pull back to the last revision (which is now the bad revision as it was just replaced with fresher deploy bits).
Having two revert processes in the toolkit and maybe a few disperse teams working to revert the issue without tight communication leads to this issue.
I think this is more likely the basis issue vs a bad merge (I assume that the root cause was broadcasted wide and large to anyone making a merge)
I think it's a testament to their Ops/Incident response teams and internal processes, it builds confidence in their ability to respond quickly when something does go wrong. Incredible work!
Don't companies have a fiduciary duty to calculate things; the reason for doing something actually cannot just be that it's a nice thing to do? Not down to the word, but at least the general decision to be this way?
Additionally, these post-mittens work for Cloudflare because they have a great reputation and good uptime. If this were happening daily or weekly, it would be a warning sign to customers.
It’s a strategy other companies could adopt, but to do it effectively requires changes all across the organization.
1. On the buy side, crappy big companies with procurement etc. may not have “actual engineers” deciding things. Maybe for something like Cloudflare that’s likely to sit within an actual technical team’s mandate
2. Not 100% foolproof but if a startup is selling its tool and has an uptime page and details of all its downtime and as a prospective customer you go and see that they have 3-4 hour downtimes once or twice a month, it should raise alarm bells.
I think their calculation (to the extent you can call it that) is that in the interest of PR and damage control, it’s better to get a thorough postmortem out quickly to stem the bleeding and keep people like us from going “I can’t wait to hear what happened at Cloudflare” for a week. Now we know, the customers have an explanation, and this bad news cycle has a higher chance of ending quickly.
Incidents - yes. But why would a post-mortem turn someone off? The incident happened regardless. Do you think anyone would be more likely turned off by reading how they solved it / plan to prevent it on the future than by silence?
IMO it is a slippery slope to see this as opportunity too strongly. Sure, doing the right thing may be net beneficial to the business in the long run...but the $RIGHT_THING should be done first and foremost because it's the right thing.
But once it happens things will change and, to be honest, likely for the worse.
edit> fix typo
Succession?
Doing incident response (both outage and security), the tactical fixes for a specific problem are usually pretty easy. We can fix a bug, or change this specific plan to avoid the problem. The search for conditions that allowed the incident to occur can be alot more time consuming, and most organizations I've worked for are happy to make a couple tactical changes and move on.
There is not. From here there's an ongoing process with a formal post-mortem, all sorts of tickets tracking work to prevent further reoccurrence. This post is just the beginning internally.
> This was delayed as network engineers walked over each other's changes, reverting the previous reverts, causing the problem to re-appear sporadically.
They are running as fast as they can and this extended the incident. There is a “slow is smooth, smooth is fast” lesson in here. I’d rather have a team that takes a day to put up the blog post, but doesn’t unnecessarily extend downtime because they are sprinting.
At around the same time, I was watching HBO’s The Wire. One scene had a high-profile shooting with a frantic police response; people were running everywhere trying to help. The lead commander on the scene gave this instruction to his sergeant: “Slow this thing down to a crawl. Give these bastards no chance to fuck up in a meaningful way.”
And then it hit me. That’s exactly how they ran the calls. I asked them about this, and they said absolutely – human nature, when things are broken badly, is biased towards action. You want to try to make a change, to reboot a system, to do something to hopefully make things better. But often if you were not careful, you risk losing information about the outage you are in. Best case scenario you luck into fixing the problem and don’t know how. Worst case? You’ve changed the state of an already broken system, and have done nothing but add more variables to unwind.
So now, every time I’m on an outage call, I try to do what Wes and Tim and Major Rawls would all do: I take control, pump the brakes, and make sure that we are capturing enough information about the current state that we don’t confuse ourselves further.
And with nowhere near the detail level of what was presented here. Typically lots of sweeping generalizations that don't tell you much about what happened, or give you any confidence they really know what happened or have the right fix in place.
The sheer level of technical competence of your engineering team continues to astound me. (Yes, they made a mistake and didn't catch an error in the diff. But your response process went exactly as it should, and your postmortem is excellent.) I couldn't even begin to think about designing or implementing something of this complexity, much less being able to explain it to a layperson after a failure. It is really impressive, and I hope you will continue to do so into the future!
Most of the companies I've worked for unfortunately don't use your services, but I've always been a staunch advocate and converted a few. Maybe the higher-ups only see downtime and name recognition (i.e. you're not Amazon), but for what it's worth, us devs down the ladder definitely notice your transparency and communications, and it means the world. I've learned to structure my own postmortems after yours, and it's really aided in internal communications.
Thank you again. I can't wait for the day I get to work in a fully-Cloudflare stack :)
Not as transparent as "post it on the internet" but at least better than the usual hand wavey bullshit
It should revert as a failsafe if not confirmed within X minutes.
> Primarily, we will be concentrating on automation improvements ... and provide an automated “commit-confirm” rollback.
$ apply new rules; sleep 10; apply original rules
If your ssh access was still working and various sites were still up during that 10sec you were probably good to go - or at least you hadn't shut yourself out.
I'm sure there were and are better ways of doing it, but it was simple enough and worked for us.
junipers allows for instance, one to do the command commit confirmed, which will apply the configuration, and revert back to the previous version if one does not acknowledge this command within a predifined time. this prevents permanent lockout out of a system.
- Cloudflare 2022 (this one)
- Facebook 2021: https://news.ycombinator.com/item?id=28752131 - this one probably had the single biggest impact, since engineers got locked out of their systems, which made the fixing part look like a sci-fi movie
- (Indirectly caused by BGP: Cloudflare 2020: https://blog.cloudflare.com/cloudflare-outage-on-july-17-202...)
- Google Cloud 2020: https://www.theregister.com/2020/12/16/google_europe_outage/
- IBM Cloud 2020: https://www.bleepingcomputer.com/news/technology/ibm-cloud-g...
- Cloudflare 2019: https://news.ycombinator.com/item?id=20262214
- Amazon 2018: https://www.techtarget.com/searchsecurity/news/252439945/BGP...
- AWS: https://www.thousandeyes.com/blog/route-leak-causes-amazon-a... (2015)
- Youtube: https://www.infoworld.com/article/2648947/youtube-outage-und... (2008)
And then there are incidents caused by hijacking: https://en.wikipedia.org/wiki/BGP_hijacking#:~:text=end%20us...
Some more:
- Google 2016, configuration management bug/BGP: https://status.cloud.google.com/incident/compute/16007
- Valve 2015: https://www.thousandeyes.com/blog/steam-outage-monitor-data-...
- Cloudflare 2013: https://blog.cloudflare.com/todays-outage-post-mortem-82515/
It just seems that most of these are local enough and the Internet resilient enough that they don't cause global issues. Maybe the exception would be AWS us-east-1 outages :-)
Sounds like the same happened here:
"Due to this withdrawal, Cloudflare engineers experienced added difficulty in reaching the affected locations to revert the problematic change. We have backup procedures for handling such an event and used them to take control of the affected locations."
But Cloudflare had sufficient backup connectivity to fix it. I'm curious how Cloudflare does that today-- the solution long ago was always a modem on an auxiliary port.
Also lets face it - the utility of a trusted security guard/staff with an old fashioned physical key is pretty hard to screw up!
Now you can use mobile Internet (4G/5G)
I'm surprised more places don't implement a "click here to confirm changes or it'll be rolled back in 5 minutes" like all those monitor settings dialogues
BGP is just a tool, it would be something else to do the same purpose.
The main difference between BGP and all other tools is that if you mess up BGP, you've done a very visible thing because BGP underpins how we get to each other's networks. But it's not a sign of BGP being fragile, just very important.
Another canonical example is C++. Some tools make it easy to blow your leg off. Some tools provide safety mechanisms to stop the saw from cutting off your finger.
I'm not blaming BGP, since it prevents far more outages than it causes, but BGP-based outages have been a thing since its beginning. And any other protocol would have outages too - BGP just happens to be the protocol being used.
Or you could go the opposite direction and risk turning something like this into a PR death spiral.
At the end of the day, everybody makes mistakes and that's okay. Everybody else also know that everybody makes mistakes. So why not accept it?
I really don't get what's wrong with accepting mistakes, learning from them, and moving on.
Some people really struggle with this (myself included) but I think it's one of the easiest "power ups" you can use in business and in life. The key is that you have to actually follow through on the "learning from them" clause.
https://appleinsider.com/articles/12/09/28/apple-ceo-tim-coo...
This one line will forever cement exactly how bad Apple Maps' release was. Thanks Mike Judge!
In addition to trying to de-Googlify my life, there was also an occurance where Google Maps literally tried to kill me: at an intersection that connects into a highway it guided me to drive straight into the opposite direction to a highway, straight onto the coming cars at 140km/h. I've quit Google Maps right there and never used it again.
Apparently so terrible that Apple apologized, perhaps for the first (and last) time for something.
If you forgive my prying, was this an implementation issue with the maintenance plan (operator or tooling error), a fundamental issue with the soundness of the plan as it stood, or an unexpected outcome from how the validated and prepared changes interacted with the system?
I imagine that an outage of this scope wasn’t foreseen in the development of the maintenance & rollback plan of the work.
Everybody at one time experiences the dreaded REJECT not being at the end of the rule stack but just too early.
Kudos to CF for such a good explanation of what caused the issue.
Even better if the tool was syntax aware so it could highlight the different types of rules in unique colors.
https://www.reddit.com/r/ProgrammerHumor/comments/vh9peo/jus...
Cloudflare is great, and I would never move away from it. But from a business continuity standpoint, is there a fallback approach that we should be prepared for during such cases?
One crude approach we were discussing is during an outage we could change the NS records in the registrar to point to for eg. Google Cloud DNS which would already be in sync in terms of the DNS records it has.
If not, a manual swap at the registrar level would be good enough.
I should also mention this approach sort of breaks with Cloudflare's proxied records which dynamically assign anycast IPs for records placed on their CDN. So if using this approach the failover NS provider would probably need to also use a different CDN, preferably one that just gives you a CNAME.
Networking talent is kind of hard to find and if you learn that your chances of employment get pretty high.
They've got some Docker examples in the README.
You can simulate arbitrarily large networks and internetworks with this, provided you have the hardware to run a large enough number of virtual appliances, but they are pretty lightweight.
Not trying to take the light away from the outage, the outage was bad. But the relative quickness to recovery is pretty impressive, in my opinion. Sounds like they could have recovered even quicker if not for a bit of toe stepping that happened.
You broke half the internet: BGP You broke half of your company's ability to access the internet: DNS
And a dry-run: "a Change Request ticket was created, which includes a dry-run of the change, as well as a stepped rollout procedure."
And a Peer review: "Before it was allowed to go out, it was also peer reviewed by multiple engineers. "
I would doubt the expertise of tech guys of cloudflare, reviewing the change. And there was a dry-run.
But is it really OK to apply the change to a spine network which would affect 50% network traffic? Just out of peer review and a dry run? No green/blue, no gray release, maybe these are not proper for a small change here. But this "small" change really got big affect. I thougt it was worth it.
And from my shallow experience, the dry-run would always have do nothing to the env. It is dry-run anyway.
And at last the three lines are found out. So I wonder how did this re-order happen? And why?
With these tiny changes, there should be some mechanism to verify their correctness, not just review and dry-run.
The specific network locations that were impacted by this change were amongst the last to see the change rolled out. One deficiency in our deployment strategy (which we will correct) is that no network locations in the affected "MCP" configuration received the change early in our rollout process. If that had been the case, we would have found the problem much earlier and the incident's impact would have been much reduced.
Another massive gap is the rollback: 6:58 – 7:42 – 44 minutes! What exactly was going on and why did it take so long? What were those back-up procedures mentioned briefly? Why engineers where stepping on each other toes? What's the story with reverting reverts?
Adding more automation, tests and fixing that specific ordering issue of course is an improvement. But that adds more complexity and any automation ultimately will fail some day.
Technical details are all appreciated. But it is going to be something else next time. Would be great to learn more about human interactions. That's where the resilience of a socio-technical system happened and I bet there is some room for improvement there.
Like, does Cloudflare have an emergency procedure for escalation? What does that look like? How does the CTO get woken up in the middle of the night? How to get in touch with critical and most important engineers? Who noticed Cloudflare down first? How do quick decisions get made and decided? Do people get on a giant zoom call? Or emails going around? What if they can't get hold of the most important people that can flip switches? Do they have a control room like the movies? CTO looking over the shoulder calling "Affirmative, apply the fix." followed by a progress bar painfully moving towards completion.
https://sre.google/resources/book-update/managing-incidents/ is Google focused, but our flavor of incident response is not too far off.
Slack: "@here need to connect to <long list of devices> to rollback change asap"
It sounds like it's a key architectural part of the system that "[...] convert all of our busiest locations to a more flexible and resilient architecture."
25 year experience and it's always the things that are supposed to make us "more flexible" and "more resilient" or robust/stable/safer <keyword> that ends up royally f'ing us where the light don't shine.
"In this time, we’ve converted 19 of our data centers to this architecture, internally called Multi-Colo PoP (MCP): Amsterdam, Atlanta, Ashburn, Chicago, Frankfurt, London, Los Angeles, Madrid, Manchester, Miami, Milan, Mumbai, Newark, Osaka, São Paulo, San Jose, Singapore, Sydney, Tokyo."
Is the term MCP synonymous with "tier 1 PoPs" (mentioned elsewhere in other cloudflare blogs from time to time) or are the two terms referring to different things?
Sounds like a good idea!
I also think that having a script that walks through the file and points out any ovibious mistakes would be good to have as well.
I'd say that if you're designing a system which has the potential to disconnect half your customers based on a misconfiguration, then you should spend at least an hour thinking about what sorts of misconfigurations are possible, and how you could prevent or mitigate them.
The cost-benefit analysis of "how likely is it such a mistake would get to production (and what would that cost us)?" vs "how much effort would it take to write and maintain a verifier that prevents this mistake?" should then be fairly easy to estimate with sufficient accuracy.
We did get alarms. Our things partially worked though so CF was not the first thing to check.
Really strange, a coincidence?
I'm not familiar with whatever language this is, but wouldn't such a construct always indicate something was being ignored?
One of their future changes is to include MCPs in their testing environments.
The problem is, we couldn't tell all our client they should change this :(
I buy more NET every time I see posts like this.
<html> <head><title>500 Internal Server Error</title></head> <body bgcolor="white"> <center><h1>500 Internal Server Error</h1></center> <hr><center>nginx</center> </body> </html>
Your DNS host needs to support being able to assign a CNAME record on your root domain to a domain provided by Cloudflare. AWS Route 53 does not let you do this which I imagine is a decent chunk of enterprise clients. AWS only lets you alias records to AWS resources not external domains.
With that said, even with enterprise in this case you would need to go all-in with Cloudlfare's nameservers or run the risk of not having DDoS protection on your root domain (ie. example.com wouldn't be protected but you could protect www.example.com since a CNAME with subdomains is a standard thing).
However it's kind of interesting because an attacker could get the real IP of your root domain's AWS load balancer which is probably the same load balancer used for the `www` version of your site too, but now that they know your load balancer's IP they can completely bypass Cloudflare and go straight to your infrastructure.
I'm pretty sure AWS doesn't let you assign an external domain with their aliases because they want you to pay them for AWS Shield Advanced instead of using Cloudflare because AWS knows getting an enterprise client to change their nameservers and all of their DNS records (potentially dozens of domains and multiple hundreds of records) is kind of a pain. It can be done but it's a friction point.
Here's a somewhat old (2016) but very impressive system at a major ISP for avoiding exactly this: https://www.youtube.com/watch?v=R_vCdGkGeSk
If it was AWS, Akamai, Google Cloud, or any of the other massive providers this comment wouldn't be made. I don't really understand the association between centralisation and CloudFlare, other than it being a Meme.
And yeah, cf is trying to get as much traffic to go through them as possible and add edge services for more opportunities - that's literally their business. Also now r2 with object storage. They're already too big, harmful (as in actually putting people in danger) and untouchable in some ways.
What do you have when you have all DNS going through them, via DoH, and all web requests going through them, if not recentralization?
Sure, they want us to think they give us the freedom to host our web sites anywhere because they're "protected" by them, but that "protection" means we've agreed to recentralize.
It's pretty dismissive to describe something as a meme just because you don't understand it, and either you're pretending to not understand it, or you truly don't.
Look at it this way: If a single company goes down for an hour, and that company going down for an hour causes half the web traffic on the Internet to fail for that hour, what is that if not recentralization?
What I don't understand, or believe, is that they want to be the sole (as in centralised) network for the internet. I don't believe they as a company, or the people running it, want that. They obviously have ambition to be one of the largest networking/cloud providers, and are achieving that.
I don't intend either to dismiss your concerns (which are a legitimate thing to have, centralisation would be very bad), my suggestion with the meme comment is that there is at times a trend to "brigade" on large successful companies in a meme-like way. That isn't to suggest you were.
It's pretty appalling that you are even being downvoted.
sure, transparency is better than "something went wrong, we take this very seriously, sorry." (although the non technical crowd couldn't care less)
only people who dont do anything make no mistakes, but doing such highly impactful changes so quickly (inside one day!) for where 50% of traffic happens seems a huge red flag to me, no matter the procedure and safety valves.
at time of writing no comment has done that except you.
For their part, they handled this very well, and are to be commended (quick fix, quick explanation of failure).
But you also can't help but see that they have a dangerous amount of control over such important systems.