Post Mortem on Cloudflare Control Plane and Analytics Outage
blog.cloudflare.com
blog.cloudflare.com
Honestly, this is rather unprofessional. I understand and support explaining what triggered the event and giving a bit of context, but the focus on your postmortem needs to be on your incident, not your vendor's.
Clearly, a lot went wrong and Flexential needs to do their own postmortem, but Cloudflare doesn't need to make guesses and do it for them, much less publicly.
It might also be an effort to get out in front of the story before someone else does the speculating.
In any case, with at least three parties involved, with multiple interconnected systems… if Cloudflare is going to effectively anticipate this cluster of failure modes in future design decisions, it's reasonable for them to want to know what happened all the way down.
Edit to add: I for one am grateful for the information Cloudflare is sharing.
It's been 2 days. I doubt PGE or Flexential even have root caused it yet, and even if they have, good communication takes time.
You don't throw someone under the bus and smear their name publicly just because they haven't replied for two days, and you certainly don't start speculating on their behalf. That's bad partnership.
You also don't publicly share what "Flexential employees shared with us unofficially" (quote from the article) - what a great way to burn trust with people who probably told you stuff in confidence.
>if Cloudflare is going to effectively anticipate this cluster of failure modes in future design decisions, it's reasonable for them to want to know what happened all the way down.
They can do all of that without smearing people on their company blog. In fact, they can do all of that without even knowing what happened to PGE/Flexential, because per their own admission they were already supposed to be anticipating this, but failed at it. Power outages and data center issues are a known thing, and is exactly why HA exists. HA which Cloudflare failed at. This post-mortem should be almost entirely about that failure rather than speculation about a power outage.
1. When you’re paying them the kind of money I imagine they’re paying and they don’t reply for 2 days, yea that’s crazy if true. I’d expect a client of this size could take to an executive on their personal number.
2. Telling the facts as you know them to be especially regarding very poor communication isn’t a smear.
Publicly casting blame based on speculation isn't something you do to someone that you want to have a good working relationship with, no matter how much money you pay them.
What are you disagreeing with OP ?
He is talking about how to behave if you continue the relationship not whether to continue it .
2 days is outrageous here, I have to imagine whoever thinks that is acceptable is approaching this from the perspective of a company whose downtime doesn't affect profits.
UPS failing early sounds like it may be a battery maintenance issue.
What???? We have 4 hour boots on the ground support with Supermicro and that's a few thousand dollars a year lol.
That doesn't make any sense for a customer as big as CF.
Two days without support communications would be a long time, but my original comment about the two day period is about the post-mortem. It's totally reasonable IMO for a company to take longer than two days to gather enough information to correctly communicate a post-mortem for an issue like this, and IMO its unreasonable for CF to try to shame Flexential for that.
> Due to circumstances beyond our control the DC lost all power. We are still working with our vendors to investigate the cause. While such a failure should not have been possible, our systems are supposed to tolerate a complete loss of a DC.
Here's what happened, here's what went wrong, here's what we did wrong, here's our plans to avoid it happening again
Seems like a standard post mortem tbh
The ordering that they list the mistakes would be a fair point to make though, in my opinion. They hinted at a mistake they made in their summary, but don't actually tell us point blank what it was until they tell us all the mistakes that their vendor made. I'd argue that was either done to make us feel some empathy for Cloudflare as being victims of the vendor's mistakes, misleading us somewhat. Or it was done that way because it was genuinely embarrassing for the author to write and subconsciously they want us to feel some empathy for them anyway. Or some combination of the two. Either way, I'll grant that I would have preferred to hear what went wrong internally before hearing what went wrong externally.
5th paragraph to the 9th are Cloudflare's "we buggered up" before they get to the power segment. They then continue with the "this is our fault for not being fully HA" after the power bit.
Each to their own, I'm going to read it as a regular old post mortem on this one.
Going into such depths on the 3rd party just shows how embarrassing this is for them.
And yes, that also means the initial event description in what they know so far.
Highly likely there will be another one https://twitter.com/eastdakota/status/1720688383607861442?t=...
And even though Cloudflare did put some of the blame, as it were, on the vendor, the post mortem recognizes that Cloudflare wasn't doing their due diligence on their vendor's maintenance and upkeep to verify that the state of the vendor's equipment is the same as the day they signed on. And that's ignoring a huge focus of the post mortem where they admit guilt at not knowing or not changing the fact that Kafka and Clickhouse were only in that datacenter.
Furthermore, we do not know that Cloudflare didn't get the vendor's blessing to submit that diagram to their post mortem. You're assuming they didn't. But for what it's worth as someone that has worked in datacenters, none of this is all that proprietary. Their business isn't hurt because this came out. This is a fairly standard (and frankly simplified for business folk) diagram of what any decently engineered datacenter building would operate like. There's no magic sauce in here that other datacenter companies are going to steal to put Flexential out of business. If you work for a datacenter company that doesn't already have any of this, you should write a check to Flexential or their electrical engineers for a consultancy.
And finally, the things that Cloudflare speculated on were things like, to paraphrase, "we know that a transformer failed, and we believe that its purpose was to step down the voltage that the utility company was running into the datacenter." Which, if you have basic electrical engineering knowledge, just makes sense. The utility company is delivering 12470 volts, of course that needs to be stepped down, somewhere along the way, probably multiple times, before it ends up coming through the 210 volt rack PDUs. I'm willing to accept that guess in the absence of facts from the vendor while they're still being tight lipped.
However, that's not to say I'm totally satisfied by this post mortem either. I am also interested in hearing what decisions led to them leaving Kafka and Clickhouse in a state of non-redundancy (at least at the datacenter level) or how they could have not known about it. Detail was left out there, for sure.
That said, the existence of the 480V labeled intermediary does suggest they have a 277/480 V outside system, and a 120/208 V rack-side system.
(none of this takes away from the mistakes that were wholly theirs that shouldn't have happened and that they should fix)
> The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04.
> A handful of products did not properly get stood up on our disaster recovery sites. These tended to be newer products where we had not fully implemented and tested a disaster recovery procedure.
So the root cause for the outage was that they relied on a single data center. I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet.
Bah, who cares about such unimportant details, what's important is that ~dev velocity~ was reaaally high right until that moment!
> We were also far too lax about requiring new products and their associated databases to integrate with the high availability cluster. Cloudflare allows multiple teams to innovate quickly. As such, products often take different paths toward their initial alpha. While, over time, our practice is to migrate the backend for these services to our best practices, we did not formally require that before products were declared generally available (GA). That was a mistake as it meant that the redundancy protections we had in place worked inconsistently depending on the product.
Complete and utter management failure. And customers apparently are sold what Cloudflare internally considers to be alpha quality software?
This has been my experience with AWS and GCP as well. Assume anything that's under 3 years old is not really GA quality no matter what they say publicly.
GCP run multi-year betas of services and features, so I'm doubtful there were still things not ironed out for GA. Do you have some examples?
Too strong. A failure certainly, but painting this as the worst possible management failure is kind of silly.
"Cloudflare Reverse Proxies Are Dumping Uninitialized Memory" - https://news.ycombinator.com/item?id=13718752
There’s a lack of awareness there.
It's amazing that they don't have standards that mandate all new systems to use HA from the beginning.
Absolute lack of faith in cloudflare rn.
This is amateur hour stuff.
It's especially egregious that these are new services that were rolled out without HA.
Tbh. As far as I can see, their data plane worked at the edge.
Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush.
Their control plane was bad though. But since most products worked, that's more redundancy than most products.
The proposed solution is simple:
- GA requires to be in the high availability cluster
- test entire DC outages
[1]: https://blog.cloudflare.com/introducing-cloudflare-stream/ [2]: https://www.cloudflare.com/press-releases/2018/cloudflare-st...
It's not just streams, image upload & Logpush.
The data plane ( which I mentioned) had no issues.
It's literally in the title what was affected: "Post Mortem on Cloudflare Control Plane and Analytics Outage"
Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that.
Source: I watched it all happen in the cloudflare discord channel.
If you know anyone that is claiming to be affected on the data plane for the services you mentioned, that would be an interesting one.
Note: I remember emails were also more affected though.
Which was still like ~12+ hours, if we check the status page.
>Eg. The status page mentioned the healthchecks not working, while everything was fine with it. There were just no analytics at that time to confirm that.
What good is a status page that's lying to you? Especially since CF manually updates it, anyway?
>Source: I watched it all happen in the cloudflare discord channel.
Wow, as a business customer I definitely like watching some Discord channel for status updates.
This wasn't about status updates going to discord only.
There is literally a discussion section on the discord, named: #general-discussions
Not everything was clear in the discord too ( eg. The healthchecks were discussed there), that's not something you want to copy-paste in the status updates...
Priority for cloudflare seemed to get everything back up. And what they thought was down, was always mentioned in the status updates.
However, I still fail to see your argument regarding Zero Trust and not being impacted. The status page literally mentioned that the service was recovered on Nov 3, so I don't understand what you mean by:
>The data plane ( which I mentioned) had no issues.
There's literally a section with "Data plane impact" on all over the status page, and ZT is definitely in the earlier ones. And this is given the fact that status updates on Nov 2 were very sparse until power was restored.
What I mentioned, was what I've seen passing by in the channel at the time.
I also saw no incoming help requests for zero trust tbh ( did some community help)
Arguable, it's best to think of the edge as a buffering point in addition to processing. Aggregation has to happen somewhere, and that's where shit hit the fan.
Cloudflare's data lives in the edge and is constantly moving.
The only thing not living in the edge ( as was noticed), is stream, logpush and new image resize requests ( existing ones worked fine) from the data plane
You're being loose in your usage of 'data'. No one is talking about cached copies of an upstream, but you probably are.
Read the post mortem a bit more closely. They explicitly state that the control plane(s) source of truth lives in core, and that logs aggregate back to core for analytics and service ingestion. Think through the implications on that one.
So do they.
75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text.
But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully recover service. This was longer than the outage, and the text just states that too many services were dependent from each other. But I'd wish they go into more detail here why the operation as a whole took that long. Are there any take-aways from the recovery process, too? Or was it really just syncing data from the edges back to the "brain" that took this long?
Also one aspect I am missing here is the lack of communication - especially to Enterprise customers. Cloudflare support was basically radio silent during this outage except for the status page. Realistically, they couldn't do much anyway. But at least any attempt at communication would be appreciated - especially for Enterprise customers, and even more especially after the post-mortem blames Flexential for a lack of communication.
While I like Cloudflare since it's a great product, I think there are still a few more things that should be taken as a conclusion for CF to take away from this incident.
That being said, glad you managed to recover, and thanks for the post-mortem.
This paragraph similarly leaves out juicy details. Exactly what services fail if logging is down? Were they built that way inadvertently? Why did no one notice?
They blame Flexential for lack of communication, but were the first one not saying anything.
DC related updates:
> Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023 - 17:08 UTC
> Identified - Cloudflare is assessing a loss of power impacting data centres while simultaneously failing over services.
We will keep providing regular updates until the issue is resolved, thank you for your patience as we work on mitigating the problem. Nov 02, 2023 - 13:40 UTC
In reality, Cloudflare's support team was essentially completely unavailable on Nov 2, leaving only the status page. And for most of the day, the updates on the status page were very sparse except "we are working on it", and "We are still seeing gradual improvements and working to restore full functionality.".
Yet clearer status updates were only giving starting on Nov 3. However, I still don't think I heard anything from support or a CSM during that time.
1) Were you affected on the data plane? Which product?
As far as I can tell, while the outage was in the core dc's. The impact was minor.
2) Both examples were exactly from 2 November. Not 3 November.
3) What method of support did you try? I thought that their support was impacted ( email?).
The status page explicitly mentioned to get in contact with your account manager for some config changes on some products, if you wanted changes.
4) I have never heard of Enterprise customers being contacted by a cloud company during an outage.
Which company does that? Do you have an example?
5) I would think it's absolutely a nogo to contact every preemptively Enterprise customer with: "hey, the product works, but if you change xyz, atm that doesn't.".
Since most customers weren't affected and some others were minorly impacted.
There is not a single cloud company that does that.
Feel free to correct me if I'm wrong...
ssl for saas -> custom hostnames are not working for new domains or changes to current ones. also page rules -> redirects are not working for new rules or changes to current rules. which are game-stoppers for our business.
we contacted via enterprise email support + ccing our managers and assigned engineers.
first they try to tell us product is working and sending us some details how to do that,this etc, after a couple of hours later they understand the issue is bigger than they thought and they said "the product is affected by api outage".
then in another email we asked them when this can be solved but only answer we got is "please follow status page for the updates".
and after a day, ssl for saas & ssl services took their places on status page. for a day nobody notices if it's working or not except customers.
so as we understand these emails even the team internally haven't got any idea what is working and what is not!
No, but we needed to make urgent changes.
>2) Both examples were exactly from 2 November. Not 3 November.
Both messages contain no clear messages about remediation and co. They also didn't state clearly which products were failed over. I noticed that at this point I could at least login to the dashboard, but most stuff was still severely broken, and I had no idea whether changes with the few semi-functional components were actually applied or not.
Updates to single products with a more clear status were given only at the end of November 2nd (UTC).
(Also one of the message states data centres - not just data center. Not sure what happened there).
>3) What method of support did you try? I thought that their support was impacted ( email?).
Emergency line + contacting our CSM. The emergency line was shut down and replaced with voice mail (WTF?), and our CSM did not reply at all (or the message somehow made it to the wrong person, I'll find out next week, I guess).
So in our case, the communication was essentially non-existent, even though I raised a support case (or wanted to).
>4) I have never heard of Enterprise customers being contacted by a cloud company during an outage. Which company does that? Do you have an example?
I can remember of Datadog reaching out to us for their 2023-03-08 incident. Not sure if it was just our CSM being nice or someone did a support request on another communication channel, but looking back in history that came without asking + the post mortem. Same case when stuff happens such as vulnerabilities in one of their packages, they reach out to us proactively and notify us.
To be fair, this is a bit of a wishlist and definitely not necessary for a 30 minutes hickup, but for a 2 day outage... I don't know.
At the bare minimum, I'd expect at least their support team to be replying and not shutting down the communication channels.
>5) I would think it's absolutely a nogo to contact every preemptively Enterprise customer with: "hey, the product works, but if you change xyz, atm that doesn't.".
I don't know... At least at the time I raise an urgent support case about an issue, I expect to be kept up-to-date.
> Since most customers weren't affected and some others were minorly impacted.
What does it mean they were not affected? Yes, their core service was still functioning (thank god - after all they advertise a 100% (!) SLA on that), but you can see on same Discord channel you mentioned people failing to renew TLS certificates, people couldn't make Vercel deployments and more. So it did affect quite a bunch of downstream customers in their products, and they might also sell SLAs to their customers...
I cannot really comment on whether that just affected us, or if other customers had better support experiences here.
But I expect better in terms of communication here. Doesn't have to be as outreaching as I did in my last message, but stuff like shutting down the emergency line and not giving any comment is not really acceptable for an Enterprise contract.
We are a service provider in ( mostly) Europe.
Our policy ( playbook) in case of an issue is updating the status page as quick as possible and customers can subscribe on RSS.
There was one issue in the past where we wanted to inform the clients. But it's not easy, as only some were impacted and we decided against it.
5 minutes later ( it was out of our hands) it was solved...
Our playbook is too update the status page as soon as possible to inform the clients something is up and we are aware.
There shouldn't be too much info on it, since sometimes you just aren't 100% sure about what's exactly going on.
We also decided that we want provide durations on it, since you then create a commitment that's possibly dependent on external factors.
Tbh. I can completely understand the approach from Cloudflare here. With an issue, support is overwhelmed. That's why you use the status page ASAP.
Technical details happen in the post-mortem. When we can be sure if any data is lost ( normally, there is nothing lost though, but it's possible we need to requeue some actions)
=> this is when we can contact our clients and brought up to date.
Depending on the SLA it's included or eg. Is paid extra ( in a lot of times, an external provider fails and we can fix something from our end, eg. Resending some data)
Instead of defending what was done and calling that good enough, Cloudflare should use this as an opportunity to commit to reevaluating the strategy for customer outreach during major service failures. If that's what Cloudflare expects from its service providers, that's what Cloudflare should provide to its customers.
You want Cloudflare to update every customer for an issue that they probably aren't affected with ( except when changing things) ?
Who even does that when you've got so many customers?
That's exactly what why the status page is there:
https://www.cloudflarestatus.com/
The DC obviously didn't have any means to update their customers.
Even those that had complete outages.
We were affected but it’s blog posts like these that make me never want to move away. Everyone makes mistakes. Everyone has bad days. It’s how you react afterwards that makes the difference.
The issue is when you start having bad days every other day though. We use and depend on CloudFlare Images heavily, it has now been down more than 67 hours over the last 30 days (22h on October 9th, 42h Nov 2 - Nov 4 and a sprinkle of ~hour long outages in between). That's 90.6% availability over the last month.
Transparency is a great differentiator between providers that are fighting in the 99.9% availability range, but when you are hanging on for dear life to stay above the one 9 availability, it doesn't matter.
Similar story for GCP.
All three of them had decades of institutional knowledge and procedures in place around running big services by the time Cloudflare was founded.
No, even at the onset AWS was an entirely-from-the-ground-up build. The only thing it could even be argued to sit on top of was the extremely crufty VMs and physical loadbalancers from the original Prod at that point, and those things were not doing anybody any favors.
I really appreciate that they're going to fix the process errors here. But as they suggested, there's a tension between moving fast and being sure. This is typically managed like the weather, buying rain jackets afterwards (not optimal). I'd be curious to see how they can make reliability part of the culture without tying development up in process.
Perhaps they can model the system in software, then use traffic analytics to validate their models. If they can lower the cost of reliability experiments by doing virtual experiments, they might be able to catch more before roll-out.
Disagree completely, it's the frank detail that makes me trust their story.
Has Flexential provided a similarly detailed, public root cause analysis? If so, maybe we can refer to it. If not, how do you expect us to read it?
Imagine if the literal power company failed, and took days to tell people what was going on. You can see why people are reading the postmortem that exists, rather than the one that doesn't.
They do both. They stated what their problem was and they stated their due diligence in picking a DC
> While the PDX-04’s design was certified Tier III before construction and is expected to provide high availability SLAs
They said the core issue: innovating fast, which led to not requiring in the high availability cluster.
Which is also a fix.
From cloudflare 's POV, part of what made it originally worse, is the lack of communication by the DC.
Which is an issue, if you want to inform clients.
Very worrying is they start by stating their intended design:
> Cloudflare's control plane and analytics systems run primarily on servers in three data centers around Hillsboro, Oregon
You need way more geographic dispersion than that, this control pane is used by people across the world. We are still on the intended design, not the flawed implementation by the way, which is wild to me.
> This is a system design that we began implementing four years ago. While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster.
I don't understand why this would ever be done in this way. If Cloudflare is making a new product for consumers shouldn't redundant design be at the forefront here? I am surprised that it was even an option. For the record I do use Cloudflare for certain systems and I use it because I assume it has great failovers if events like this occur making me not have to worry about these eventualities, but now I will be reconsidering this, how do I actually know my cloudflare workers are safe from these design decisions?
> When services were turned up there, we experienced a thundering herd problem where the API calls that had been failing overwhelmed our services.
Yeh I'll bet, its because Cloudflares core design is not redundant.
Really disappointed in this blog post trying to shift the blame to Flexential when this slapdash architecture should be the main problem on show. As a customer I don't care if Flexential disappears in an earthquake tomorrow, I expect Cloudflare to handle it gracefully.
Is placing the entirety of such a critical cluster in a known earthquake and tsunami zone a good idea? It looks like their disaster recovery to Europe didn't really work either...
By the way, assuming an ideal control plane (in contrast to data plane) would be 3 DCs at a distance of about 20-40 miles, are there any mitigation techniques so that a seismic event which destroys a single DC doesn't also sever the comms between the remaining two?
That is a painful lesson, but unless you are physically powering off the dc or at least disconnecting the network from the outside world you are not testing a real disaster.
You can point fingers at the facility operators, but at the end of the day you have to be able to recover from a dc going completely offline and maybe never coming back. Mother Nature may wipe it off the face of the earth.
Business was mostly running as usual.
The OVH outage was immediate downtime.
I liked this - the human element is underemphasised often in these kinds of reports, and trying to fix a major outage while overly tired is only going to add avoidable mistakes.
I don’t know how it would work for an org of Cloudflare’s size, but I know we have plans for a significant outage for staff to work/sleep in shifts, to try to avoid that problem as well.
Issue there is that you need a way to hand over the current state of the outage to new staff as they wake up/come online.
Like Mike Tyson says, everyone has a plan until they get punched in the face.
If you don't do that, you're still going to be scrambling.
"Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04."
Bingo, there we have it.
> When services were turned up there, we experienced a thundering herd problem where the API calls that had been failing overwhelmed our services. We implemented rate limits to get the request volume under control.
This seems not to be mentioned in the bullet points at the end of the text (which are otherwise reasonable).
And now I’m curious—how do you design cold failover when the system is complex enough to be metastable[1] and you can’t afford to test it on live traffic? I can guess which techniques you could use to build it, it’s the design and testing part (knowing the techniques actually work in your situation) that’s the problem.
One other thing that seems to have gone completely unmentioned:
> Beginning on Thursday, November 2, 2023, at 11:43 UTC Cloudflare's control plane and analytics services experienced an outage. [... W]e made the call at 13:40 UTC to fail over to Cloudflare's disaster recovery sites located in Europe.
Why did the decision take so long? I can imagine it can’t be made lightly, but two hours seems like too much hesitation, even if there was an expectation that power would be restored imminently for most of that time. There has to be a (predetermined?) point when you hit the switch regardless of any promises. Was it really set that far?
[1] http://charap.co/metastable-failures-in-distributed-systems/
The problem here, which is 100% Cloudflare, is that their systems were not resilient across geography.
> While, over time, our practice is to migrate the backend for these services to our best practices, we did not formally require that before products were declared generally available (GA).
I really like the model where a single team in a company, with Product + Dev, can quickly ship, iterate on a new product, and prove market demand without going through layers and layers of internal bureaucracy (Ops/Infra, Security, Privacy/Legal, Finance approval for production-scale), with the main stipulation being that such work is marked as alpha/beta/preview, and only going through the layers of internal bureaucracy once it's ready to go GA. But most companies really struggle with this, especially with ensuring that customers are never exposed to a/b/p software by default, requiring opt-in from the customer, allowing the customer to easily opt-out, and ensuring that using a/b/p software never endangers GA features they depend on. Building that out, if it's even on a company's internal Platform/DevX backlog, is usually super far down as a "wishlist" item. So I'm super interested to see what Cloudflare can build here and whether that can ever get exposed as part of their public Product portfolio as well.
> We need to use the distributed systems products that we make available to all our customers for all our services so they continue to function mostly as normal even if our core facilities are disrupted.
Super excited to see this. Cloudflare Workers is still too much of an "edge" platform and not a "main datacenter" platform, at least because D1 is still in beta and even if it wasn't, Postgres is far more feature-ful, and that pulls more software into a traditional single-datacenter model. So if Cloudflare can really succeed at this, then it'll be a much stronger statement in favor of building out software in an edge-only model.
Between the Pages outage and the API outage happening in one week, I was considering selling my NET stock, but reading a postmortem like this reminds me why I invested in NET in the first place. Thanks Matt.
Speaking from personal experience, what you're claiming as 'good', for CF meant SRE- usually core, but edge also suffered- got stuck with trying to fix a fundamentally broken design that was known faulty- and called faulty repeatedly- but forced through.
Nothing about this is desirable or will end well.
This reckoning was known and raised by multiple SRE near a decade before this occurred, and there were multiple near misses in the last few years that were ignored.
The part that's probably funny- and painful- for ex-CF SRE is that the company will do a hard pivot and try to rectify this mess. It's always harder to fix after, rather than building for, and they've ignored this for a long while.
>> Super excited to see this. Cloudflare Workers is still too much of an "edge" platform and not a "main datacenter" platform, at least because D1 is still in beta and even if it wasn't, Postgres is far more feature-ful, and that pulls more software into a traditional single-datacenter model. So if Cloudflare can really succeed at this, then it'll be a much stronger statement in favor of building out software in an edge-only model.
On the other when a company dogfoods its own products you end up in a dependency hell like AWS apparently is in where a single Lambda cell hitting full capacity in us-east-1 breaks many services in all regions.
I'm sure there is a right way to manage end to end dependencies for 100% of your services past, present, and future but increasingly I'm of the opinion that it's not possible in our economic system to dedicate enough resources to maintain such a dependency mapping system since that takes away developer time from customer facing products that show up in the bottom line. You just limp along and hope that nothing happens that takes out your whole product.
Maybe companies whose core business is a money printing machine (ads) can dedicate people to it but companies whose core business is tech probably don't have the spare cash.
Security is what keeps a single service getting breached from causing the whole company to get breached.
> Privacy/Legal
Cloudflare doesn't get indemnification from the law just because a customer agrees to mutually break the law.
I suspect that in end it's just easier to put everything into single declarative formal verification system and see if new change to the system passes, transition between configurations passes etc.
If the three data centers are all around Hillsboro, Oregon, an earthquake could probably take out all three simultaneously.
Is it west of I5?
(yes)
Oh yeah, they all gone.
Cascadia Subduction Zone - https://pnsn.org/outreach/earthquakesources/csz
Between that, and being ~50 miles inland - I'd say there's ~zero threat of Cascadia quakes or tsunamis directly knocking out those DC's. (Yeah, larger-scale infrastructure and social order could still be killers.)
OTOH - Mt. St. Helens is about 60 miles NNE of Hillsboro. If that really went boom, and the wind was right...how many cm's of dry volcanic ash can the roofs of those DC's bear? What if rain wets that ash? How about their HVAC systems' filters?
https://www.oregon.gov/oem/Documents/Cascadia_Rising_Exercis...
50% of roads and near 75% of bridges damaged on the west coast and the I5 corridor.
Refer to PDF page #93 where over 70% of power generation is highly damaged on the I5 corridor and 60% in the coastal areas with 0% undamaged.
Highly damaged - "Extensive damage to generation plants, substations, and buildings. Repairs are needed to regain functionality. Restoring power to meet 90% of demand may take months to one year."
"In the immediate aftermath of the earthquake, cities within 100 miles of the Pacific coastline may experience partial or complete blackout. Seventy percent of the electric facilities in the I-5 corridor may suffer considerable damage to generation plants, and many distribution circuits and substations may fail, resulting in a loss of over half of the systems load capacity (see Table 22). Most electrical power assets on the coast may suffer damage severe enough as to render the equipment and structures irreparable"
The two big problems I'd see would be (1) Social Order and (2) Internet Connectivity. DC's are not fortresses, and internet backbone fibers/routers/etc. are distributed & kinda fragile.
*After all the large-scale power outages & near-outages of recent decades, Cloudflare has no excuse if they lack really-good backup generators at critical facilities. And with their size, Cloudflare must support enough "critical during major disaster" internet services to actually get such generators.
They're just going to straight up lie like that? We definitely weren't able to get "traffic through [their] network" through the outage at many different random points.
So if the CF team is under the impression traffic was not impacted, dig deeper.
Actually, this is the CF version, maybe Flexential will come out with a different one.
BTW, if you design a system to survive a DC failure, you cannot blame the DC failure.
Obviously I'm joking because they are blaming an external company (Flexential) to which they are surely paying big money for the DC space.
I'm surprised that CF are renting space in colocation facilities. I would have expected a business of their size to have their own DCs. Is this common practice for cloud providers?
Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane.
I have no idea how many DCs they have or operate in. Where does "300" come from?
> Colo is much more flexible, cheaper and quicker to start. Definitely since they sit close to the end-user on the data plane.
I understand that, but it has the disadvantage of reduced control and observability - particularly in the event of an outage such as that described in the blog post.
I kind of assumed that top-tier cloud platforms like AWS/Azure/GCP operate out of dedicated DCs, and that CF are similar because of their well-known scale of operations. Since my original comment has been downvoted†, someone presumably thinks this it was a naive or trivial question - although I don't understand why.
(† I don't much care about downvotes, but I do take them to be a signal.)
They want DC's close to every big city. I think most of us knew that they can't launch > 300 DC's in such a short amount of time.
The many amount of DC's is mentioned a lot ( social networks, blogs, here).
There is a distinction between eg. AWS / Azure / ... Which work with a couple of big DC's, while cloudflare operates more spread across more locations.
You're comment did made me realize it may may not be that clear from an outsider viewpoint though ( fyi, I'm an outsider too)
Google for one has both. Some GCP regions [0] are in colos, while others are in places where we already had datacenters [1]. We also use colo facilities for peering (and bandwidth offload + connection termination).
I'm under the impression that most AWS Cloudfront locations are also in colo facilities.
And here we are. My trust in them has hit zero.
I feel like a lot of people in this thread are commenting under the impression that all of Cloudflare was down for 24 hours when in reality I wouldn't be surprised if a lot of customers were unaffected and unaware of the incident.
I wouldn't even have known of the outage had it not been for HN..
Cloudflare outage – 24 hours now - https://news.ycombinator.com/item?id=38112515
Cloudflare Dashboard Logins Failing - https://news.ycombinator.com/item?id=38112230
Ask HN: Cloudflare Workers are down? - https://news.ycombinator.com/item?id=38074906
Cloudflare API, dashboard, tunnels down - https://news.ycombinator.com/item?id=38014582
Cloudflare Intermittent API Failures for Cloudflare Pages, Workers and Images - https://news.ycombinator.com/item?id=37819045
Cloudflare Issues with 1.1.1.1 public resolver and WARP - https://news.ycombinator.com/item?id=37762731
Cloudflare – Network Performance Issues - https://news.ycombinator.com/item?id=37604609
Cloudflare Issues Passing Challenge Pages - https://news.ycombinator.com/item?id=37336743
Even honest engineers cannot foresee the exact cascading consequences effects of such outages. Sales reps are not paid to be either competent on such issues nor to be honest.
The DC almost certainly advertises the redundant power supplies, generator backups and battery failover in order to get the customers. But probably doesn't do the legwork or spend the money to make those things truly reliable. It's a bit like having automated backups - but never testing them and discovering they're empty when they're really needed.
The TLDR is that CF runs in multiple data centers, one went down, and the services that depend on it went down with it.
The interesting question would be why those services did depend on a single data center.
They are pretty vague about it
Cloudflare allows multiple teams to innovate quickly. As such,
products often take different paths toward their initial alpha.
If I was the CEO, I would look into the specific decisions of the engineers and why they decided to make services depend on just one data center. That would make an interesting blog post to me.Designing a highly available system and building a company fast leads to interesting tradeoffs. The details would be interesting.
In my experience, no engineers really decided to make services depend on just one data center. It happened because the dependency was overlooked. Or it happened because the dependency was thought to be a "soft dependency" with graceful degradation in case of unavailability but the graceful degradation path had a bug. Or it happened because the engineers thought it had a dependency on one of multiple data centers, but then the failover process had a bug.
Reminds me of that time when a single data center in Paris for GCP brought down the entire Google Cloud Console albeit briefly. Really the same thing.
Partially true in this case; I can't speak to modern CF (or won't, moreso) but a large amount of internal services were built around SQL db's, and weren't built with any sense of eventual consistency. Usage of read replicas was basically unheard of. Knowing that, and that this was normal, it's a cultural issue rather than an "oops" issue.
Flipping the whole DC data sources is a sign of what I'm describing; FAANG would instead be running services in multiple DC's rather than relying on primary/secondary architecture.
But probably we should. It's an immensely larger coordination problem, but frankly, it's probably the more common failure mode.
> This is a system design that we began implementing four years ago. While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster.
and
> It [PDX-04] is also the default location for services that have not yet been onboarded onto our high availability cluster.
And the product team defining requirements
And IT/governance/architecture teams for not properly cataloging dependencies
And the sales and marketing team not clearly articulating what they're selling (a beta/early access product that's not HA)
Both are non-GA products, and the point is that non-GA are not part of the HA cluster (yet)
While Cloudflare should have been better prepared for this, it seems to be amateur hour in that particular Portland data-center. Other customers (Dreamhost, etc) were impacted too, and I can't imagine they don't also have some very pointed questions.
[1] https://www.dreamhoststatus.com/pages/incident/575f0f6068263... [2] https://www.cloudflarestatus.com/incidents/hm7491k53ppg
"So the root cause for the outage was that they relied on a single data center.". No. Root cause was that data centre operator didn't manage the outage properly and didn't have systems in place in which case they could have avoided it + some systems knowingly and unknowingly had dependencies on the centre that went down because CF did have systems in place to allow that centre to fail.
"Cloudflare has a shit reputation in my eyes, because their terrible captchas". You don't like one product so they have a shit reputation? Enough said.
"but unless you are physically powering off the dc or at least disconnecting the network from the outside world you are not testing a real disaster." If you have ever had to do this, you know that it is never a good feeling. On-paper, yes, you should try your DR but in reality, even if it works, you lose data, you get service blips, you get a tonne of support calls and if it doesn't work, it might not even rollback again. On top of that, it isn't a case of just disconnecting something, most problems are more complicated. System A is available but not system B. Routers get a bad update but are still online, and on top of all of that, you would need some way to know that everything is still working and some problems don't surface for hours or until traffic volume is at a certain level etc. If you trust that a data centre can stay online for long periods of time and that you would then be able to migrate things at a reasonable rate if it doesn't, then you have to trust that to an extend.
All-in-all, CF are not attempting to blame someone, even though a lot is down to Flexential, the last paragraph of the first section says, "To start, this never should have happened...I am sorry and embarrassed for this incident and the pain that it caused our customers and our team."
Well done CF
I mean you're contradicting yourself in the same sentence. Had CloudFlare had such a system in place that would allow that particular center to fail, there would be no outages in the service. The truth is that they didn't account for it , and because they missed it, that center became a single point of failure which is what brought the whole CloudFlare service down. Power outage was just a trigger to discover a weakness in their system design and not a root cause.
Many of these comments sound like they’re coming from some mythical alternate universe where bugs don’t exist and people and orgs have 100% flawless execution every time.
It reminds me a little of someone sitting at a sports bar yelling about a “stupid” play or otherwise criticizing a 0.0001% athlete who is playing at a level they can’t possibly fathom.
Monday Morning quarterbacking.
This seems misleading. Their own status page said the data plane was impacted across many services.
A minor point but this feels like not the most efficient way to manage an emergency. Having some form of staggered shifts or other approach versus just having everyone pile on. If a lot of knowledge resided in specific individuals so they are vital to an effort like this and cannot be substituted then that seems like a risk in it's own.
Does Cloudflare have a plan to move 200+ racks in Oregon if that supplier decides just not to renew that deal whenever it comes up next? Are Cloudflare claiming they were able to build a technical plan which gets their architecture away from this site being a SPOF before the deal is up to renew, or is the CEO making a gamble here again?
Cloudflare have demonstrated their willingness to create reputational issues for suppliers by publicly shaming two of them recently, and here, in only about 2 days from incident. One interpretation of this blog would be Cloudflare are a very unreasonable customer and one who is willing to post incomplete or informal information from their suppliers. Cloudflare also chose to focus the first half of a lengthy postmortem on blaming the supplier and only then on their own culpability for the outage, despite it clearly being a shared responsibility.
One of the diagrams Cloudflare have posted is clearly marked "Proprietary and Confidential". Do Cloudflare have permission to post that? It's not clearly stated that they do. Should other suppliers expect when the sh*t hits the fan that any sensitive information they've shared will be part of a blog?
Most of the "Lessons and Remediation" section is stuff Cloudflare could have worked on at any point in advance of a major incident, and Cloudflare's senior management have quite clearly chosen not to prioritize that work until today, when forced to by this major incident.
When signing large deals, Cloudflare will frequently have to complete 'Supplier Disclosures', and they are also making claims through industry-standard certifications [1] like ISO, SOC and FedRAMP. Most of those will ask questions about the disaster recovery and business continuity plans and Cloudflare will have (repeatedly) attested they are adequate, something that this blog clearly demonstrates was a misrepresentation of their true capabilities.
Will there be an SEC disclosure coming out of this considering it could have material impacts on the business, which is publicly traded? Was there any requirement that the SEC disclosure come first, or be concurrent with a blog?
[1] https://www.cloudflare.com/trust-hub/compliance-resources/
I've run a ClickHouse cluster with hundreds of bare-metal machines distributed across three datacenters in two countries at my previous job. It survived power failures (multiple), a flood (once), and network connectivity issues (regular). This cluster was used for logging and analytics :)
Regardless of any and all corporate spin on the issue, a newbie was dumped into an event at the worst possible time.
I really really really hope he gets a decent bit of counselling to make sure that is fully aware that the issue with the data centre had NOTHING to do with him unplugging the coffee maker to plug in his recharger for his iPhone. Absolutely nothing at all.
I for one am and will always be a cloudflare customer
Yes it is hard and very expensive to do these types of tests. And doing it regularly is even more $$$ and time.
As most customers we seem to be okay with a cheap price hidden behind a facade of "high availability" since I don't really want to pay for true HA. Because if I knew the real cost it would be too expensive.
Are the data centers compensated or anything for this? I'd imagine generator-only might cost more in terms of fuel and wear-and-tear/maintinaince/inspections.
edit:
> DSG allows the local utility to run a data center's generators to help supply additional power to the grid. In exchange, the power company helps maintain the generators and supplies fuel
Interesting.
Ensuring all of their services are fully distributed is now top of mind at CF.
Ultimately, customers win if CF executes.
Having read the post mortem, I do not think it could have been handled any better. I think the decision to extend the outage in order to provide rest was absolutely correct.
I always enjoy reading these reports from Cloudflare as they are the best in the business.
We should absolutely blame them, just as "victims" of ransomware should be blamed. Hardening against system failure is the same process as security hardening.
> Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04
What do you mean you discovered? How could you not know? Surely when you were setting this high availability cluster up years ago and migrating services over, you double checked that all crucial dependencies had also been moved, right? And surely, since you had been "implementing" this for four years now, you've TESTED what would happen if one of the three DCs went completely offline, right???
They discussed this. They had been running tests were they disabled the high availability cluster in any of (and of two of) the three DCs. That test didn't involve disabling the rest of the (non-HA) services from PDX-04 DC (oops).
believing that this doesn't or can't happen to another vendor is being naive.
it has happened to all of them and it'll happen again. can only hope it's super rare.
Centralising the web is such a great move.
The first group of people have been to war. The second have not.
How exactly do you imagine that working while inside a data center operated by a third party?
It's not like they let you stick some solar panels on the roof and run an extension cord to your rack.
It doesn't change anything fundamentally. A complex product is only as good as the weakest link. I have worked with various employers, some world leaders at the time. All of them had seriously weak links.