Today's Outage Post Mortem
blog.cloudflare.com
blog.cloudflare.com
With all of that said, as a Cloudflare customer and also having a call with them tomorrow scheduled already over the WAF stuff, I find it a bit... frustrating that this is occurring now and such a kind of mistake.
Edit: As an aside, I wonder if the Puppet module for Junos will be extended to support route statements. That would make this kind of deployment much easier.
Didn't see this when I posted my comment. So there is a team monitoring not a single person?
That's an assumption on your part. Many negative events happen because of assumptions people make and things they take for granted. I ask a valid question. Do you know for a fact how many people monitor their network or what systems they have in place? Or whether they even simulate events like this to see if other members of the team are even reachable? Have you ever seen the constant testing that goes on at large organizations (such as the military) to make sure battle systems are ready to deploy when necessary?
Note also that cloudflare charges $3000 per month for enterprise plans (in addition to free) and that their own website was taken offline.
Many businesses call roles a "team," even when it's just one person. I believe this is the case here, since the linked CloudFlare blog post states:
> Someone from our operations team is monitoring our network 24/7.
Deleted comment
"Someone from our operations team is monitoring our network 24/7."
"Someone" seems to indicate "1 person". Not "people are monitoring" but "someone". That's it, one person monitors the network? Like the single night guard at the warehouse?
And my use of "not impressive" was in reply to someone who said "impressive" but more importantly thought it was "impressive" that they put up a post mortem within hours. That's nice but it doesn't answer the question that I had.
I stand behind my comment and re ask the question (since the info is ambiguous we have jgramhmc saying "small team who monitor things" and we have the blog post saying "Someone from our operations team is monitoring our network 24/7."
I don't think it's unreasonable (in the interest of transparency) to know exactly the structure and # bodies of who monitors the network at any given time. What is the human point of failure in the system?
I don't depend on cloudflare. But if I was running a mission critical operation and depended on them I might setup a site visit to actually get a feel of what is going on.
As an aside back when the .org registry got started one of the dns servers sat in an open unguarded office under a desk accessible by the cleaning person. I saw it when I did a site visit. And of course if you've been around long enough you know there was a time when the root dns servers sat unguarded in university offices.
It seems that you care more about how many people are actively waiting for something to break vs. how long their response takes. Also, (those people | that person) probably (is | are) the first responder. I really believe that one "first mate" watching the automated ship sail at night ready to triage a technical problem is better than 10 guards who will promptly fall all over themselves when the bits hit the fan.
Paying 5 people to be up all night staring at dashboards isn't generally (required|a good use of people).
Deleted comment
(And I don't agree with that anyway a night guard can call 911 and get the police pretty quick.)
The issue is whether there is a single person monitoring the network or several or even two. And is the coverage different during "working" hours? And what about the skills of the person monitoring at 1am vs. during the day?
Seems that after something happens people wise up to the weak points. Remembering the case of a single air traffic controller in some towers and after something went wrong there was such shock that only one person was on duty with no backup.
Shouldn't that always say - CloudFlare currently runs in 23 data centers worldwide?
Or is that just how one would phrase that if you rent multiple racks or a cage in a datacenter? ...because I've seen that a bunch of times before from just about everyone.
Just curious.
To be clear I share your desire for precise speech, perhaps the could have said 'Cloudflare currently uses 23 data centers worldwide' but given that the person writing this is doing it on a Sunday, probably after a long night and having missed all the things they normally would have been doing on Sunday, I'm willing to cut them some slack.
Don't agree. If you own the data center you have more control over it. We had a case where the UPS systems in a data center had bad batteries and equipment went down because the batteries failed to kick in. Since we don't own the data center we have no realistic way to make inspections and make sure the right thing happens or that the batteries (or the generators) are cycled and maintained. We just have to trust. [1]
Now this may or may not matter with the way they have their redundancy setup. But owning a data center does give you more control over more things.
"not whether they own the buildings where the servers are kept."
Owning the data center and owning the buildings are two different things. Owning the building is owning real estate. Owning the data center is owning the security setup, backup systems etc. Two different things.
[1] So as not to contradict myself with other things I have said I should not say "trust" because you can always put some things in place to verify the right thing is happening (inspections, logs etc.) if you want. But if you own the place it's easier. If I own my home I can decide when to replace the HVAC so it doesn't fail in the middle of the summer. If I rent that's up to the landlord.
Part of being a business is that you have to make tradeoffs in the real world rather than game-theoretical perfect moves. In many cases this means carefully writing contracts because you can't afford the certain expense and distraction of doing it in-house in the hope that the results might be slightly better.
The difference between renting space in a datacenter versus running an entire datacenter is very big, and has ramifications for their uptime, security of their data and disaster recovery. Not sure why they aren't clearer about this.
It appeals to my limited knowledge and non-existant experience that this would be a solution to the prevention of this occurring again in the future?
My 2cents:
Any change should be considered dangerous, and be tested first, time weighted to it's level of change. (An internal policy that could be communicated publicly. One that I apply to all my staff)
It would be also good to have data center clusters (preferably a datacenter clusters are sharded among regions) which would allow these changes to happen as necessary. A random cluster being the "first cluster" with a roll back in place if fails, or proceed through to other clusters progressively until all are live.
The sharding should hopefully alleviate any corner of the world taking any massive hit due to degraded performance.
If they did some colos as vendor J and some colos as vendor C, I think it would be manageable, but I don't really know how much of the cross colo traffic is actually their routers talking to their routers. Homogeneity in networks makes things easier to manage, until a platform fault breaks everything at the same time. In this case, at least it was related to a change they had made and happened quickly, so it was easy to determine the cause; other platform faults may not be as easy to determine, but if only your vendor X colos fell over, at least you'd have your vendor C colos up and something to look for.
Not sure what your timeline shows but the "cloudflare is down" post hit the #1 spot on the front page just a few minutes after they went down. About 40 minutes after that, the services started to come back online for me.
That's a significant outage. That's not reflecting on the job they did bringing things back online but your statement made it seem like a minor outage.
Deleted comment
However you are most likely right, CloudFlare should have test ed it before rolling out.
See also R.Reagan: "trust, but verify"
One thing that occurred to me though is that performing a hard reboot of the routers required calling people to physically access the devices and took some time to perform (as you would expect). Although I wouldn't expect it to be needed very often, I'm sort of surprised CloudFlare doesn't have out-of-band remote power cycle capabilities.
There may be some factor I'm not considering that would make that an unattractive option, but it does seem like it could cut down an already quick response time even further for any similar events in the future.
I've seen smaller routers, CSU/DSU, etc. type devices in branch offices on cyclers, though.
I think it's mostly that the routers usually have both good OOB management and good watchdog (reboot on freeze) behavior, and that the PSUs in the bigger routers tend to exceed the per-port power limits of most of the external power cyclers.
It may be a good idea, though.
I'm looking at the Verizon Private-IP thing (an outsourced private network over Verizon's cell infrastructure) for OOB management of lots of CPE; the cost per device per month is low, and then you pay for bandwidth across all of them. Makes initial provisioning easier, plus ongoing monitoring/maintenance.
e.g., from simple things like hard drives in a raid from different vendors, to n-version programming in safety critical systems (like airplanes).
Same with Juniper. (there aren't really other viable options besides those two)
You could build the same site fully independently with all-Cisco on one, and all Juniper on another, and potentially get some better isolation from vendor faults, but at very high expense.
You end up with much worse reliability if you have a mixed Cisco/Juniper network without a lot of additional isolation otherwise.
Total misconception. BGP, OSPF, ISIS, LISP, etc. are all non proprietary. Sure, the root cause of this particular problem is that CF is using something specific to Juniper, however router interoperability is not predicated on components like that. This example was a tool CF operationalized, and likely had little to do with their routing with the exception of it being a metric they may have influenced routes with.
People who have all Cisco or all Juniper shops namely do it from a cost perspective. Sure, there are some reasons outside of that but it's likely the big driver. The more you buy, the more you save. And the network sales realm is royally messed up to begin with. I've seen Juniper give 90% discounts on hardware just to break into a Cisco shop. But, the reality of the situation is that all of this gear is marked up well into the thousands of percent. So if you're not getting, minimally 50% then your probably not doing yourself due diligence.
There are some routing protocols which interoperate (which is how different sites on the Internet can talk to each other), but most of the protocols used for HA or management of a given set of routers, or, more importantly, most tested/debugged implementations of HA and device management, are Cisco or Juniper specific.
No big deal announcing routes to your upstream if you use Juniper and they use Cisco. Big deal if you have Cisco+Juniper and want to do HSRP (Cisco-only).
I've been in network engineering for 12+ years and I fundamentally disagree with a lot of what is said about "networking" and interop by many programmer-types (not casting here, but) on HN. Yes, yes, you may understand system DevOps to a point, however I'm not sure you've spent a significant amount of time studying Dijkstra's algorithm or truly have an idea of how to deploy a global IPv6 overlay. I'm also not trying to be snide here but I feel that, often times, many things that come up on HN are just fundamentally designed wrong from PHY all the way up until the devs get a hold of the rest. I've been in a very successful startup (think one of the top online backup services) wherein their network was run on commodity junk hardware. They were asking me how I'd troubleshoot this, that and the other thing - obviously with no debug (this guy said that with a grin). First and foremost, you designed it wrong - I can show you inefficiency in about 10 minutes of performance engineering that I would have designed around without thinking about those things. So, yes, I can waste time tracking down a bad NIC on your network, but if you feel that you've earned geek cred because you fired up Wireshark and parsed through a few simplistic ARP tables - you haven't impressed anyone but yourself. That's when I realized I was working with professional developers, and not network architects.
Your simplistic view of FHRP is trivial at best. Maybe if you were talking about how you'd design fault tolerance into a virtual link, say an LSP, with something like BFD in your design I'd be more impressed than conversations about proprietary redundancy protocols of which most network engineers won't touch for a variety of other reasons than the big "C".
</endrant>
Similarly very few developers have to solve open CS problems in writing a CRUD application (or I guess more comparably to ops, come up with a novel implementation of a complex algorithm).
This is progress, though.
"<redacted> takes your security very seriously." - right. That's a statement, not information regarding the thought or implementation. There's not even a mention of technology. <sigh>
I've been hearing this for a decade. It's still not true. I'm not sure why, either.
I agree the right choice today is almost certainly a C or J router and probably C or A switches, but e.g. hardware load balancers like F5 seem to be losing out to software in most deployments (increasingly).
I built a decent sized network with Zebra 15y ago, which was pretty obviously the wrong tech, but interesting.
Hello, VRRP!
There are open standards for pretty much every Cisco proprietary protocol. Even EIGRP is now available as an informational RFC.
most service provider networks manage via the CLI (generally scripted, for better or worse) and occasionally a vendor-specific API.
while I will agree juniper and cisco are generally the best choice for core/edge routers, there are other 'viable' options depending on your requirements. if you need in excess of 2500 BGP sessions on a single chassis, there aren't many viable options besides a Cisco 7600.
I certainly do not mean to be rude, but I feel you're attempting to speak from experience you don't fully have (yet, hopefully!)
You seem to be confusing hardware with software. Juniper's gross margin was 64.25% for the quarter ending Dec. 31, 2012, and in that ballpark for previous quarters back to inception.
Back around 2000 this was a big deal. Cisco slacked on gigabit routers, and Juniper didn't have a comprehensive product portfolio, so while SP networks could be all J (but maybe with some switches from Extreme, etc,), enterprise networks were a lot more likely to have juniper and Cisco mixed, if they needed juniper performance in the core. Juniper ended up broadening their portfolio and Cisco improved their high performance offerings a few years later.
Here's an example of a large provider with parallel infrastructure each powered by a different hardware provider(Brocade/Extreme). One failed and one kept working. I seem to recall a more detailed RFO but my Google-fu has failed me this morning.
Running an anti-DDoS/CDN service which handles traffic like Cloudflare does would be vastly more difficult.
It's certainly possible to do, but I think the given ~reasonable engineering resources, the net reliability of a heterogenous J/C version of CloudFlare would be less, and performance worse, than what they have now.
Switch fabric is a lot closer to the "run different models of hard drives" (although, you don't do that WITHIN a RAID group either -- you do it on separate RAIDs and possibly separate chassis), than routing infrastructure (which is like running a 777 with 1 GE engine and 1 RR engine. At best, you can turn it back into a 747 and run 2 GE engines and 2 RR engines.)
Unfortunately, amortized over the lifetime of a computer system, risk is not reduced in this manner. There is no hedging of vendor vs vendor in a technical portfolio; what happens instead, for any tech of significance, is the internal development of an abstract control plane that can communicate with both, and that control plane is then the single source of defects. In the meantime your engineers have to become world-class experts on two platforms rather than one. In practice the divided loyalties will turn one world-class engineer into two half-assed ones.
Domino-effect failures, or global misconfiguration failures like those experienced by Cloudflare are edge cases in my experience and not something you should optimise for. When they happen they tend to be catastrophic, but worse is the insidious decline in quality caused by carrying too much technical debt.
Cloudflare's scenario is not comparable to the installation of a RAID set. They are more comparable to a developer of RAID controllers. The experience curve for such is very, very long.
Not saying they couldn't have done other things to make this situation less catastrophic, but diversity of core technology portfolio isn't a winning ticket.
Hope lessons are learnt and your next generation is less prone to these attacks
Wonder if Cloudflare need to do some tests along the lines of:
A) List up all the types of rules we usually use to mitigate these situations. B) Run those rules on a test router with wildly unusual input values, as was the case in this situation. C) Send test traffic using that wildly unexpected input to see what happens.
Basically a bit of manual fuzzing
Time-consuming and maybe not worthwhile, but it could save against another full system death.
1. That video of the BGP routes disappearing is awesome, and
2. A 40 minute outage sounds bad, but consider the following timeline (based on the writeup):
> T+0: route change made, propagates
> T+10: Response team online, attempting local fixes
> T+30: Routers across 23 data centres in 14 countries hard reset and networks coming back up.
I know very little of networking, but this seems to be a recurring pattern that aggravates many major outages. What surprises me is that this so often seems to be a scenario not accounted for.
I bet that most of the time the domino effect happens to internet services in general it's with nodes that are accepting most requests. They allow themselves to be overloaded. An active HTTP session uses orders of magnitude more resources than simply denying the initial packet and forgetting about it forever.
>". I would be willing to bet that outside of heavy-DDoS conditions that even a tiny fraction of Cloudflare's network could handle the incoming tcp connections and deny all of them." depends on the attack.
>"You can send a tiny error page. You can let X% of requests get through and be fully served." Not usually that easy.
I call BS on saying it's not easy to limit the number of served connections and RST the rest. Isn't this something every web server can do by itself it's so easy?
There will always be potential SPOF.
The core problem with CloudFlare is that they seem to have a highly-centralized take on what is normally a massively-decentralized solution-space, with large numbers of value-adds they encourage customers to use without making it clear that they treat in a haphazard manner, doing very little testing before deploying pushing-the-envelope features while simultaneously having very little in-house debugging expertise to handle serious issues.
(As a concrete example of that last complaint, Cydia was crippled for an entire day due to ModMyi turning on CloudFlare's "preloader" transformation, which apparently caused many WebKit-based browsers--including both MobileSafari and Cydia--to entirely lock up; CloudFlare seemed to go the entire day without noticing, which I continue to be utterly shocked by, and it was only after I told them how to fix it that they were able to acknowledge the issue.)
http://www.saurik.com/id/14 <- When "Dumb Pipes" Get Too Smart, an extensive analysis of this bug
http://blog.cloudflare.com/cloudflare-fastest-free-dns-among...
This should raise a red flag, as it must be impossible. Ethernet NICs would just bail out on packets longer than what you've set the MTU to, and ethernet frames would just come from the next hop in most cases. And IP packets have a max length field of 16 bit.
> An optional feature of IPv6, the jumbo payload option, allows the exchange of packets with payloads of up to one byte less than 4 GiB
As a consequence, fragmented ipv6 packets are error for use in DoS attacks. This is not a "weird" occurrence, but rather an expected one, and since end points are not required to accept such huge packets, I am surprised Cloud flare want already doing all it could to advertise to upstream sources that IPv6 fragments longer than a much smaller than 90K should be dropped, at least if rooted to their DNS. I am also surprised that when their software came up with that kind of a response without first validating that it wouldn't cause the exact memory problem it did. Rules on v6 fragmented packets that can't match on a single fragment are inherently dangerous. It is only reasonable to have safe guards already in place for them.
I am also not sure this is really a bug in Juniper software. I imagine the memory problem only shows up with high traffic and in the midst of a DoS attack. That is kind of a given when you put a rule like that in that kind of a situation.
they didn't. and they paid the price. good on 'em for the quick and honest post-mortem. regardless, it was a dumb move.
I don't see how that would solve anything here.
Downside: priced accordingly.
In this case: Development. Versioned change. Test or staging environment. Tests pass. Production.
Yes, I fully agree that for things like software and standard network maintenance the above is good. But as someone else mentioned in this thread. DDoSes that require quick resolution put you between a rock and a hard place in terms of doing things "right"
However, at least how I read it, your comment was about testing rules in consistently in dev -> staging -> prod when you create one which I think is not viable in this situation since you are on a very tight deadline with immediate impact on your customers.
It's probably more reasonable to split your network into a few more independent sections and never do updates which affect everything, but unless you're building the space shuttle (and can accept vastly higher costs and lower performance), it's probably better to pick one hardware platform, at least now.
For standalone units that don't interact, n-version redundancy is good.
But if they have to interact and somebody has to troubleshoot, n-version redundancy is a nightmare.
The only reason it got fixed this quickly is because they had an intimate knowledge of a single vendor's products. That would be much more difficult with multiple vendors.
Ouch if a single host activity took down ~750k websites - whether deliberate and direct or not.
"Oh noes! Something went wrong."
They have my sympathy: So, they typed in a 'rule'. At one time I was working in 'artificial intelligence' (AI), actually 'expert systems', based on using 'rules' to implement real time management of server farms and networks. Of course, in that work, goals included 'lights out data centers', that is, don't need people walking around doing manual work but not the case of 'lights out' as in the CloudFlare outage, and very high reliability.
Looking into reliability, that is, putting into a few, broad categories the causes of outages, a category causing a large fraction of the outages was humans doing system management, or as in the words of the HAL 9000, "human error". Yup.
And the whole thing went down? Yup: One example we worked with was system management of a 'cluster'. Well, one of the computers in the cluster "went a little funny, a little funny in the head" and was throwing all its incoming work into its 'bit bucket'. So, the CPU busy metric on that computer was not very high, and the load leveling started sending nearly all the work to that one computer and, thus, into its bit bucket and, thus, effectively killed the work of the whole cluster.
As one response I decided that real time monitoring of a cluster, or any system that is supposed to be 'well balanced' via some version of 'load leveling', should include looking for 'out of balance' situations.
So, let's see: Such monitoring can have false positives (false alarms) and false negetives (missed detections). So, such monitoring is necessarily essentially a case of some statistical hypothesis testing, typically with the 'null hypothesis' that the system is healthy, applied continually in near real-time. So, for monitoring 'balancing', we will likely have to work with multi-dimensional data. Next, our chances of knowing the probability distribution of that data, even in the case of a healthy system, is from slim down to none. So we need a statistical hypothesis test that is both multi-dimensional and distribution-free.
So, CloudFlare's problems are not really new!
I went ahead and did some work, math, prototype software, etc. and maybe someday it will be useful, but it wouldn't have helped CloudFlare here if only because they needed no help noticing that all their systems around the world were crashing.
In our work on AI, at times we visited some high end sites, and in some cases we found some extreme, high up off the tops of the charts, concern and discipline for who, what, or why any humans could take any system management actions. E.g., they had learned the lesson that can't let someone just type in a new rule in a production system. Why? Because it was explained that one outage in a year, and the CIO would lose his bonus. Two outages and he would lose his job. Net, we're talking very high concern. No doubt CloudFlare will install lots of discipline around humans taking system management actions on their production systems.
Net, I can't blame CloudFlare. If my business gets big enough to need their services, they will be high on the list of companies I will call first!
Except that it wasn't human error, at least not in the sense that the decision to enter the rule, or the rule itself was in error. The human error was with the bug in the Juniper firmware that caused this rule to crash this router, and arguably with the CloudFlare process that allows rules to be propagated to all routers concurrently, rather than segmenting the network and testing for success before further deployment.
Had any of these steps not gone wrong there likely would not have been an outage. It was a combination of failures that caused it.
I thought that the rule was in error: I couldn't read the rule clearly on the screen, but it seemed, or I guessed, that the problem with the rule was that the "humans" omitted the decimal points and, thus, asked for blocking packets with lengths 1000 larger than intended. The Juniper software got sick, i.e., allocated too much memory, only because it was trying to swallow working with such absurdly large packet sizes. But, then, I couldn't clearly read the screen capture with the rule.
Regardless, the Juniper should either have rejected the rule, or accommodated it.
I would imagine that while we have seen a public reason for outage statement that there is a lot more work going on inside CloudFlare as far as the post-mortem is concerned. There are a lot of angles to cover here to really understand how the _system_ failed.