Update about the October 4th outage
engineering.fb.com
engineering.fb.com
It looks like Zuckerberg doesn't have a personal Twitter though, nor does Jack Dorsey have a public Facebook page (or they're set not to show up in search).
He does: https://twitter.com/finkd
It was amazing. I’m sad he remove it.
hint: "some people".
I hope we get something more substantial and informative than that over the next couple of weeks, but it doesn't seem (at least from my searching) that Facebook is in the business of publicly posting in-depth post-mortems for their outages, which I personally find unfortunate.
vs
How tech savvy are people that Facebook profits from?
Gotta target your audience, every communication is PR...
I'm sure the world would quickly adapt by re-adopting these things called "websites" and "email" but in the meantime, it's highly self-centered to think this "didn't matter".
This is Hacker News, so the distinction between network performance, server performance and application performance should matter.
"The Internet" did not slow down. "The Internet" infact probably had more available capacity as a result of Facebook's outage, as all those bits of outrage and cats ceased to be transferred for the duration.
Some applications may have seen performance hits, as a result of poorly thought out dependencies on an external service without graceful failure.
Some applications may have seen increased load and suffered due to server resourcing constraints, caused by applications like the above failing to fail gracefully, and instead polling more aggressively.
> They don't have a right to privacy here, and we are all owed an explanation
Morally / ethically, you're right. The fact that Facebook exists in it's current form tells me that morals and ethics aren't particularly important to the real world.
https://www.theverge.com/2021/10/4/22709123/facebook-outage-...
I had Cloudflare's woes in mind when I wrote that.
This is referring explicitly to users of 1.1.1.1, which is likely not the same infrastructure as domains hosted by cloudflare dns.
[0] https://blog.cloudflare.com/october-2021-facebook-outage/
Why? Why couldn't you just post that the RCA is still ongoing and that proper updates will follow? Otherwise all you get is meaningless fluff.
This is not meaningless fluff. It may not provide info to technical persons, but valid info to other persons, as others noted (was hacked: yes/no).
Getting down to root cause takes time. It is usually multiple-things-at-once that caused X to happen. And then they must also make a decision on how to prevent X to happen again. All that must be written into RCA. It takes days not hours.
An example from my life: Service has intermittent disruptions. Antivirus activity correlated 100% with disruptions. Upon further investigation turns out that AV was just doing its job when there was less load. (And before anyone points out why on earth there is AV on such service, well, because it deals with user uploaded files)
So what should I have called out - AV is the guilty one? And then say: oh, no, false info.
Its much better for some random committee in Congress to debate antitrust forever, instead of bigger committees and agencies debating national security threats
It's not a full root cause analysis, to be sure, and leaves many open questions, but I definitely wouldn't describe it as painfully vague.
DNS is related to BGP only that without the right BGP routes in the routers, no packets can get to the facebook networks and thus the facbook DNS servers.
That their DNS servers were taken out was a side affect of the root issue - they withdrew all the routes to their networks from the rest of the Internet.
Not picking on you - but there has been a lot of confusion around DNS that is mostly a red herring and people should just drop it from the conversation. Everything on facebook networks disappeared, not just DNS. The main issue is they effectively took a pair of scissors to every one of their internet connections - d'oh!
https://mobile.twitter.com/mikeisaac/status/1445196576956162...
DR downtime was about an hour, but the bank fired him anyway.
Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications.
Sidenote: I asked the bank infrastructure team why the DR site was in the same earthquake zone, and they thought I was crazy. They said if there's an earthquake we'll have bigger problems to deal with.
Prostrate, he came before the COO expecting to be canned with much malice. The COO just asked if he learned his lesson and said all is forgiven.
The US, not even once.
The guy should have had "reload in 10", an outage window and config review. There must be more to this story than it being a firable offence for causing a P2 outage for an hour.
> Sidenote: I asked the bank infrastructure team why the DR site was in the same earthquake zone, and they thought I was crazy. They said if there's an earthquake we'll have bigger problems to deal with.
I bring this sort of thing up all the time in disaster planning. There are scales of disaster so big that business continuity is simply not going to be a priority.
Unless there was some kind of nefarious intent, it's very unlikely anyone will be 'punished'. The likely ramifications will be around changes to processes, tests, automations, and fallbacks to 1) prevent the root sequence of events from happening again and 2) make it easier to recover from similar classes of problems in the future.
Organizational failures require organizational solutions. That seems pretty obvious.
In many ways we're wired to do it and it FEELS GOOD to do it. Even the industries that are championed for focusing on organizational process over human blame (ex: airlines) are often lulled into initially falling back on the emotional knee-jerk of "pilot error" (see: the early days of the 737 max debacle).
Companies have to be very intentional, usually top-down, about focusing on the context that allowed humans to fail instead of the human themself. That's often easier said than done.
But yeah, likely not.
I wasn't aware that Stanley Kubrick was now in NetOps. /s
Whoops! Never attribute to malice that which can more easily be explained by stupidity and all that.
Will a real postmortem follow? Or is this the best we are gonna get?
It’s not really surprising to me that Facebook is writing comms that most users will understand right now, rather than publishing detailed post-mortems straight away. You have to speak the same language as your users initially in these comms.
Although I wouldn’t be surprised if we see a post-mortem in the days ahead, but Facebook probably will want to say why it happened (not just what happened, but why did it not get detected during testing, was the configuration change correct but there is an underlying bug on the routers etc) and what new mitigation’s will be put in place to stop it happening again, and these might not be known yet.
Facebook sells ad space, retail. The impact on their customers of the outage is ‘sorry, you couldn’t buy ads for a few hours.’
Demanding a public RCA for this is like demanding an RCA from Costco because they’re out of stock of tinned beans.
https://engineering.fb.com/2021/08/09/connectivity/backbone-...
The badge system should be local to the building. There are few actual reasons (sure, besides "efficiency") of why badge control should be centralized. Even less reasons for it to be a subdomain of fb. Another option would be to keep the system but make it failsafe (but it seems the newer generation doesn't know what that means). If the network goes down keep it at the last config. Badge validation should be offline first and added/removed ones should be broadcast periodically.
This is the same issue with smartlocks times the number of employees. Do you really want to add another point of failure between yourself and your home?
Akso, it's likely not on an fb subdomain, but something like office.security.fb-infra.com (example). It just happens to be that fb-infra.com is using the Facebook DNS server.
You might need a break-glass account/badge somewhere. Sure, the angle-grinder works, but probably cost you 2h maybe?
> it's likely not on an fb subdomain, but something like office.security.fb-infra.com
Thanks, yeah, makes sense
It's just more expensive and another thing to maintain, and still doesn't account for _all_ failure modes (what if you sync really frequently and a bad change was made deleting all accounts?)
Still, a bit unexpected behaviour though.
Also IIRC an employee was posting on Reddit saying the incident started shortly after a network update was posted this morning.
When you know more about the tech and systems involved a mistake seems infinitely more likely than sabotage.
Seems like the new system having an unanticipated flaw is a far more likely scenario than a malicious actor.
More boring - but usually the boring stuff is the far more reasonable.
Some of them seem to actually say it's likely, not just possible.
- Facebook was hacked
- They did it on purpose to bury the whistleblower story
- No one could access Facebook offices
- They had to cut open servers with angle grinders
- Disgruntled employees changed DNS records
- Lots of made up numbers for how much money Facebook/the rest of the economy was losing (or gaining)
They probably rushed out this blog post just to dispel some of these rumors.
And they probably "fixed it" by putting someone near the door to let people in.
My curiously downvoted point was, I’m surprised an internet problem would stop an high security access control system from functioning - they’re supposed to be designed to cope with that and continue to work autonomously in emergencies.
There was an Indian opposition Member of Parliament blaming the current government that it blocked FB due to some protests being held in the Capital.
> Correction: Oct. 4, 2021. An earlier version of this article misstated a Facebook team’s means of getting access to server computers at a data center in Santa Clara, Calif. The team did not have to cut through a cage using an industrial angle grinder.
Another bit of misreporting (e.g. The Verge) is that this was "one of its main US data centers in California" when there's no such thing. The seriously big heaps of hardware are elsewhere and they're not shared, so there are no cages to force entry into. I've been in one; no cages in sight. I know which facility they're talking about, and its only distinguishing characteristic is that it's close to where relevant people live.
I want to believe.
Sure, my friends and I wondered if it was a malicious insider. It takes surprisingly few people in an organization to cause chaos.
Knowing IoT, it isn't unbelievable that badge readers could be offline.
Knowing division of duties, it isn't hard to believe that the network engineers, domain admins and datacenter ops people may have hustled to a DC to get things back online.
Never did I see a large number of people take anything as fact that didn't seem to be substantiated.
It amazing how wrong people can be and how confident they are about being right.
Even sometimes fighting _me_ about things _I_ designed and built.
It’s quite sobering; taught me not to believe all the speculation I read.
No, that just happens during uptime.
This could be anything, potentially.
I'm not very knowledgeable in computer networking, but this could be as trivial as an incorrect update to a DNS record, right?
I had question: This is what we can only perceive through internet/routing table entries right?
Internal to FB, we don't know what had caused issues that led to the BGP UPDATE.
That's kind of what has been confusing me - there's a lot of speculation around FB's data center design and what actually happened, but we actually don't know for sure until they post an RCA - please correct me if I'm wrong here.
It strikes me it's like DNS when you get a SERVFAIL, why not try the prior IP address. The similarity in the design here suggests there may be common reasoning??
When the announcement is revoked, you fall back to a less specific prefix if present, or your default route.
If you've got a full BGP table, then you tend not to have a useful default route (you should have specific routes for everything) and it might be useful to fallback to the last known value. But many participants have an intentional default route and then get announcements for special traffic --- dropping the announcement would mean to send it on the default route instead. It's hard to know what the right thing to do is, so better to do what you were told by the authority.
DNS is a bit different, but again, the authority told you to use some data and how long to keep it (ttl), if they're not there to tell you a value again later, what else can you do but report an error? Some DNS servers have configurable behavior to continue using old data while fetching new data or when new data is unavailable.
But the expectation is if you can't keep your BGP up and your DNS up, your server probably isn't up either. Note that in this case, bypassing DNS and going to the FB Edge PoPs that were still network available (because of different BGP announcements, that weren't withdrawn) resulted in errors, because they weren't able to connect to the upstream data centers. (Or so it seems)
It happened to also kill the announcements for anycast DNS.
I can completely picture a world in which many people bought some ads yesterday morning (say, to promote an event that occured yesterday evening), the ads were never displayed to anyone, and FB will keep the money, thank you.
By having a large centralized and monolithic system, aren't they guaranteeing that mistakes cause huge splash damage and don't separate concerns?
Of course, we can't really know for sure until/if they release exactly what caused this issue.
The anti-anti-pattern would be to have country specific AS's, kind of like a franchised restaurant pattern, where each country's version of fb owns their own AS that loosely forms back into the facebook mothership org, but from infrastructure point of view points to their own set of AS netspace.
I recall major ISP's screwing up their routing tables in the past but never globally on this level.
https://www.bleepingcomputer.com/news/technology/ibm-cloud-g... (this one isn't clear, maybe BGP hijacking, and if so, not sure who the responsible party was)
https://www.catchpoint.com/blog/vodafone-idea-bgp-leak (not sure how major this one was)
You can practically search ISP bgp outage and get news about the last couple times they screwed up BGP and caused a big problem. Or service BGP and get a 50/50 chance of the service screwing up BGP or an ISP/country hijacking their routes and causing a big problem.
BGP is one of the best ways to break things at scale.
For all it’s flaws, BGP is the piece of the Internet that truly makes it decentralized. Without it, there would be a centralized routing table of some sort.
You are effectively asking is why is there a single routing table for the internet.
To put in simple terms having a single routing table is what it makes it the internet we can share, otherwise it would just be a bunch of independent networks.
It certainly does not. If I peer with you, neither of us (generally) announce that route to our other peers, but often announce to our customers. There are many routes that are not visible to everyone, and there is no single routing table for the internet. Each BGP speaker ends up with their own routing table, although there are a lot of similarities.
BGP in this context implicitly meant external BGP . Yes, no single router necessarily sees all the routes, but all routers combined generally see the internet as a single network of networks was my point.
Convergence in this context is how most ASN will resolve on where/how to route a specific ASN traffic.
It is hard to peer with someone and not trust the routing table they publish, that is why Pakistan could by mistake block YouTube for everyone few years back.
In this case if you peered with Facebook, and they published incorrect routing for their AS you would accept it.
This doesn't mean FB couldn't have used multiple ASNs did some rolling updates etc, however without knowing what exactly fb screwed up for five hours it is hard to say what they could done differently.
No, they're asking why a single set of routers is in charge of announcing BGP routes for all of facebook. If you have multiple ASes, with independent configuration sources and independent routers broadcasting them, it's a lot harder to break everything at once.
However if even one set of routers were misconfigured and it was announcing incorrect routes for all their ASes as result of the issue then their peers will not typically drop that set alone automatically.
BGP doesn't have a Paxos / Raft style smart consensus algorithms, it runs on trust. Either their peers had to trust what FB published or they won't be peering with FB ASes in the first place.
That's what I was meant when I said it comes down to one network web of trust, and if there is a breakdown of that redundancy cannot typically help
The config database would reject routes for the wrong ASes. The router would reject it. The peers would be told "add filters so you only accept these ASes from these routers".
Maybe they have all that and it somehow broke anyway? But what it looks like from the outside is that all the ASes are controlled by the same system.
> their peers will not typically drop that set alone automatically.
I'm not sure what this sentence means.
> router would reject it.
Any of these hardware could have bugs, if one of them announces wrongly it will be propagated wrongly by all other ASes peering with them and to the next level so on and on. That is the point, at this level it is possible to fuck up globally.
> peers would be told "add filters so you only accept these ASes from these routers".
There are hundreds of ISPs , all of them peer cannot directly with each other. Routes are propagated downstream and upstream it is a web of networks running on trust.
While filtering is built into most implementations ( sadly not the protocol itself), practically ISPs have no easy way to determine which AS can actually originate which other AS'es traffic, so they don't actually implement a lot of filtering. Remember traffic can have more than 2 hops. Effectively that means you would be routing traffic for AS 3/4 hops away. Neither you nor your peer would know anything about it or whether you can trust it etc.
Even if some ISPs do drop/block the announcements, unless every single AS also implements the block there won't be an impact. Traffic would route through ASes which don't have filtering and announce the routes incorrectly . For example say AT&T blocks an incorrectly announced FB route, but British Telecom does not, BGP is designed to assume that FB has lost peering with AT&T and route all traffic for FB via British telecom.
If filtering was robustly possible we wouldn't have periodic BGP hijacking incidents as we do whether accidental or maliciously. The famous Pakistan Telecom Youtube hijacking [2] or as recently as April-2021 [3] or incidents over the last few years usually authoritative governments (such as China/Russia etc) but with impact well beyond their networks.
[1] http://www.bgpexpert.com/article.php?article=145
[2] https://www.ripe.net/publications/news/industry-developments...
[3] https://blog.apnic.net/2021/04/26/a-major-bgp-route-leak-by-...
This is partially due to design limitations of BGP, and partially due to it being nearly impossible to eliminate all sources of large scale failures in any highly complex system, and increasing the uptime of a system that already has a few nines costs an additional order of magnitude for each new nine. At some point you set your risk tolerance and have catastrophic failures now and then.
Although it only covers their API and business apps, not the site itself.
It looks like the status page is hosted on CloudFront though, so it got part of the way. (Of course, the other question is if it was updatable / updated during the outage)
Is there any way to keep DNS up in case BGP goes down for any reason? Like a fallback nameserver hosted elsewhere/not affected by Facebook's ASs?
Is it technically impossible or did Facebook just assume something like yesterday would never happen and kept things simple instead of complicating things?
At least one of those networks they accidentally removed also happened to contain the DNS servers; DNS being unavailable was a symptom - but not part of the root problem. Any focus on DNS at this point is a red herring.
Think of routes as street directions - they tell routers where to ship packets. If you erase all your addresses and directions to them from the outside world at at large, then there literally is no way for network packets to get from the global Internet to Facebooks networks (where I imagine the DNS servers were up and probably twiddling their thumbs wondering where everyone went).
An easier way to think of it - they essentially took a pair of scissors and cut the cable connections to the Internet - which is why it was so catastrophic.
They only way to mitigate that is to have an identical infrastructure managed by different tooling so a bad configuration setting from one environment wouldn't pollute the second in the same way. Not exactly an easy thing to do and might cause more other problems than it's worth. And you would have to do that for all services, not just DNS. Let's say Facebook used Cloudflare for their DNS. Great - DNS can resolve your request for fb.com to the IP address of the facebook datacenter - there still is no path for your packets to get to that facebook datacenter because they accidentally purged the routes to their networks.
It's easier to just not cut your connection to the Internet :) I'm sure there are all kinds of internal discussions picking this incident apart and formulating ways to either prevent it, or more realistically - have improved procedures to speed recovery when it inevitably happens again. BGP is not known for its inherent robustness or security. But since it's at the core of the Internet, any changes to it would have to be done on a massive internet-wide scale in perfect unison or the "cure" would be a lot worse than the current problems with it.
Murphy was indeed an optimist! (search "Murphy's Law" for those unfamiliar with the idiom)
If it's a FB managed server, run on someone else's network, you still have a lot of the FB software risk (FB's software stack and development mantra make it easy to push changes, some of which break everything, including the ability to push further changes); even if not FB, there's a similar risk.
If it's not a FB managed server, like a 3rd party DNS provider, it's difficult to get that synchronized considering all the fun geographic loadbalancing FB is doing at the DNS level. That's generally hard once you start doing this; and it's why you don't see many dual-provider DNS setups.
Really, the status page should be not on a core domain, so that the DNS can just be external.
FB DNS breaking yesterday almost doesn't matter in the scheme of things, because the BGP breakage broke everything anyway. Would it have been a bit nicer to get http error messages instead of DNS not found messages, sure; but mostly nothing was working anyway.
Which was and is a lame assumption. Stuff happens. SMTP wouldn't even be phased by this; it would just pick up where it left off.
I've seen far too many applications fail in bizzare ways because people make unrealistic assumptions like "X will ALWAYS be there". Sure it's highly unlikely, but when you have multiple things making the same dumb assumptions, on the inevitable day when multiple things that need X and X is suddenly no longer there then you start to get cascading effects of Y that relied on something that relied on X not being there when it is assumed that it would always be there so now Y fails, and then something dependent in the same way on Y unexpectedly fails and so on.
One should never assume that anything will "always" be available. That's an incredibly unrealistic assumption; and the more interconnected things become, the chances of these really nasty dependency chains/cascade failures skyrocket - leading to far worse outages and longer recovery times.
NO CARRIER
I found a paper that describes the process in detail. See page 10-11:
https://web.archive.org/web/20211005034928/https://research....
Phase Specification
P1 Small number of RSWs in a random DC
P2 Small number of RSWs (> P1) in another random DC
P3 Small fraction of switches in all tiers in DC serving web traffic
P4 10% of switches across DCs (to account for site differences)
P5 20% of switches across DCs
P6 Global push to all switches
We classify upgrades in two classes: disruptive and non-disruptive, depending on if the upgrade affects existing forwarding state on the switch. Most upgrades in the data center are non-disruptive (performance optimizations, integration with other systems, etc.). To minimize routing instabilities during non-disruptive upgrades, we use BGP graceful restart (GR) [8]. When a switch is being upgraded, GR ensures that its peers do not delete existing routes for a period of time during which the switch’s BGP agent/config is upgraded. The switch then comes up, re-establishes the sessions with its peers and re-advertises routes. Since the upgrade is non-disruptive, the peers’ forwarding state are unchanged.
Without GR, the peers would think the switch is down, and withdraw routes through that switch, only to re-advertise them when the switch comes back up after the upgrade. Disruptive upgrades (e.g., changes in policy affecting existing switch forwarding state) would trigger new advertisements/withdrawals to switches, and BGP re-convergence would occur subsequently. During this period, production traffic could be dropped or take longer paths causing increased latencies. Thus, if the binary or configuration change is disruptive, we drain (§3) and upgrade the device without impacting production traffic. Draining a device entails moving production traffic away from the device and reducing effective capacity in the network. Thus, we pool disruptive changes and upgrade the drained device at once instead of draining the device for each individual upgrade. Push Phases. Our push plan comprises six phases P1-P6 performed sequentially to apply the upgrades to agent/config in production gradually.
We describe the specification of the 6 phases in Table 4. In each phase, the push engine randomly selects a certain number of switches based on the phase’s specification. After selection, the push engine upgrades these switches and restarts BGP on these switches. Our 6 push phases are to progressively increase scope of deployment with the last phase being the global push to all switches. P1-P5 can be construed as extensive testing phases: P1 and P2 modify a small number of rack switches to start the push. P3 is our first major deployment phase to all tiers in the topology.
We choose a single data center which serves web traffic because our web applications have provisions such as load balancing to mitigate failures. Thus, failures in P3 have less impact to our services. To assess if our upgrade is safe in more diverse settings, P4 and P5 upgrade a significant fraction of our switches across different data center regions which serve different kinds of traffic workloads. Even if catastrophic outages occur during P4 or P5, we would still be able to achieve high performance connectivity due to the in-built redundancy in the network topology and our backup path policies—switches running the stable BGP agent/config would re-converge quickly to reduce impact of the outage. Finally, in P6, we upgrade the rest of the switches in all data centers.
Figure 7 shows the timeline of push releases over a 12 month period. We achieved 9 successful pushes of our BGP agent to production. On average, each push takes 2-3 weeks
Hey what is our internal BGP called again? AS32934?
"Yeah"
"OOK."
We YOLO'd our BGP experiment to prod. It failed.
https://web.archive.org/web/20210626191032/https://engineeri...
Networks have grown so large and complex that the only reasonable way of managing them is through SDN, and a small mistake in configuration might results in a cascading effect on the whole infrastructure.
That's also true for the entire (western) internet. We've ended up with a centralized market where a few key players, e.g. cloud providers/CDNs/DNS (Amazon/Google/Microsoft/Akamai/Fastly/Cloudflare) can easily break large parts of the internet. See Akamai outage in July.
I am not sure why they had to mention this specifically. This makes it sound like an external attack.
After all the scandals, leaks, whistleblowers etc it would take more than a DNS record wipe to take down the Facebook mafia.