Reconnect and the GLBs fall apart under load as the entire world’s cadre of recursive resolvers hit you.
Fix that. Reconnect again. This time your LBs have marked half the servers as offline because their heartbeats have been failing.
Fix that. Reconnect again. Now all the memcache data is hours old and so the site business logic fetches straight from databases, knocking them over.
Fix the databases. Reconnect again. Ad nauseam.
I can’t imagine trying to restart something as big as facebook…
Facebooks no longer an innovator, just a mining operation with a dwindling population of hateful elderly and bots.
Having been a part of Google the only thing more awe inspiring than the sheer complexity of production is the fact that it all worked so well.
This is not a dig at current Googlers but entropy is cruel and uncaring. Perhaps parts of the stack which have been kept fit & fresh in people's minds due to constant rewrites will last longer but there are tons of places in the depot that are unowned despite serving production query traffic and the number of engineers that have any context to support it grows smaller over time.
No idea if Facebook is similar.
1: https://image.slidesharecdn.com/datareportal20210308gd001dig...
2: https://image.slidesharecdn.com/datareportal20210308gd001dig...
3: https://wearesocial.com/blog/2021/01/digital-2021-the-latest...
It's pretty inexcusable that FB wasn't able to use OOB management.
I suppose with something like BGP it would be very difficult to get such a fallback working given how distributed the system is, and even more difficult to keep it exercised and tested.
https://www.juniper.net/documentation/us/en/software/junos/c...
But Facebook's network is certainly much more complex and automated than just doing one commit on one device.
But I think they do still uses Juniper devices at the edge.
<snark> Software defined networking. In PHP.
How do you know if you waited long enough to see negative effects (think caches, your own and caches of others).
Waiting too long with a bad config also costs you.
iptables-apply reverts to previous config instead of flushing all the rules.
Whereas to make an announcement, the entire internet (or at least all routers between the AS and the user) need to pickup the new announcement.
(Note: I still need to read the article)
I think it's not so simple because authoritative DNS systems are involved.
So it's not just a BGP error. It's a BGP error which disconnected authoritative DNS for all facebook. I'm not quite sure why that makes it so slow to fix. is it just because internal difficulties due to having no DNS at all?
It's an IP address management (IPAM) solution that also just happens to be a fantastic, federated (if you want) DNS management system too. Indeed a previous org I worked at bought it strictly to tame the DNS beast - local sys admins could control DNS for their subnets but not affect anything else. If we wanted, we could have had approval processes on top of the change requests - the system supported that too.
I think the security teams finally woke up to the IP address management functionality and were slowly starting to integrate that into the rest of the infrastructure - but I was leaving around then. It was a fantastic system. One of the best hierarchical role-based access control systems in an application I have ever seen; the granularity was amazing yet it was easy to understand/administer. Not an easy trick!
I would find it a bit surprising if Facebook didn't have OOB access to their data centers, however.
Next time FB save you passwords in OneDrive and Google Drive as a backup LOL. facebook-oob-password@gmail.com
Laptop + mobile tethering + serial cable to the router + teamvewier for the remote admin to get the access solves problems like this in minutes.
Breaking a gajillion security policies by doing that is a different story though.
I don’t know enough about BGP to make an informed decision; but at the point the outage is noticed it’s entirely possible that the system has been unavailable for quite some time already.
https://ns1.com/resources/dns-propagation#:~:text=DNS%20prop....
(There's a corner case related to DNSSEC that can make it go higher, but that's being worked on, and isn't relevant here.)
In this situation, the nameservers were just down. I haven't done exhaustive research, but the resolvers I'm aware of cache that kind of thing for no more than 15 minutes.