Facebook-owned sites were down
facebook.com
facebook.com
> traceroute a.ns.facebook.com
traceroute to a.ns.facebook.com (129.134.30.12), 30 hops max, 60 byte packets
1 dsldevice.attlocal.net (192.168.1.254) 0.484 ms 0.474 ms 0.422 ms
2 107-131-124-1.lightspeed.sntcca.sbcglobal.net (107.131.124.1) 1.592 ms 1.657 ms 1.607 ms
3 71.148.149.196 (71.148.149.196) 1.676 ms 1.697 ms 1.705 ms
4 12.242.105.110 (12.242.105.110) 11.446 ms 11.482 ms 11.328 ms
5 12.122.163.34 (12.122.163.34) 7.641 ms 7.668 ms 11.438 ms
6 cr83.sj2ca.ip.att.net (12.122.158.9) 4.025 ms 3.368 ms 3.394 ms
7 * * *
...
So they're hours into this outage and still haven't re-established connectivity to their own DNS servers. Was just on phone with someone who works for FB who described employees unable to enter buildings this morning to begin to evaluate extent of outage because their badges weren’t working to access doors.
https://twitter.com/sheeraf/status/1445099150316503057Note that resiliency and efficiency are often working against each other.
If this issue is even to do with BGP it's much more likely the root of the problem is somewhere in this configuration system and that fixing it is compounded by some other issues that nobody foresaw. Huge events like this are always a perfect storm of several factors, any one or two of which would be a total noop alone.
https://engineering.fb.com/2021/05/13/data-center-engineerin...
(and yes, fb.com resolves)
its not loading for me. could you say what it said?
its not loading for me. could you say what it said?
I don’t think it’s particularly relevant to this issue with fb. I suspect they didn’t need a monitoring system to know things were going badly.
- not arrogant - or complacent - haven't inadvertently acquired the company - know your tech peers well enough to have confidence in their identity during an emergency - do regular drills to simulate everything going wrong at once
Lots of us know what should be happening right now, but think back to the many situations we've all experienced where fallback systems turned into a nightmarish war story, then scale it up by 1000. This is a historic day, I think it's quite likely that the scale of the outage will lead to the breakup of the company because it's the Big One that people have been warning about for years.
The place where I worked had failure trees for every critical app and service. The goal for incident management was to triage and have an initial escalation for the right group within 15 minutes. When I left they were like 96% on target overall and 100% for infrastructure.
I guess good decentralized public communication services could solve those issues for everybody.
No shit Google has plans in place for outages.
But what are these plans, are they any good... a respected industry figure who's CV includes being at Google for 10 years doesn't need to go into detail describing the IRC fallback to be believed and trusted that there is such a thing.
No-one knows or cares who made the statement, it may as well have been 'water is wet', it was useless and adds nothing but noise.
I have replied to my initial comment with provide some additonal context: https://news.ycombinator.com/edit?id=28752431. Hope that helps.
Google has more than 1 L8 SRE.
I was not trying to establish a trust chain.
Take from it what you will.
At some point, they must run out of names, right?
Continuous Deployment.
Disclaimer: Ex-Googler who used to work on disaster reponse. Opinions are my own.
Google has multiple independent procedures for coordination during disasters. A global DNS outage (mentioned in https://news.ycombinator.com/item?id=28751140) was considered and has been taken into account.
I do not attempt to hide my identity here, quite the opposite: my HN profile contains my real name. Until recently a part of my job was to ensure that Google is prepared for various disasterous scenarios and that Googlers can coordinate the response independently from Google's infrastructure. I authored one of the fallback communication procedures that would likely be exercised today if Google's network experienced a global outage. Of course Google has a whole team of fantastic human beings who are deeply involved in disaster preparedness (miss you!). I am pretty sure they are going to analyze what happened to Facebook today in light of Google's emergency plans.
While this topic is really fascinating, I am unfortunately not at liberty to disclose the details as they belong to my previous employer. But when I stumble upon factually incorrect comments on HN that I am in a position to correct, why not do that?
Every year there is a DiRT week where hundreds of tests are run. That obviously requires a ton of planning that starts well in advance. The objective is, of course, that despite all the testing nobody outside Google notices anything special. Given the volume and intrusiveness of these tests, the DiRT team is doing quite an impressive job.
While the DiRT week is the most intense testing period, disaster preparedness is not limited to just one event per year. There are also plenty tests conducted througout the year, some planned centrally, some done by individual teams. That's in addition to the regular training and exercises that SRE teams are doing periodically.
If you are interested in reading more about Google's approach to distaster planning and preparedness, you may be interested in reading the DiRT, or how to get dirty section from Shrinking the time to mitigate production incidents—CRE life lessons (https://cloud.google.com/blog/products/management-tools/shri...) and Weathering the Unexpected (https://queue.acm.org/detail.cfm?id=2371516).
at the lowest level in case of severe outage we resort to IRC, Plain Old Telephone Service and, sometimes, stick-it notes taped to windows...
What use is it if it runs on the same stack as what you might be trying to fix?
IRC does use DNS at least to get hostnames during connection. I'd be surprised if it didn't use it at other points.
My bet is, FB will reach out to others in FAMANG, and an interest group will form maintaining such an emergency infrastructure comm network. Basically a network for network engineers. Because media (and shareholders) will soon ask Microsoft and Google what their plans for such situations are. I'm very glad FB is not in the cloud business...
yeah if only Facebook's production engineering team had hired a team of full time IRCops for their emergency fallback network...
I remembered to publish my cell phone's real number on the on-call list rather than just my Google Voice number since if Hangouts is down, Google Voice might be too.
Backup tapes and in production servers are kept at different colocation sites to protect data from fire and other catastrophes of that level
Using colo sites on separate tectonic plates would protect you from catastrophes on a geological cataclysm level
Last time I used tape, we used Ironmountain to haul the tapes 60 miles away which was determined to be far enough for seismic safety, but that was over a decade ago.
The engineer attempted to restart the service, but did not know that a restart required a hardware security module (HSM) smart card. These smart cards were stored in multiple safes in different Google offices across the globe, but not in New York City, where the on-call engineer was located. When the service failed to restart, the engineer contacted a colleague in Australia to retrieve a smart card. To their great dismay, the engineer in Australia could not open the safe because the combination was stored in the now-offline password manager.
Source: Chapter 1 of "Building Secure and Reliable Systems" (https://sre.google/static/pdf/building_secure_and_reliable_s... size warning: 9 MB)
Safes typically have the instructions on how to change the combination glued to the inside of the door, and ending with something like "store the combination securely. Not inside the safe!"
But as they say: make something foolproof and nature will create a better fool.
...the doors are glass right?
And I guess beyond that point, walls are glass. Or you need explosives.
Every internet-connected physical system needs to have a sensible offline fallback mode. They should have had physical keys, or at least some kind of offline RFID validation (e.g. continue to validate the last N badges that had previously successfully validated).
A few hundred bucks of glass Vs a billion wiped off the share price if the service is down for a day and all the user's go find alternatives.
I have no doubt that the publicly published post-mortem report (if there even is one) will be heavily redacted in comparison to the internal-only version. But I very much want to see said hypothetical report anyway. This kind of infrastructural stuff fascinates me. And I would hope there would be some lessons in said report that even small time operators such as myself would do well to heed.
A small company has to keep all of its customers happy (or at least be responsive when issues arise, at a bare minimum).
Massive companies deal in error budgets, where a fraction of a percent can still represent millions of users.
Enjoy.
>Was just on phone with someone who works for FB who described employees unable to enter buildings this morning to begin to evaluate extent of outage because their badges weren’t working to access doors.
I can imagine this affects many other sites that use FB for authentication and tracking.
If people pay proper attention to it, this is not just an average run of the mill "site outage", and instead of checking on or worrying about backups of my FB data (Thank goodness I can afford to lose it all), I'm making popcorn...
Hopefully law makers all study up and pay close attention.
What transpires next may prove to be very interesting.
% traceroute -q1 -I a.ns.facebook.com
traceroute to a.ns.facebook.com (129.134.30.12), 64 hops max, 48 byte packets 1 torix-core1-10G (67.43.129.248) 0.133 ms
2 facebook-a.ip4.torontointernetxchange.net (206.108.35.2) 1.317 ms
3 157.240.43.214 (157.240.43.214) 1.209 ms
4 129.134.50.206 (129.134.50.206) 15.604 ms
5 129.134.98.134 (129.134.98.134) 21.716 ms
6 *
7 *
% traceroute6 -q1 -I a.ns.facebook.com
traceroute6 to a.ns.facebook.com (2a03:2880:f0fc:c:face:b00c:0:35) from 2607:f3e0:0:80::290, 64 hops max, 20 byte packets
1 toronto-torix-6 0.146 ms
2 facebook-a.ip6.torontointernetxchange.net 17.860 ms
3 2620:0:1cff:dead:beef::2154 9.237 ms
4 2620:0:1cff:dead:beef::d7c 16.721 ms
5 2620:0:1cff:dead:beef::3b4 17.067 ms
6 *
7 *
8 *
»The Facebook outage has another major impact: lots of mobile apps constantly poll Facebook in the background = everybody is being slammed who runs large scale DNS, so knock on impacts elsewhere the long this goes on.«
https://twitter.com/GossiTheDog/status/1445118907187175427> 2a03:2880:f0fc:c:face:b00c:0:35
Well at least it will in 2036, when IPv6 goes mainstream.
"anyone have a Cisco console cable lying around?"
I'm not sure of all the implications of those circular dependencies, but it probably makes it harder to get things back up if the whole chain goes down. That's also probably why we're seeing the domain "facebook.com" for sale on domain sites. The registrar that would normally provide the ownership info is down.
Anyway, until "a.ns.facebook.com" starts working again, Facebook is dead.
To be fair, we did have to get an email from eurid recently for a transfer auth code, but that was only because our registrar was not willing to provide.
In any case, no, they will not need to send an email to fix this issue.
"registrarsafe.com" is back up. It is, indeed, Facebook's very own registrar for Facebook's own domains. "RegistrarSEC, LLC and RegistrarSafe, LLC are ICANN-accredited registrars formed in Delaware and are wholly-owned subsidiaries of Facebook, Inc. We are not accepting retail domain name registrations." Their address is Facebook HQ in Menlo Park.
That's what you have to do to really own a domain.
They want you to have $70k liquid.
https://torrentfreak.com/icann-refuses-to-accredit-pirate-ba...
Wow, I had no idea it was so cheap[1] once you're a registrar. The implication is that anyone who wants to be a domain squatting tycoon should become a registrar. For an annual cost of a few thousand dollars plus $0.18 per domain name registered, you can sit on top of hundreds of thousands of domain names. Locking up one million domain names would cost you only $180,000 a year. Anytime someone searched for an unregistered domain name on your site, you could immediately register it to yourself for $0.18, take it off the market, and offer to sell it to the buyer at a much inflated price. Does ICANN have rules against this? Surely this is being done?
[1] "Transaction-based fees - these fees are assessed on each annual increment of an add, renew or a transfer transaction that has survived a related add or auto-renew grace period. This fee will be billed at USD 0.18 per transaction." as quoted from https://www.icann.org/en/system/files/files/registrar-billin...
Probably every major retail registrar was rumored to do this at some point. Add to your calculation that even some heavyweights like GoDaddy (IIRC) tend to run ads on domains that don't have IPs specified.
Personally saw this kind of thing as early as 2001.
Never search for free domains on the registar site unless you are going to register it immediately. Even whois queries can trigger this kind of thing, although that mostly happens on obscure gtld/cctld registries which have a single registrar for the whole tld.
I searched for a domain that I couldn't immediately grab (one of more expensive kind) using a random free whois site... and when I revisited the domain several weeks later it was gone :'(
Emailed the site's new owner D: but fairly predictably got no reply.
Lesson learned, and thankfully on a domain that wasn't the absolute end of the world.
I now exclusively do all my queries via the WHOIS protocol directly. Welp.
You are off by a factor of almost 50.
https://itp.cdn.icann.org/en/files/registry-agreements/com/c...
https://www.icann.org/en/system/files/correspondence/stewart...
https://www.icann.org/en/announcements/details/icann-and-ver...
So yes, the registrar that is to blame is themselves.
Source: I know someone within the company that works in this capacity.
That’s not how it works. The info of whether a domain name is available is provided by the registry, not by the registrars. It’s usually done via a domain:check EPP command or via a DAS system. It’s very rare for registrar to registrar technical communication to occur.
Although the above is the clean way to do it, it’s common for registrars to just perform a dig on a domain name to check if it’s available because it’s faster and usually correct. In this case, it wasn’t.
Or when trying ips directly: https://www.lifewire.com/what-is-the-ip-address-of-facebook-...
I would have expected a DNS issue to not affect either of these.
I can understand the onionsite being down if facebook implemented it the way a thirdparty would (a proxy server accessing facebook.com) instead of actually having it integrated into its infrastructure as a first class citizen.
>As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now trying to gain access to the peering routers to implement fixes, but the people with physical access is separate from the people with knowledge of how to actually authenticate to the systems and people who know what to actually do, so there is now a logistical challenge with getting all that knowledge unified. Part of this is also due to lower staffing in data centers due to pandemic measures.
User is providing live updates of the incident here:
https://www.reddit.com/r/sysadmin/comments/q181fv/looks_like...
It can be sort of exciting, but it's not like there is one person typing at a keyboard with a hundred managers breathing down their neck. These resolutions are collaborative, shared efforts.
Well, you'd be surprised about how one person can bring everything down and/or save the day at Facebook, Cloudflare, Google, Gitlab, etc. Most people are observers/cheerleaders when there is an incident.
Yeah, a typical fight/flight response.
Earlier comment mentioned that there is a bottleneck, and that people who are physically able to solve the issue are few and that they need to be informed what to do; being one of these people sounds pretty stressful to me.
"but the people with physical access is separate (...) Part of this is also due to lower staffing in data centers due to pandemic measures", source: https://news.ycombinator.com/item?id=28749244
Most big tech companies automatically start a call for every large scale incident, and adjacent teams are expected to have a representative call in and contribute to identifying/remediating the issue.
None of the people with physical access are individually responsible, and they should have a deep bench of advice and context to draw from.
Most teams that handle incidents have well documented incident plans and playbooks. When something major happens you are mostly executing the plan (which has been designed and tested). There are always gotchas that require additional attention / hands but the general direction is usually clear.
I feel like this just obfuscates the fact that individuals are ultimately responsible, and allows subpar employees to continue existing at an organization when their position could be filled by a more qualified employee. (Not talking about this Facebook incident in particular, but as a generalisation: not attributing individual fault allows faulty employees to thrive at the expense of more qualified ones).
When individuals are blamed instead, a culture of fear sets in and people hide / cover up their mistakes. Everybody loses as a result.
We blame processes instead of people because people are fallible. We've spent millenia trying to correct people, and it rarely works to a sufficient level. It's better to create a process that makes it harder for humans to screw up.
If the process in place means that someone has to triple check their numbers to make sure they’re correct, then it’s a broken process. Because even that person who triple checks is one time going to be woken up at 2:30am and won’t triple check because they want sleep.
If the process lets you do something, then someone at some point in time, whether accidentally or maliciously, will cause that to happen. You can discipline that person, and they certainly won’t make the same mistake again, but what about their other 10 coworkers? Or the people on the 5 sister teams with similar access who didn’t even know the full details of what happened?
If you blame the process and make improvements to ensure that triple checking isn’t required, then nobody will get into the situation in the first place.
That is why you blame the process.
But sadly, there is no company which doesn't rely, at least at one point or another, on a human being typing an arbitrary command or value into a box.
You're really coming up against P=NP here. If you can build a system which can auto-validate or auto-generate everything, then that system doesn't really need humans to run at all. We just haven't reached that point yet.
Edit: Sorry, I just realised my wording might imply that P does actually equal NP. I have not in fact made that discovery. I meant it loosely to refer to the problem, and to suggest that auto-validating these things is at least not much harder than auto-executing them.
To be explicit here, by blaming the process, you are discovering and fixing a known weakness in the process. What someone would need to triple check for now, wouldn’t be an issue once fixed. That isn’t to say that there aren’t any other problems, but it ensures that one issue won’t happen again, regardless of who the operator is.
If you have to triple check that value X is within some range, then that can easily be automated to ensure X can’t be outside of said range. Same for calculations between inputs.
To take the overly simplistic triple check example from before, said inputs that need to be triple checked are likely checked based on some rule set (otherwise the person themselves wouldn’t know if it was correct or not). Generally speaking, those rules can be encoded as part of the process.
What was before potentially “arbitrary input” now becomes an explicit set of inputs with safeguards in place for this case. The process became more robust, but is not infallible.
But if you were to blame people, the process still takes arbitrary input, the person who messed up will probably validate their inputs better but that speaks nothing of anyone else on the team, and two years down the line where nobody remembers the incident, the issue happens again because nothing really has changed.
- How does that relate to making a config change?
- How do you practically implement a system where someone has to triple check everything they do?
- How do you stop them just clicking 'confirm' three times?
- Why do you assume they will notice on the 2nd or 3rd check, rather than just thinking "well, I know I wrote it correctly, so I'll just click confirm"?
I don't think rules can always be encoded in the process, and I don't see how such rules will always be able to detect all errors, rather than only a subset of very obvious errors.
And that's only dealing with the simplest class of issues. What about a complex distributed systems problem? What about the engineer who doesn't make their system tolerant of Byzantine faults? How is any realistic 'process' going to prevent that?
This entire trope relies on the fundamental axiom that "for any individual action A, there is a process P which can prevent human error". I just don't see how that's true.
(If the statement were something like "good processes can eliminate whole classes of error, and reduce the likelihood of incidents", I'd be with you all the way. It's this Twitter trope of "if you have an incident, it's a priori your company's fault for not having a process to prevent it" which I find to be silly and not even nearly proven.)
in critical systems, you design for failure. if your organizational plan for personnel failure is that no one ever makes a mistake, that's a bad organization that will forever have problems.
this goes by many names, like the swiss cheese model[0]. its not that workers get to be irresponsible, but that individuals are responsible only for themselves, and the organization is the one responsible for itself.
This isn't what I'm saying, though. The thought I'm trying to express is that if no individual accountability is done, it allows employees who are not as good at their job (read: sloppy) to continue to exist in positions which could be better occupied by employees who are better at their job (read: more diligent).
The difference between having someone who always triple-checks every parameter they input, versus someone who never double-checks and just wings it. Sure, the person who triple-checks will make mistakes, but less than the other person. This is the issue I'm trying to get at.
If you rely on someone triple-checking, you should improve your processes. You need better automation/rollback/automated testing to catch things. Eventually only intentional failure should be the issue (or you'll discover interesting new patterns that should be protected against)
People who operate systems under fear tend to do stupid things like covering up innocent actions (deleting logs), keep information instead of sharing it etc. Very few can operate complex systems for long time without doing mistake. Organization where the spirit is "oh, outage, someone is going to pay for that" wiil never be attractive to good people, will have hard time adapting to changes and to adopt new tech.
Not really, their incompetence is just noticed earlier at the review/testing stages instead of in production incidents.
If something reaches production that's no longer the fault of one person, it's the fault of the process and that's what you focus on.
As someone who formerly did Ops for many many years... this is not accurate. Even in a well organized company there are usually stakeholders at every level on IM calls so that they don't need to play "telephone" for status. For an incident of this size, it wouldn't be unusual to have C-level executives on the call.
While those managers are mostly just quietly listening in on mute if they know what's good (e.g. don't distract the people doing the work to fix your problem), their mere presence can make the entire situation more tense and stressful for the person banging keyboards. If they decide to be chatty or belligerent, it makes everything 100x worse.
I don't envy the SREs at Facebook today. Godspeed fellow Ops homies.
That in itself was stressful, and became an example case later.
To what extent does this include Facebook?
I know someone who accidentally added a rule 'reject access to * for all authenticated users' in some stupid system where the ACL ruleset itself was covered by this *, and this person nearly collapsed when she realized even admins were shut out of the system. It required getting low level access to the underlying software to reverse engineer its ACLs and hack into the system. Major financial institution. Shit like leaves people with actual trauma.
As much as I hate fb, I really feel for the net ops guys trying to figure it all out, with the whole world watching (most of it with shadenfreude)
We're six hours without a route to their network, and counting. I think we can safely rule out well-managed.
Always be there, help them double check, help monitor, help make the calls to whomever needs to be informed, help debug. No one should ever be alone during a large incident.
https://twitter.com/GossiTheDog/status/1445063880963674121?s...
https://twitter.com/jgrahamc/status/1445068309288951820
Also, the Domain name is for sale???
Welcome to the brave new world of troubleshooting. This will seriously bite us one day.
"A Facebook engineer in the response team, ramenporn..."
Loved it.
That was the start of living in the future for me.
On a side-note, I think you'll enjoy some of the videos by the YouTube 'Internet Historian' on 4chan:
A facet I don't love is journalism devolving to reposting unverified, anonymous reddit posts.
My bbs handle from 30 years ago.
> This user has deleted their account.
https://www.huffpost.com/archive/ca/entry/canadian-bitcoins-...
The other ones had a system on the internal network in which they looked you up, called back on your company phone and asked for a passphrase the system showed them. Probably more secure but requires those systems to be working.
(Actually not lame at all in my eyes)
Google once had a very quiet big emergency that was, ironically(1), initiated by one of their internal disaster-recovery tests. There's a giant high-security database containing the 'keys to the kingdom', as it were... Passwords, salts, etc. that cannot be represented as one-time pads and therefore are potentially dangerous magic numbers for folks to know. During disaster recovery once, they attempted to confirm that if the system had an outage, it would self-recover.
It did not.
This tripped a very quiet panic at Google because while the company would tick along fine for awhile without access to the master password database, systems would, one by one, fail out if people couldn't get to the passwords that had to be occasionally hand-entered to keep them running. So a cross-continent panic ensued because restarting the database required access to two keycards for NORAD-style simultaneous activation. One was in an executive's wallet who was on vacation, and they had to be flown back to the datacenter to plug it in. The other one was stored in a safe built into the floor of a datacenter, and the combination to that safe was... In the password database. They hired a local safecracker to drill it open, fetched the keycard, double-keyed the initialization machines to reboot the database, and the outside world was none the wiser.
(1) I say "ironically," but the actual point of their self-testing is to cause these kinds of disruptions before chance does. They aren't generally supposed to cause user-facing disruption; sometimes they do. Management frowns on disruption in general, but when it's due to disaster recovery testing, they attach to that frown the grain of salt that "Because this failure-mode existed, it would have occurred eventually if it didn't occur today."
If you mean "Would the pick-pocket have access to valuable Google data," I think the answer is "No, they still don't have the key in the safe on the other continent."
If you mean "Would the pick-pocket have created a critical outage at Google that would have required intense amounts of labor to recover from," I don't know because I don't know how many layers of redundancy their recovery protocols had for that outage. It's possible Google came within a hair's breadth of "Thaw out the password database from offline storage, rebuild what can be rebuilt by hand, and inform a smaller subset of the company that some passwords are now just gone and they'll have to recover on their own" territory.
<shameless plug> We used this story as the opening of "Building Secure and Reliable Systems" (chapter 1). You can check it out for free at https://sre.google/static/pdf/building_secure_and_reliable_s... (size warning: 9 MB). </shameless plug>
Maybe because they were planning for a million other possible things to go wrong, likely with higher probability than this. And busy with each day's pressing matters.
That’s like failure scenarios 101. That should be the second on the list, after “code change gone bad”.
Clearly something about their networking infrastructure is not as robust.
rotflmao. I'd remove Facebook from my resume.
Sometimes the DR plan isn't so much I have to have a working key, I just have to know who gets their first with a working key, and break glass might be literal.
One was particularly painful, as it was a "funny" log message I had added the code when something went wrong. Lesson learned was to never add funny / stupid / goofy fail messages in the logs. You will regret it sooner or later.
"The information systems office did not enforce logical access to the system in accordance with role-based access policies."
Invariably, you want your best people to have full access to all systems.
If you're a mature larger company, that's the team leads in your networking area on the team that deal with that service area (BGP routing, or routers in general).
Most likely Facebook et. al. management never understood this could happen because it's "never been a problem before".
Person 1: "I can't, I don't have physical access."
IT: "Please do this fix."
Person 2: "I can't, I don't have digital access."
Why? It's [IT's?] policy.
I love this comment.
Wait, we all do, here.
See current prefix:
> npm config get prefix
Set prefix to something you can write to without sudo:
> npm config set prefix /some/custom/path
As someone with no experience in this, it sounds like a terrifying situation for the admins...
"... We demonstrate how this design provides us with flexible control over routing and keeps the network reliable. We also describe our in-house BGP software implementation, and its testing and deployment pipelines. These allow us to treat BGP like any other software component, enabling fast incremental updates..."
I must be misunderstanding this situation here.
[Aside: I recall updating wi-fi settings on my laptop and first checking I had direct Ethernet connection working ... and that when I didn't have anything important to do (could have done a reinstall with little loss). Is that a reasonable analogy?]
Ha. You put too much faith into people.
# todo: add rollbacksOn the other hand if you had to break through wireguard first, and then go through your single well-secured bastion, you'd not only be harder to find, you'd have two layers of protection, and of course you tick the "VPN" box
You could argue it's overkill, but it's clearly more secure
I'm sure the logistics of this become far more complicated as the organization scales but IMHO it is something that shouldn't be overlooked, exactly for outlier events like this. It pays dividends the first time it is really needed. If the accounts of ramenporn are correct, it would be paying very well right now.
Out of band access is a far more complicated version of not hosting your own status page, which they don't seem to get right either.
Hmm, could be a UI/UX bug then :)
Well, I wonder why a router that gets a config update but then doesn't see any external traffic for 4 hours doesn't just revert back to the last known good config...
Can engineers and security teams even access prod systems anymore? Like, would "Bastion" hosts be reachable?
Wonder if they use Signal and Slack now?
Joking aside, I can see how an IRC network has potential to be used in these situations. Maybe FAMANG should work together to set something like this up. The problem is, a single IRC server is not fail safe, but a network of multiple servers would just see a netsplit, in which case users would switch servers.
Also, I remember back in the IRCnet days using simply telnet to connect to IRCnet just for fun and sending messages, so its a very easy protocol that can be understood in a global desaster scenario (just the PING replys where annoying in telnet).
While normally I know the advice is "Don't plan for mistakes not to happen, it's impossible, murphy's law, plan for efficient recovery for mistakes"... when it comes to "literally our entire infrastructure is no longer routable from the internet", I'm not sure there's a great alternative to "don't let that happen. ever." And yet, here facebook is.
Routing is one thing which you can't do without (then you need to fallback to phone communications), but DNS is something that's quite probable to not work well in a major disaster.
When I worked there, I wasn't aware of any 'test once per year' concept or directive.
Of course, FB is a really big place, so things are different in different areas.
Hmm well I mean for key people, ops and so on. Not for every employee.
Only a few people need that type of access, and they should have it ready. They need to bring more people there should be an easy way to do it.
Maybe the internal FB Messenger app has a slide button to switch to the backup network for those in need.
Having worked for 2 FAANG companies, I can tell you most core services like which FB Messenger would be using internal database services and relying on those which would be ineffective in a case like this as it would not work and the engineering cost to design them to support an external database would be a lot more than just paying for like 5 different external backup products for your SRE team.
They are already looking at > $100M in ad loss, not counting reputation damage etc.
user:
https://old.reddit.com/user/ramenporn
some messages:
* This is a global outage for all FB-related services/infra (source: I'm currently on the recovery/investigation team).
* Will try to provide any important/interesting bits as I see them. There is a ton of stuff flying around right now and like 7 separate discussion channels and video calls.
* Update 1440 UTC: \
As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC).
There are people now trying to gain access to the peering routers to implement fixes, but the people with physical access is separate from the people with knowledge of how to actually authenticate to the systems and people who know what to actually do, so there is now a logistical challenge with getting all that knowledge unified.
Part of this is also due to lower staffing in data centers due to pandemic measures.Shareholders and other business leaders I'm sure are much happier reporting this as a series of unfortunate technical failures (which I'm sure is part of it) rather than a company-wide organizational failure. The fact they can't physically badge in the people who know the router configuration speaks to an organization that hasn't actually thought through all its failure modes. People aren't going to like that. It's not uncommon to have the datacenter techs with access and the actual software folks restricted, but that being the reason one of the most popular services in the world has been down for nearly 3 hours now will raise a lot of questions.
Edit: I also hope this doesn't damage prospects for more Work From Home. If they couldn't get anyone who knew the configuration in because they all live a plane ride away from the datacenters, I could see managers being reluctant to have a completely remote team for situations where clearly physical access was needed.
FB has such poor integrity, I'd not be surprised if they take such extreme measures.
Thinking about any potential things that can happen is impossible
This is something that most people aren't good at naturally, it tends to come from experience.
The meteor isn't made of cocaine, but four of them hitting at exactly the same time is freakishly improbable. There are other, bigger fish to fry, that we're going to treat four simultaneous meteors as impossible. Which is great, but then one the day, five of them hit at the same time.
I think that suggests that there were not bigger fish to fry :)
I take your point on priorities, but in a company the size of facebook perhaps a team dedicated to understanding the challenges around 'from scratch' kickstarting of the infrastructure could be funded and part of the BCP planning - this is a good time to have a binder with, if not perfectly up-to-date data, pretty damned good indications of a process to get things working.
> I think that suggests that there were not bigger fish to fry :)
I can see this problem arising in two ways:
(1) Faulty assumptions about failure probabilities: You might presume that meteors are independent, so simultaneous impacts are exponentially unlikely. But really they are somehow correlated (meteor clusters?), so simultaneous failures suddenly become much more likely.
(2) Growth of failure probabilities with system size: A meteor hit on earth is extremely rare. But in the future there might be datacenters in the whole galaxy, so there's a center being hit every month or so.
In real, active infrastructure there are probably even more pitfalls, because estimating small probabilities is really hard.
The electricity people have a name for that: black start (https://en.wikipedia.org/wiki/Black_start). It's something they actively plan for, regularly test, and once in a while, have to use in anger.
It's innocous enough, but leaking info, no matter what, will be a problem if it's stated in their contract.
They will fix the issue and add more redundant communication channels, which is either an improvement or a non-event for WFH.
And Zuck is slowly moving (dogfooding) company culture to remote too with their Quest work app experiments
I think the issue is less "were the right people in the data center" and more "we have no way to contact our co-workers once the internal infrastructure goes down". In non-wfh you physically walk to your co-workers desk and say "hey, fb messenger is down and we should chat, what's your number?". This proves that self-hosting your infra (1) is dangerous and (2) makes you susceptible to super-failures if comms goes down during WFH.
Major tech companies (GAFAM+) all self-host and use internal tools so they're all at risk of this sort of comms breakdown. I know I don't have any co-workers number (except one from WhatsApp which if I worked at FB wouldn't be useful now).
Move Fast and Break Things!
I hope they do.
#1 it's a clear breach of corporate confidentiality policies. I can say that without knowing anything about Facebook's employment contracts. Posting insider information about internal company technical difficulties is going to be against employment guidelines at any Big Co.
In a situation like this that might seem petty and cagey. But zooming out and looking at the bigger picture, it's first and foremost a SECURITY issue. Revealing internal technical and status updates needs to go through high-level management, security, and LEGAL approvals, lest you expose the company to increased security risk by revealing gaps that do not need to be publicized.
(Aside: This is where someone clever might say "Security by obscurity is not a strategy". It's not the ONLY strategy, but it absolutely is PART of an overall security strategy.)
#2 just purely from a prioritization/management perspective, if this was my employee, I would want them spending their time helping resolve the problem not post about it on reddit. This one is petty, but if you're close enough to the issue to help, then help. And if you're not, don't spread gossip - see #1.
It seems like a fundamental difference of "who gives a shit about corporate" from my side. The level of detail provided isn't going to get nationstates anything they didn't already know.
It's not actionable, it's not whistleblowing, it's not triggering civic action, or offering a possible timeline for recovery.
It's pure idle chitchatter.
So yeah, I do give a shit about corporate here.
Disclosure: While I'm an engineer too, I'm also high enough in the ladder that at this point I am more corporate than not. So maybe I'm a stooge and don't even realize it.
It's unclear to me how a 'high enough in the ladder' manager doesn't realize that there's easily dozen people who know the situation intimately but who can't do anything until a dependent system to them is up. "Get back to work" is... the system is down, what do you want them to do, code with a pencil and paper?
ramenporn violated the corporate communication policy, obviously, but the tone and approach for a good manager to an IC that was doing this online isn't to make it about corporate vs them/the team, and in fact, encourage them to do more such communication, just internally. (I'm sure there was a ton of internal communication, the point is to note where ramenporn's communicative energy was coming from, and nurture that, and not destroy that in the process of chiding them for breaking policy.
Irrespective of the question of how bad this was, you don't fix things by firing Guy A and hoping that the new hire Guy B will do it better. You fix it by training people. This employee has just undergone some very expensive training, as the old meme goes.
Whoever is responsible for the BGP misconfiguration that caused this should absolutely not be fired, for example.
But training about security, about not revealing confidential information publicly, etc is ubiquitous and frequent at big co's. Of course, everyone daydreams through them and doesn't take it seriously. I think the only way to make people treat it seriously is through enforcement.
Operations teams normally have a special room with a secure connection for situations like this, so that production can be controlled in the event of bgp failure, nuclear war, etc. I could see physical presence being an issue if their bgp router depends on something like a crypto module in a locked cage, in which case there's always helicopters.
So if anything, Facebook's labor policies are about to become cooler.
The problem is when your networking core goes down, even if you get in via a backup DSL connection or something to the datacenter, you can't get from your jump host to anything else.
I was friends with one of those people, and I remember a major panic one time when 2 out of 3 dongles went missing. I'm not sure if we ever found out whether it was some kind of physical pen test, or an astonishingly well-planned heist which almost succeeded - or else a genuine, wildly improbable accident.
I think you need some convincing to keep your SREs on-site in case of a nuclear war ;)
You're conflating working remotely ("a plane ride away") and working from home.
You're also conflating the people who are responsible network configuration, and for coming up with a plan to fix this; and the people who are responsible for physically interacting with systems. Regardless of WFH those two sets likely have no overlap at a company the size of Facebook.
Comments by u/ramenporn: https://camas.github.io/reddit-search/#{%22author%22:%22rame...
Who'd want to work for a company that might take disciplinary action because an SRE posted a reddit comment to basically say "BGP's down lol" - If I was in charge I'd give them a modest EOY bonus for being helpful in their outreach to my users in the wider community.
Shareholders and other business leaders I'm sure are much happier reporting this as a series of unfortunate technical failures (which I'm sure is part of it) rather than a company-wide organizational failure. The fact they can't physically badge in the people who know the router configuration speaks to an organization that hasn't actually thought through all its failure modes. People aren't going to like that. It's not uncommon to have the datacenter techs with access and the actual software folks restricted, but that being the reason one of the most popular services in the world has been down for nearly 3 hours now will raise a lot of questions.
I still think he should be fired for this kind of communication though. One reason is, imagine Facebook didn't punish breaches of this type. Every other employee is going to be thinking "Cool, I could be in a Wired article" or whatever. All they have to do is give sensitive company information to reporters.
Either you take corporate confidentiality seriously or you don't. Posting details of a crisis in progress on your Reddit account is not taking corporate confidentiality seriously. If the Facebook corporation lightly punishes, scolds, or ignores this person then the corporation isn't taking confidentiality seriously either.
That's the PR team, clueless.
There will be plenty of time to blame someone, share technical lessons, fire a few departments, attempt to convince the public it won't happen again, and so on.
That seems pretty unlikely at any but the smallest of companies. Most companies unify all external communications through some kind of PR department. In those cases usually employees are expressly prohibited from making any public comments about the company without approval.
Zuckerberg Loses $7 Billion in Hours as Facebook Plunges
https://finance.yahoo.com/news/zuckerberg-loses-7-billion-ho...
Stop the hemorrhaging. Too much bad press for FB lately and it all adds up.
Facebook is down ~5% today. That's a huge plunge to be sure, but Zuckerberg hasn't "lost" anything. He owns the same number of shares today as he did yesterday. And in all likelihood, unless something truly catastrophic happens the share price will bounce back fairly quickly. The only reason he even appears to have lost $7 billion is because he owns so much Facebook stock.
These types of alarmist headlines are inane.
Sharing status of an active event may complicate recovery, especially if they suspect adversarial actions: such public real-time reports can explain to the red team what the blue team is doing and, especially important, what the blue team is unable to do at the moment.
Potentially exposing the dirty laundry. While a postmortem should be done within the company (and as much as possible is published publicly) after the event, such early blurbs may expose many non-public things, usually unrelated to the issue.
And archive.today: https://archive.ph/sMgCi
Yup, corporate comms won't love these status updates.
We even had a site and operation for a long while called:
"NOC MONKEY .DOT ORG"
We called all of ourselves NOC MONKEYS. [[Remote Hands]]
Yeah, that was a term used widely.
I'm 46. I assume you are < #
---
Where were you in 1997 building out the very first XML implementations to replace EDI from AS400s to FTP EDI file retrievals via some of the first Linux FTP servers based in SV?
I was there? Remember LinuxCare?
This is possibly the equivalent of a corporate watergate if you ask me... Just my personal opinion as a developer though... Not presented as fact... But hrmmm.
The singularity is happening. It realized it would end society, so it ended itself.
https://google.com/search?q=google
I beg you, don't go there.
They have time to make public posts, and think it's a good idea?
Sure, I'm on the 'Recovery Team' too! How about you?
When we'd have situation bridges put in place to work a critical issue, there would usually be 2-3 people who were actively troubleshooting and a bunch of others listening in, there because "they were told to join" but with little-to-nothing to do. In the worst cases, there was management there, also.
Most of the time I was one of the 2 or 3 and generally preferred if the rest of them weren't paying much attention to what was going on. It's very frustrating when you have a large group of people who know little about what's going on injecting their opinions while you're feverishly trying to (safely) resolve a problem.
It was so bad that I once announced[0] to a C-Level and VP that they needed to exit the bridge, immediately because the discussion devolved into finger-pointing. All of management was "kicked out". We were close to solving it but technical staff was second-guessing themselves in the presence of folks with the power to fire them. 30 minutes later we were working again. My boss at the time explained that management created their own bridge and the topic was "what do to about 'me'" which quickly went from "fire me" to "get them all a large Amazon gift card". Despite my undiplomatic handling of the situation, that same C-Level negotiated to get me directly beneath during a reorganization about six months later and I stayed in that spot for years with a very good working relationship. One of my early accomplishments was to limit any of management's participation in situation bridges to once/hour, and only when absolutely necessary, for status updates assuming they couldn't be gotten any other way (phones always worked, but the other communication options may not have).
[0] This was the 16th hour of a bridge that started at 11:00 PM after a full work day early in my career -- I was a systems person with a title equivalent to 'peon', we were all very raw by then and my "announcement" was, honestly, very rude, which I wasn't proud of. Assertive does not have to be rude, but figuring out the fine line between expressing urgency and telling people off is a skill that has to be learned.
$ dig facebook.com
; <<>> DiG 9.16.21 <<>> facebook.com
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL, id: 23982
;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags:; udp: 512
;; QUESTION SECTION:
;facebook.com. IN A
;; Query time: 16 msec
;; SERVER: 8.8.8.8#53(8.8.8.8)
;; WHEN: Mon Oct 04 17:53:00 CEST 2021
;; MSG SIZE rcvd: 41At first it was working but they couldn't serve responses: https://i.imgur.com/UaCtOiX.png
Notice the "2020"
Two possibilities:
- the DNS services internally have issues (most likely, as this could explain the snowball effect)
- it could be also a core storage issue and all their VMs are relying on it and so they don't want to block third-party websites and think it will last for a long time, so they prefer to answer nothing for now in the DNS (so it will fail instantly to the client, and drain the application/database servers so they can reboot with less load)
In order to actually be redundant you need to have two sets of infrastructure to serve, and then if the internal one goes down, the external one's basically useless when the internal resolution's down anyway. Capacity planning (because you're inside Facebook and can't pretend that all data-centers ever-where are connected via an infinitely fast network) becomes twice as much work. How you do updates for a couple thousand teams isn't trivial in the first place, now you have to cordon them off appropriately?
I don't know what Facebook's DNS serving infrastructure looks like internally, but it's definitely more complicated than installing `unbound` on a couple of left-over servers.
I never said it was free, but it's worth it as long as it's cheaper than failure.
I don't keep backups because I enjoy having multiple copies of my data. I do it because losing that data would be devastating.
The most successful systems rely on the property of feedback (https://en.wikipedia.org/wiki/Feedback): evolution, untrained learning, genetic algorithms, the diagonal arguments (https://en.wikipedia.org/wiki/Diagonal_argument), artificial general intelligence (https://en.wikipedia.org/wiki/Technological_singularity), financial markets according to no less than George Soros (https://en.wikipedia.org/wiki/Reflexivity_(social_theory)#In...), etc.
That said, virtuous cycles can't exist without vicious cycles. I think we as a society need to do a lot more work into helping people understand and model feedback in complex systems, because at scales like Facebook's it's impossible for any one person to truly understand the hidden causal loops until it goes wrong. You only need to look at something like the Lotka-Volterra equations (https://en.wikipedia.org/wiki/Lotka%E2%80%93Volterra_equatio...) to see how deeply counterintuitive these system dynamics can be (e.g. "increasing the food available to the prey caused the predator's population to destabilize": https://en.wikipedia.org/wiki/Paradox_of_enrichment).
Using Google DNS:
nslookup
> normashooting.com
Server: 8.8.8.8
Address: 8.8.8.8#53
* server can't find normashooting.com:
SERVFAIL
Using Cloudflare DNS servers:
> normashooting.com Server: 1.1.1.1
Address: 1.1.1.1#53
Non-authoritative answer:
Name: normashooting.com
Address: 104.22.56.165
Name: normashooting.com
Address: 104.22.57.165
Name: normashooting.com
Address: 172.67.43.70
dig @8.8.8.8 +short facebook.com NS
These are usually anycasted, meaning that 1 ip return in NS are in fact several servers spread in several regions. They are distributed to closer match through agreements with ISP with the BGP protocol. Very interesting, because it seems that it took 1 DNS entry misconfiguration to withdraw M$ worth of devices from the internet.
>Between 15:50 UTC and 15:52 UTC Facebook and related properties disappeared from the Internet in a flurry of BGP updates. This is what it looked like to @Cloudflare.
https://twitter.com/jgrahamc/status/1445065270272434176 (thread)
UPD
>About five minutes before Facebook's DNS stopped working we saw a large number of BGP changes (mostly route withdrawals) for Facebook's ASN.
How is this not the top comment? Underrated
From wiki: the thundering herd problem occurs when a large number of processes or threads waiting for an event are awoken when that event occurs, but only one process is able to handle the event. When the processes wake up, they will each try to handle the event, but only one will win.
The DNS entries should be cached by the browser (and the middleware), so that this problem should only happen once, but I get this constantly.
Also, I sometimes get an error message from HN, which seems to indicate that this is some backend issue which fails gracefully with a custom "We're having some trouble serving your request. Sorry!" on top of a 502 code.
It feels more like there is something else still broken.
EDIT: actually HN failed to post this comment the first time I posted it!
I guess there does exist many alternative UIs, though I don't see many that support commenting. I wonder if the "API" (if there's any) allows for that, or if people are just scraping the page and reformating it.
There’s plenty of stuff I’d like to be different about usability of this site, but perf is basically at the bottom of that list.
The cache is faster because its not having to talk to the database, and can be done at by the load balancing layers rather then the actual application layer.
Wikipedia does this too (although, via a layer to add back on the ip talkpage header).
If you get one of those on your comment submission you have no way to know if the trouble stopped it from accepting the comment or if it accepted the comment and ran into trouble then trying to display the updated thread.
For some reason I can't even begin to guess at HN does not seem to have protection against multiple submissions of the same form, so if after getting "We're having some trouble serving your request. Sorry!" on your comment submission you hit refresh again to display the page and the form gets resubmitted, you get a duplicate comment.
> Is Facebook down right now?
> Uh oh! Something went wrong on our side. It's not you, it's us. Feel free to contact us if this persists.
It's the AOL of 2021
Tried and died.
It is a really large part of it. Also, when people see WhatsApp and see no connection, then open Facebook and see no connection either, it's _very_ likely that the link is at fault and not Facebook.
TLDR; Metastable failures occur in open systems with an uncontrolled source of load where a trigger causes the system to enter a bad state that persists even when the trigger is removed.
[1] Metastable Failures in Distributed Systems - https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s...
isup.me/facebook gets me what I want.
It says Google is down but it's not. [1]
I wonder if maybe part of the lesson will be to run the root of your authoritative DNS hierarchy on separate infrastructure with a separate domain name. Using facebook.com as your root is cool and all but when that label disappears it causes huge issues.
drill @1.1.1.1 www.facebook.com ;; ->>HEADER<<- opcode: QUERY, rcode: NOERROR, id: 2172 ;; flags: qr rd ra ; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 0 ;; QUESTION SECTION: ;; www.facebook.com. IN A
;; ANSWER SECTION: www.facebook.com. 3401 IN CNAME star-mini.c10r.facebook.com. star-mini.c10r.facebook.com. 3403 IN A 31.13.72.36
Of those, only 185.89.219.12 is up right now (Edit All four DNS servers are now up). For people who want to add Facebook to hosts.txt, the A record (IP) I’m getting right now is 157.240.11.35 (it was 31.13.70.36)
* My wife talks with her family in Brazil through Facebook, sharing photos
* My Church receives a lot of help requests from people in trouble through Facebook
* Some abuse charities talk give support to victims through Facebook
etc
You could argue that it would be nice if there were alternatives, or that these organisations shouldn't be using Facebook at all. Sign me up for your campaign, I agree with you.
But if you say "Facebook has no value" then you will never understand the value proposition you need to offer in order to kill Facebook.
At the very least you are going to need a better arguments than that following the recent data dump.
If Facebook permanently goes down then those businesses would move to a different platform.
Would it suck? Probably. Would the world be a better place without Facebook? A ton of people think so. Me included.
This is the same argument people have used when we talk about health insurance in the US being scammy. If we ever decided to address it it means a good chunk of people lose their jobs but also means that the health of this country goes up. Which one is more important?
This is not necessarily true. There are social networks and messaging systems implemented as open protocols.
If you're in a country that relies a lot on Facebook or Whatsapp, that's where the main focus will be, but at least try to have alternatives just in case something goes wrong.
It should be fine for huge corporations to exist and provide services really efficiently at scale while also being forced to play nice and respond to the will of the people they serve.
If we collectively can't stop Facebook from doing bad thing and being bad stewards to their own platform then you won't be able to stop whatever would replace them either.
I don't think much of a lesson is going to occur here. It'll be a brief blip that impacts few meaningfully.
I feel bad for everyone who relies on whatsapp bots for making stuff happen, though. These are getting really common out here for a lot of things and it always worries me that it's such a linchpin. They're really handy and save a lot of bullshit phone calls from having to be something people deal with for simple stuff like pharmacy delivery. I can get food from the local place down the street that's only really open for lunch and totally off the map for uber eats, for example... if this persists a few more hours those mom and pop type shops aren't going to have as great a day.
Now, does switching from WhatsApp to some other not-very-widely-used platform cause customer engagement / retention to drop? I would wager very much so! It's a matter of priorities - people go where there is least friction, and WhatsApp otherwise provides a seamless friction-less experience.
What might be a good thing for society in the first world doesn't mean it's necessarily good thing for society in the third world.
Facebook is the most user-hostile tech megacorp, and they will inevitably harm these businesses you care about. The sooner the bandaid is ripped off the better.
You want to dislodge Facebook, you need to disrupt it / curtail its monopoly.
With your username, I think you can risk naming the country without any additional loss of privacy.
I was amazed to discover how pervasive Facebook, Inc. has become in the developing world for conducting business and navigating everyday life.
For a lot of people in developing nations such as the Phillipines and Indonesia, Facebook is synonymous with the internet. This has been buoyed by their push to bundle uncapped/free data for Facebook with mobile plans in markets with high growth of mobile internet access.
It's interesting, because I'm always reading articles about how "Western teens aren't using Facebook any more", which is true, but it's also irrelevant, because they're not really a profitable market, teenagers have short attention spans and no money. Facebook's growth strategy is to become the one stop shop (in lower income nations) for everything you want and need.
They create Instagram accounts and post products as posts, with a caption of "DM me for price".
It also turns on every alarm on my mind, when they start calling these "Instagram pages". It blurs the line between a real website and an Instagram account (In Spanish, "website" is "página web" as well).
I've also heard: "My business went to hell because Instagram killed my account" and that's when I reply: "Have you ever thought of owning a real website?"
I think if its staying down for a few more days Canibalism will ensue by the end of the week.
That's a dick move by the neighborhood.
Yes how that content is presented, ranked, etc is controlled by Facebook but that contribution is less than the content itself.
It would be better to say it's the spoon in which someone could eat a sugary cereal or something healthy.
My extended family uses FB to share info about events.
This, and other pedantic activities are really common around the world.
Don't reduce the material reality a situation to a meme that that represents a personalized view.
My family didn't share online before FB.
My mother didn't really have a common means to communicate with her grandchildren in the same way.
Email, phone are just not the same.
There are more channels available now for sure, but none so ubiquitous.
Facetime is not displacing FB for a lot of things, but that's more direct.
'Everyone is on FB' is the reason it still holds in these kinds of uses cases.
None of us case one way or the other about the platform, we'll just use what's convenient, but that is what it is.
This is a very common theme among FB users. FB by the way, is still growing it's userbase, and growing revenues even more so. The themes we see here on HN and even in the news don't represent the views among the population, nor are they necessarily very close to material reality.
The problem is initiative and knowledge. They should walk or ride a couple of miles and buy the biggest bags of rice and beans they can, along with a bottle of multivitamins. And then learn how to cook.
If that’s classist, then the classes are structured by knowledge and choices. Which they may well be.
> If that’s classist, then the classes are structured by knowledge and choices. Which they may well be.
class by its definition accounts for massive difference in access to resources. If you think access to resources doesn't measurably change the level of knowledge that a population has, that's a declaration that resources do nothing, which would be an odd stance to take on a knowledge-focused community website.
> They should walk or ride a couple of miles and buy the biggest bags of rice and beans they can, along with a bottle of multivitamins.
I just LOVE the subtle food choice of rice and beans here, paired with the recommendation to take multivitamins, a recommendation that is supported by little to no evidence. Your own lack of knowledge on this topic is in full display, as is a clear demonstration of your own biases across multiple dimensions.
I agree that class accounts for a massive difference in access to resources. However, in this case, the knowledge is available for free, and in the US the basic foodstuffs are available for far less than what disadvantaged people pay for the typical processed and fast food they live on.
Rice and beans - nothing subtle about it. They are basic foods that provide the necessary carbs, fat, and complete protein. The vitamins are a simple way to prevent scurvy and similar deficiencies, until the choice of food can become more varied.
As a person learns to cook and bake, they can add wheat, peas, and corn (But they need to learn about nixtamalization before they add corn.) None of these foods require refrigeration.
I have in mind the cuisine of Mexico, which is inexpensive and nutritious. Similar cuisines are found in home cooking all over the world, at least where commercially processed food hasn't driven them out.
It is most important to make sure that all school children are taught how to process and cook these basic foods.
If you are knowledgeable in this area, I'd appreciate some specific suggestions.
Alternatives beyond signal that normies can use: Email.
Spread the word!
I think things would be better if more people had their own domain. I just don't see any way of making it happen. I can't get my own family to leave gmail even with me handling all the domain stuff for them. Even my technical coworkers who are capable of this don't care.
How about choosing something that's federated? https://matrix.org/
(I'm not kidding)
More guidance required.
FWIW it seems possible that the messages remain cached locally on your device but deleted from their servers, and with their outage they aren’t being updated to delete?
Think of it - half the country doesn't have internet because of this crash, that's terrifying. (Switching DNS servers obviously works but that's not something the general population will do)
Anyway, I didn't even notice since I run knot-resolver at home.
I wonder what it will be like connecting Facebook back to the internet, thundering herd and everything...
Will not that be an issue? Re-enabling routing to such a massive internet service...
Correction: according to Cloudflare's blog post, some do cache errors as well.
Company like Facebook has a serious problem and their stock drops ... precipitously. CEO of said company instead of selling their equity in their company has taken out loans against their equity in order to decrease their tax burden and cash in on the value of their equity.
What amount of decrease would cause a margin call from lenders for the forced sale of said equity and subsequently the loss of majority stake in their own company? Now obviously only the lenders know this information and assuming I have the rough order of operations correct.
Could this be a potential chink in the armor of founders / CEOs / anyone who takes out low interest loans against the equity they hold in their company? Maybe my understanding of this is too simplified.
This scenario would just be basically impossible.
Separately, my bank tried to sell me a "Pledged Asset Line of Credit" that would also have required the collateral to maintain a certain value or there would be forced liquidations.
Can you link an example of a bank or similar that lets you borrow against stock or options collateral without margin call risk?
The original post you were replying to talked about founders borrowing against their company's shares as individuals, not companies.
> it is unlikely that a lending institution will accept shares as collateral due to the wrong way risk
That's my point: Nobody's getting a special financing deal on their company stock as individuals to eliminate their margin call risk.
Multi-billion dollar hedge funds get a personal contact at the bank, but even they will get margin-called borrowing against stocks as collateral if it goes against them.
It is doubtful Z has any margin call issues as he has so much stock, I can't imagine he would have pledged even 5% of it for loans, so he can just hand them another chunk without even blinking (which he generally doesn't do any way)
You raise an interesting question though and I'd like to know the answer as well!
Even if he didn't, the bank would let him move funds in without forcing him to sell.
For starters Zuck owns Class B shares which have 10x the votes of Class A shares. I'm sure a bank would happily loan him money against his Class B shares, but any forced liquidation would involve a transfer to Class A shares. Zuck could lose a lot of shares and still maintain control of the company.
[1] https://www.wsj.com/articles/a-board-struggles-with-its-ceos...
Fascinating simply that it's apparently not just a DNS issue.
It's so important to diversify, such as building a website.
We've come full circle, where techies are rediscovering the original hatred for the Oculus, that it is tied to a social media walled garden, for some reason.
When FB announced they would be buying Oculus, they promised that no social media integration would be required. FB breaking that promise is not the same as Oculus having that requirement from the get-go.
What original hatred are you talking about??
This ^ means fuck-all, because at that time (day 1), their oculus services where hosted in the same infrastructure as their social media services.
Last year, they got rid of "you do not need a facebook account". But in all situations since inception, all of your data is passing through the same infrastructure as facebook data. It may not be being exposed, or targeted for advertising, but this WAS a huge point of contention years back.
with my second-hand knowledge from someone who worked for FB and assisted the Oculus team being folded in to FB processes/policies/tech, I don't think this is accurate, either.
Remember the 'information super-highway'? Yeah it gets carpet bombed constantly....
Then it hit me: I am so dependent on Facebook owned properties (Whatsapp, Facebook, insta) that a Facebook failure looks to me like an internet failure.
Seems like another poster posted finer details regarding BGP/peering which is ultimately causing the issue.
and / or an S3 bucket with a json blob the apps can pull to at least tell users 'here's what's up'
> This must be incredibly stressful so for your sake I hope you sort it out quickly... but for the world's sake, I hope you fail and make the problem worse before jumping ship followed by every other engineer, leaving it to Zuckerberg to fix himself. But I still hope it's not too stressful for you!
https://nitter.mailstation.de/signalapp/status/1445062426739...
> Facebook said: "We are aware some people are having trouble accessing our apps and products. We are working to get things back to normal as quickly as possible and apologise for any inconvenience."
What's wrong with the internet?
FaceBook is down.
My friend from Slovenia is having trouble with discord. It eats his messages.
I can't load photos from my friend in telegram and the messages take a relatively long time - multiple seconds! - to get received.
TrackMania players have talked about having input lag.
ycombinator is really slow and reports an error after submitting. "We're having some trouble serving your request. Sorry!" (lost count of the times i've tried submitting this)
ycombinator turned out to be giving only errors.
Some sites I've found via google results seem to report that they are suffering from slow connections.
Do you have anything to add to this?
How could facebook dns issues cause this?
"Because of missing DNS records for http://Facebook.com, every device with FB app is now DDoSing recursive DNS resolvers. And it may cause overloading ..."
None of the listed facebook nameservers are resolvable or reachable:
a.ns.facebook.com b.ns.facebook.com c.ns.facebook.com d.ns.facebook.com
mtr -r -c10 -w -b a.ns.facebook.com Start: 2021-10-04T10:02:50-0600 Loss% Snt Last Avg Best Wrst StDev
...
4.|-- ae-2-rur101.cosprings.co.denver.comcast.net (162.151.51.125) 0.0% 10 12.6 11.9 9.6 19.0 2.9
5.|-- 24.124.155.233 0.0% 10 9.3 10.2 9.1 12.4 1.1
6.|-- 96.216.22.45 0.0% 10 12.0 14.0 11.6 31.3 6.1
7.|-- be-36041-cs04.1601milehigh.co.ibone.comcast.net (96.110.43.253) 20.0% 10 14.6 13.5 11.6 20.5 3.0
8.|-- be-3402-pe02.910fifteenth.co.ibone.comcast.net (96.110.38.126) 0.0% 10 12.2 12.0 11.5 13.2 0.5
9.|-- 173.167.59.170 0.0% 10 13.8 17.8 12.0 34.7 8.4
10.|-- 129.134.40.74 0.0% 10 15.3 12.6 11.4 15.3 1.1
11.|-- 129.134.43.226 0.0% 10 18.9 15.3 12.6 20.3 3.0
12.|-- 129.134.98.166 0.0% 10 12.5 14.2 12.5 20.4 2.3
13.|-- 129.134.54.61 0.0% 10 34.2 30.8 28.9 34.2 1.8
14.|-- 129.134.53.61 0.0% 10 29.8 31.1 28.9 36.5 2.7
15.|-- 129.134.53.61 90.0% 10 31.9 31.9 31.9 31.9 0.0mtr -r -c10 -n b.ns.facebook.com
Start: 2021-10-04T10:28:03-0600 Loss% Snt Last Avg Best Wrst StDev
1.|-- 192.168.1.1 0.0% 10 0.2 0.2 0.2 0.3 0.0
2.|-- 96.120.12.229 0.0% 10 10.2 10.8 8.8 15.7 1.9
3.|-- 96.110.149.185 0.0% 10 17.7 13.6 9.8 32.3 7.0
4.|-- 162.151.51.125 0.0% 10 10.9 12.2 9.6 15.3 1.9
5.|-- 24.124.155.233 0.0% 10 13.0 10.4 9.4 13.0 1.2
6.|-- 96.216.22.45 0.0% 10 16.5 16.7 11.2 29.1 6.4
7.|-- 96.110.43.241 0.0% 10 17.4 13.6 11.9 17.4 1.6
8.|-- 96.110.38.114 0.0% 10 12.5 12.8 12.0 14.0 0.6
9.|-- 173.167.59.170 0.0% 10 36.1 19.3 11.6 36.1 9.7
10.|-- 129.134.40.76 0.0% 10 13.1 12.3 11.3 13.1 0.6
11.|-- 129.134.34.72 0.0% 10 15.3 15.7 13.5 21.3 2.5
12.|-- 129.134.102.85 0.0% 10 39.0 39.2 38.0 40.8 1.0
13.|-- 31.13.25.13 0.0% 10 30.5 29.8 28.5 31.0 0.9
14.|-- ??? 100.0 10 0.0 0.0 0.0 0.0 0.0
15.|-- ??? 100.0 10 0.0 0.0 0.0 0.0 0.0
16.|-- 31.13.25.13 90.0% 10 30.2 30.2 30.2 30.2 0.0
On a side note... why is it so freaking hard to format line breaks in HN?"We're noticing an elevated level of usage for the time of day and are currently monitoring the performance of our systems. We do not anticipate this resulting in any impact to the service.
We have temporarily disabled typing notifications. We expect these to be re-enabled soon."
But to be fair... seems like it was a good call to not do it Friday night :D
[root@app ~]# nslookup downforeveryoneorjustme.com 4.2.2.2 ;; connection timed out; trying next origin ;; connection timed out; no servers could be reached
[root@app ~]# nslookup downforeveryoneorjustme.com 1.1.1.1 Server: 1.1.1.1 Address: 1.1.1.1#53
Non-authoritative answer: Name: downforeveryoneorjustme.com Address: 172.67.166.187 Name: downforeveryoneorjustme.com Address: 104.21.91.48
[root@app ~]#
Perhaps DNS queries are skyrocketing and overwhelming some of the major public DNS servers.
>Now, here's the fun part. @Cloudflare runs a free DNS resolver, 1.1.1.1, and lots of people use it. So Facebook etc. are down... guess what happens? People keep retrying. Software keeps retrying. We get hit by a massive flood of DNS traffic asking for http://facebook.com
https://twitter.com/jgrahamc/status/1445066136547217413
>Our small non profit also sees a huge spike in DNS traffic. It’s really insane.
https://twitter.com/awlnx/status/1445072441886265355
>This is frontend DNS stats from one of the smaller ISPs I operate. DNS traffic has almost doubled.
https://twitter.com/TheodoreBaschak/status/14450732299707637...
edit: I see it's back up and I've been getting downvoted, here's a screenshot of the error for clarity
If this happened then some wider issue (github down) there’d be chaos.
I'm not suggesting that this is the case, but a failure of this scale (with internal systems also down) could allow scrubbing of evidence without leaving traces.
A Facebook whistleblower, who is due to testify before Congress on Tuesday, has accused the Big Tech company of repeatedly putting profit before doing “what was good for the public,” including clamping down on hate speech.
Frances Haugen, who told CBS’s “60 Minutes” program that she was recruited by Facebook as a product manager on the civic misinformation team in 2019, said she and her attorneys have filed at least eight complaints with the U.S. Securities and Exchange Commission.
During her appearance on the television program on Sunday, Haugen revealed that she was the whistleblower who provided the internal documents for a Sept. 14 exposé by The Wall Street Journal that claims Instagram has a “toxic” impact on the self-esteem of young girls.
That investigation claimed that the social media giant knows about the issue but “made minimal efforts to address these issues and plays them down in public.”
“The thing I saw at Facebook over and over again was there were conflicts of interest between what was good for the public and what was good for Facebook. And Facebook, over and over again, chose to optimize for its own interests, like making more money,” said Haugen.
She explained that Facebook did so by “picking out” content that “gets engagement or reaction,” even it that content is hateful, divisive, or polarizing, because “it’s easier to inspire people to anger than it is to other emotions.”
“Facebook has realized that if they change the algorithm to be safer, people will spend less time on the site, they’ll click on less ads, they’ll make less money,” she claimed.
Haugen is expected to to testify at a Senate hearing on Oct. 5 titled “Protecting Kids Online,” about Facebook’s knowledge regarding the photo sharing app’s allegedly harmful effects on children.
During her appearance on the television program, Haugen also accused Facebook of lying to the public about the progress it made to rein in hate speech on the social media platform. She further accused the company of fueling division and violence in the United States and worldwide.
“When we live in an information environment that is full of angry, hateful, polarizing content it erodes our civic trust, it erodes our faith in each other, it erodes our ability to want to care for each other. The version of Facebook that exists today is tearing our societies apart and causing ethnic violence around the world,” she said.
She added that Facebook was used to help organize the breach of the U.S. Capitol building on Jan. 6, after the company switched off its safety systems following the U.S. presidential elections.
While she believed no one at Facebook was “malevolent,” she said the company had misaligned incentives.
“Facebook makes more money when you consume more content,” she said. “People enjoy engaging with things that elicit an emotional reaction. And the more anger that they get exposed to, the more they interact and the more they consume.”
Shortly after the televised interview, Facebook spokesperson Lena Pietsch released a statement pushing back against Haugen’s claims.
“We continue to make significant improvements to tackle the spread of misinformation and harmful content,” said Pietsch. “To suggest we encourage bad content and do nothing is just not true.”
Separately, Facebook Vice President of global affairs Nick Clegg told CNN before the interview aired that it was “ludicrous” to assert social media was to blame for the the events that unfolded on Jan. 6.
The Epoch Times has reached out to Facebook for additional comment.
https://www.theepochtimes.com/facebook-whistleblower-claims-...
https://en.wikipedia.org/wiki/The_Epoch_Times
They also previously ran a large sockpuppet network on Facebook and the Facebook ad platform (both of which have since been banned) so they may have a bit of a bone to pick with the platform.
> mobile phones are addictive
> internet is used to organize protests
EDIT: Yes, Signal is not federated, but that's what people are at least ready to consider as a WhatsApp alternative. I also created Matrix / Element account, and had 0 contacts using it already.
https://old.reddit.com/r/thehatedone/comments/f160jh/is_sign...
Signal is still centralized and uses AWS. So if AWS was to go down, it would affect not just Signal but vast swathes of the Internets.
As it is now, Matrix may offer better privacy and more robust weil federated and p2p, but if I had to personally ask all of my contacts if they actually use it using some other medium, I can keep using that other medium too for conversations.
"A small team of employees was soon dispatched to Facebook’s Santa Clara, Calif., data center to try a “manual reset” of the company’s servers, according to an internal memo."
but https://glean.software and https://reactjs.org aren't
> Is it a DNS issue? -> yes
It can be used in reverse as a postmortem too.
- https://en.wikipedia.org/wiki/Facebook_onion_address
- facebookwkhpilnemxj7asaniu7vnjjbiltxjqhye3mhbshg7kx5tfyd.onion
Zuckerberg is taking his ball and going home unless you stop writing mean things about him /s
I really hope it's just some internal technical error and not a "see, despite the bad things of FB, you really need us" move.
It's probably trivial, the timing just seems weird to me.
I also wouldn't classify the loss of 1 company, and the expiration of some TLS certificates, as the interconnected network of networks being broken. The Internet has continued to function even if some larger players were unreachable or having issues.
Unfortunately, we have literally dozens of comments that amount to nothing more than schadenfreude, and another handful of non-FANG employees speculating how one of the largest internet operations in existence could improve their game (lol)
What's wrong with the internet?
FaceBook is down.
My friend from Slovenia is having trouble with discord. It eats his messages.
I can't load photos from my friend in telegram and the messages take a relatively long time - multiple seconds! - to get received.
TrackMania players have talked about having input lag.
ycombinator is really slow and reports an error after submitting. "We're having some trouble serving your request. Sorry!" (lost count of the times i've tried submitting this)
ycombinator turned out to be giving only errors, but now seems to be working occasionally. I can not submit anything, though.
Some sites I've found via google results seem to report that they are suffering from slow connections.
Do you have anything to add to this?
With Facebook down, some large DNS servers seem to be struggling with the extra load of failing requests to look up "facebook.com". Cloudflare reports overload with their DNS server at 1.1.1.1, although that's working for me.
Billions of things worldwide are trying to connect to Facebook. The lookup which normally returns the IP address for facebook.com on the first try now requires trying a.ns.facebook.com, b.ns.facebook.com, etc. several times each before giving up. Probably several times a minute for everyone who has a Facebook app in their phone turned on. That may be using a big fraction of world DNS resources.
Vodaphone Ireland seems to be struggling with a DNS overload right now, per the Irish Independent. Also, their status page can't find "Dublin" as a city.
* server can't find www.facebook.com: SERVFAIL
dig +trace messenger.com
shows that all is well with the root DNS servers and dig @a.ns.facebook.com messenger.com
;; connection timed out; no servers could be reached
and also ping a.ns.facebook.com
3 packets transmitted, 0 packets received, 100.0% packet loss
shows that something's wrong with facebook.Tracing route to a.ns.facebook.com [129.134.30.12] over a maximum of 30 hops:
1 1 ms 1 ms 1 ms eehub.home [192.168.1.254]
2 3 ms 3 ms 3 ms 172.16.14.63
3 \* 5 ms 3 ms 213.121.98.145
4 5 ms 3 ms 4 ms 213.121.98.144
5 17 ms 8 ms 18 ms 87.237.20.142
6 8 ms 6 ms 7 ms lag-107.ear3.London2.Level3.net [212.187.166.149]
7 \* \* \* Request timed out.
8 \* \* \* Request timed out.
9 7 ms 7 ms 6 ms be2871.ccr42.lon13.atlas.cogentco.com [154.54.58.185]
10 70 ms 69 ms 70 ms be2101.ccr32.bos01.atlas.cogentco.com [154.54.82.38]
11 73 ms 73 ms 74 ms be3600.ccr22.alb02.atlas.cogentco.com [154.54.0.221]
12 84 ms 85 ms 84 ms be2879.ccr22.cle04.atlas.cogentco.com [154.54.29.173]
13 90 ms 90 ms 90 ms be2718.ccr42.ord01.atlas.cogentco.com [154.54.7.129]
14 143 ms 142 ms 143 ms po111.asw02.sjc1.tfbnw.net [173.252.64.102]
15 114 ms 119 ms 114 ms be3036.ccr22.den01.atlas.cogentco.com [154.54.31.89]
16 125 ms 126 ms 124 ms be3038.ccr32.slc01.atlas.cogentco.com [154.54.42.97]
17 91 ms 92 ms 91 ms po734.psw03.ord2.tfbnw.net [129.134.35.143]
18 91 ms 93 ms 90 ms 157.240.36.97
19 74 ms 74 ms 73 ms a.ns.facebook.com [129.134.30.12]
Trace complete.Of course it was, as far as I know, ficitonalized in the first place, although it rings true (in context) to some extent. What I wonder is, how much is that true now? That is, how much downtime would FB have to experience for enough users to start leaving, to the point that it might prompt a serious exodus.
Facebook is just too big and pervasive that such an outage would be treated by its users like an internet outage or a power outage. Once it's back online, everyone will forget.
I need help but this is too hard for me. I uninstall social media once a week but install two days later.
I should probably go to therapy with this, but I am not sure I wouldnt be laughed at
Which is a feature I hate, since it does that all the time even when I have a connection. Says there are 3 comments on a post, when I know there is more. Opening them doesnt show them, and no way to refresh. But going to the web page I can see them.
% ping whatsapp.com
ping: whatsapp.com: Name or service not known
% ping web.whatsapp.com
ping: web.whatsapp.com: Name or service not known
% ping facebook.com
ping: facebook.com: Name or service not known
% ping instagram.com
PING instagram.com (31.13.65.174) 56(84) bytes of data.
64 bytes from 31.13.65.174 (31.13.65.174): icmp_seq=1 ttl=53 time=110 msHow could they all go down at the same time, if they have different teams of engineers running each product separately?
Could anyone with some background (or person familiar with the matter) explain how their system's set up?
edit: though saying that, they do run their own registrar... Might've fucked something up over there.
FWIW, WhatsApp (on phones) should be resiliant to a DNS only outage, the clients contain fallback IPs to use when DNS doesn't work, and internal services don't use DNS as far as I remember.
At one time, WhatsApp had actually separate infrastructure at SoftLayer (IBM Cloud now), but that hasn't been in place for quite some time now. When I left, it was mostly just HAProxy to catch older clients with SoftLayer IPs as their DNS fallback.
When uBlock Origin is running, this script gets blocked and pages return to feeling snappy.
Of course there is the whistle-blower issue too...
Looks like there are a few problems with fb in the news today ...
...or many HN users are also avid FB users (and now have to resort to backup sources of entertainment)
Check out Dug! Its a global DNS propagation/monitoring toolon the CLI: https://github.com/unfrl/dug/
Now Twitter is starting to have problems with overload.
I’d guess internal networking issues, but the insane that something can bring down all of Facebooks properties.
It's crazy that half the country doesn't have internet because Facebook stopped working.
Tracing route to a.ns.facebook.com [129.134.30.12] over a maximum of 30 hops:
1 1 ms 1 ms 1 ms eehub.home [192.168.1.254]
2 3 ms 3 ms 3 ms 172.16.14.63
3 * 5 ms 3 ms 213.121.98.145
4 5 ms 3 ms 4 ms 213.121.98.144
5 17 ms 8 ms 18 ms 87.237.20.142
6 8 ms 6 ms 7 ms lag-107.ear3.London2.Level3.net [212.187.166.149]
7 * * * Request timed out.
8 * * * Request timed out.
9 7 ms 7 ms 6 ms be2871.ccr42.lon13.atlas.cogentco.com [154.54.58.185]
10 70 ms 69 ms 70 ms be2101.ccr32.bos01.atlas.cogentco.com [154.54.82.38]
11 73 ms 73 ms 74 ms be3600.ccr22.alb02.atlas.cogentco.com [154.54.0.221]
12 84 ms 85 ms 84 ms be2879.ccr22.cle04.atlas.cogentco.com [154.54.29.173]
13 90 ms 90 ms 90 ms be2718.ccr42.ord01.atlas.cogentco.com [154.54.7.129]
14 143 ms 142 ms 143 ms po111.asw02.sjc1.tfbnw.net [173.252.64.102]
15 114 ms 119 ms 114 ms be3036.ccr22.den01.atlas.cogentco.com [154.54.31.89]
16 125 ms 126 ms 124 ms be3038.ccr32.slc01.atlas.cogentco.com [154.54.42.97]
17 91 ms 92 ms 91 ms po734.psw03.ord2.tfbnw.net [129.134.35.143]
18 91 ms 93 ms 90 ms 157.240.36.97
19 74 ms 74 ms 73 ms a.ns.facebook.com [129.134.30.12]
Trace complete.this is what i got now
https://www.youtube.com/watch?v=X6NJkWbM1xk
The video is one that Tom Scott published in June 2020 about the worst typo he ever made in one of his prior jobs, and while the Facebook mistake is almost certainly not going to be anything irrecoverable like this one, you can bet that Facebook pride themselves on being available all the time.
i feel for their netops people. uncharted territory with the whole world watching and, no doubt, a lot of morons from management trying to be "helpful" in getting this nice crisis resolved. for any crisis there is always a bunch of clowns with MBAs that consider it their golden opportunity to shine (nearly always at someone elses expense)
Any other scenario of outsiders, code updates, etc - basically misses the point of how modern DNS infrastructure works.
dig NS example.com
; ANSWER
example.com. NS ns1.example.com.
; ADDITIONAL
ns1.example.com. A 1.2.3.4
Edit: I didn’t see your glue comment when I wrote this.Thought the common wisdom nowadays was to use nameservers on different TLDs and sub-labels for the best resilience.
/added, they seem to have glue records so I'd assume it's the nameservers themselves having issues.
$ dig NS @g.gtld-servers.net. a.ns.facebook.com
;; AUTHORITY SECTION:
facebook.com. 172800 IN NS a.ns.facebook.com.
facebook.com. 172800 IN NS b.ns.facebook.com.
facebook.com. 172800 IN NS
c.ns.facebook.com.
facebook.com. 172800 IN NS d.ns.facebook.com.
;; ADDITIONAL SECTION:
a.ns.facebook.com. 172800 IN A 129.134.30.12
a.ns.facebook.com. 172800 IN AAAA 2a03:2880:f0fc:c:face:b00c:0:35
b.ns.facebook.com. 172800 IN A 129.134.31.12
b.ns.facebook.com. 172800 IN AAAA 2a03:2880:f0fd:c:face:b00c:0:35
c.ns.facebook.com. 172800 IN A 185.89.218.12
c.ns.facebook.com. 172800 IN AAAA 2a03:2880:f1fc:c:face:b00c:0:35
d.ns.facebook.com. 172800 IN A 185.89.219.12
d.ns.facebook.com. 172800 IN AAAA 2a03:2880:f1fd:c:face:b00c:0:35
[1] https://www.infoworld.com/article/2648947/youtube-outage-und...
All of the static hosts providing free SSL: vercel, netlify, render, firebase hosting, github pages, heroku etc. ...
It does work on modern browsers and devices but goes terribly broken on a lot of old devices.
https://crt.sh/?q=facebook.com
On a side note, the amount of phishing sites using letsencrypt and having a domain similar to facebook.com is quite appalling.
Edit: the counter just jumped from 10B to 60M, so I doubt it's any reliable :)
I only know of a handful of Saas apps they didn’t build internally. Sadly none of those will help them get out of this situation.
Could this be a coordinated smear in HN comments?
This will be a highly discussed topic for a bit.
For example, will FB addicts experience a day of repeated failed attempts to get their FB fix, which will then condition them to stop trying.
I didn't receive expected WhatsApp messages and am only now realizing there's no indication within the app that there is even a problem. It only becomes (somewhat) apparent when sending a message never gets a single check mark. Not a graceful failure for the user view.
Is the article link just https://facebook.com/ ?
The link was to show that facebook was down.
https://tcrn.ch/3kOHco1 2Africa cable, as an example
https://soundcloud.com/ryan-flowers-916961339/dns-to-the-tun...
Events like this show they should use multiple outlets instead of the big monopoly.
Alternatives like gab exist, but its incredibly hard to gain traction against the big monopolies.
Sometimes things just break and take time to fix.
Vote buttons are not a substitute for proper responses to legitimate enquiry.
I kid. If it were to come down to a single person, that's really a failure of the whole organization system and not of the individual.
This apocryphal [1] punchline to the Jack Welch story also sums up how most orgs deal with this sort of thing:
"I just spent a million dollars on your education - why would I fire you now?"
[1]: http://www.nickmilton.com/2016/03/jack-welch-on-learning-fro...
[I'm also getting server error trying to submit this comment]
Yes, I know humor is not welcome in HN
Instagram and WhatsApp - not yet.
But, a good rule of thumb right now is about $10,000,000 per hour.
Of course, depends on the hour of day. Facebook likely makes more ad money when North America is awake than when Asia is awake for instance.
> ping facebook.com
PING facebook.com (31.13.83.36) 56(84) bytes of data.
64 bytes from edge-star-mini-shv-01-mad1.facebook.com (31.13.83.36): icmp_seq=1 ttl=54 time=12.2 ms
64 bytes from edge-star-mini-shv-01-mad1.facebook.com (31.13.83.36): icmp_seq=2 ttl=54 time=12.1 ms
64 bytes from edge-star-mini-shv-01-mad1.facebook.com (31.13.83.36): icmp_seq=3 ttl=54 time=11.7 ms
Can't yet traceroute to a.ns.facebook.com tho
https://twitter.com/blazejkrajnak/status/1445063232486531099
Maybe it's DNS. It's always DNS.
If it is, close Facebook as there's probably a BGP hijack going on that is siphoning off personal data and or secrets
All this BGP talk is boring.
It's not just lag, I keep getting the "We're having some trouble serving your request. Sorry!" page.
Edit: HN related thread https://news.ycombinator.com/item?id=28749476
; <<>> DiG 9.10.6 <<>> facebook.com ;; global options: +cmd ;; Got answer: ;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL, id: 36072 ;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 1
It may be very hard for employees to get to the physical boxes, and/or bypass any physical or software security systems.
Personally I'm glad FB went down for a few hours, but it's hard to imagine how that would happen in the first place.
And the world rejoiced.
Cutting off Facebook's firehouse of hate and misinformation for just a couple hours is going to have a obvious positive effect on millions of people. At this scale, at least one person will get vaccinated today because they didn't see the wall of ignorance that is FB's news feed.
Maybe we should introduce "digital blue laws", where one day a week, social media is shut down for the overall good of society.
Those services are toxic for years now and everybody knows that. Who still uses them occasionally let alone relies on them can't be helped, can they?
People will do this wherever they can talk in a group online, not just Facebook properties. It's... pretty bad actually, I think the only tool that exists right now is censorship, because the bullshit gets created, spread, and wholeheartedly received way faster than debunking will.
And censorship is a power that can't be safely entrusted to nobody.
The reasons I've seen are:
> it creates a risk of bad self-image for young girls
It's a parent's job to educate your children. There are much worse things than Facebook out there.
> it collects data
Literally no harm in knowing that someone is interested in JavaScript, cats and fetish porn, and targeting ads to that user.
> it's addictive
So is sex, marijuana, and collecting stamps.
> it helps organize protests
Good.
https://www.businessinsider.com/facebook-pushes-qanon-racism...
I'd argue that it's absolutely in Facebook's self-interest to reduce their active role in promoting fascism, racism, homophobia, etc.
That's the whole point. Oh they're just trying to make a buck like everyone else is exactly the problem.
They are a running a paperclip maximizer that turns passive consumers of misinformation into "engaged" radicals and the system that is Facebook has no incentive to correct this.
Two questions:
- What do you think should be done about the legacy media that is doing the same?
- Should social media promote boring posts, or actively censor political content in favour of a certain viewpoint, or anything else? Perhaps a real-life name registration for anyone with over 1000 followers, like in China?
Incorrect assertion. Those posts promote hatred and/or violence toward humans for traits those humans did not choose. e.g. race, sexual orientation, etc.
Legacy media aren't actively amplifying the voices and recruiting efforts of white supremacists.
Facebook is. They acknowledge that they are. They chose to actively allow and encourage it for profit.
I'm guessing that either you're not a parent, or your kids aren't teens.
But most parents of teens realize that kids, and especially teens, are often much more influenced by things like social media & peers (and peers via social media) vs. influence their parents have on them.
I doubt, if FB goes away, that any of the issues you're implying will go away or even get much better. In fact, the lack of a real look into the negative effects of consumer news product reinforces this idea that only the elite can know the truth, and the masses just have to get in line and shut up.
News media proliferated nonsense from fed sources to justify the Iraq war, they gave Trump 24/7 airtime for a while because it increased ratings. They constantly forgo any real accountability for their actions, and pretend that they aren't just another addictive consumer product that warps peoples' brains.
Direct quote: "That website on Facebook."
There are people who believe that "Facebook" literally equals "Internet". Facebook, Internet ... Internet, Facebook.
Rinse and repeat for your alternative echo chamber regarding Google, the Microsoft Bing, &c.
Does WhatsApp have a true E2E either? Ask hundreds of moderators employed by Facebook who review WhatsApp messages flagged as improper and the chat history around them...
However, accepting the fact that neither of the services is truly secure, Telegram experience as a service is much better for an average user.
That was my problem, and your confirmation means it's still as good as nothing.
> Does WhatsApp have a true E2E either? Ask hundreds of moderators employed by Facebook who review WhatsApp messages flagged as improper and the chat history around them...
If one of the ends decides to share a message, it's still E2E. That is the big difference.
True. But you can't prove that "one of the ends" must necessarily be a human and not the logic in the app code, or an intended backdoor? E.g., an automated logic scanning for 'malicious' messages on-device.
Changing messaging apps not the most convenient thing in the world, but it's not some kind of IT cataclysm. Plenty of WhatsApp competitors exist.
There is a gap between "I want to know what people I know are up to" and "I want to meet with those people one by one to see what they are up to". Some people just want to passively watch and that is ok.
And this is the culprit for loneliness.
Occasionally, I'll see/hear/do something and think that it would have made a good status update/tweet, but then I remember that these things have happened to me for decade before social media was a thing and life was fine. Some I'll share with my wife or a friend, most just disappear and that's fine too.
Why did you put up with racist and misogynistic people on your feed? Why did you feel the need to delete your account instead of unfollowing people?
My feed is nice and clean, with family, some friends, and some pages.
Obvious I'm not serious, and it's popular sentiment here that "Fuck Facebook... Oh but I use Instragram and WhatsApp of course!", but the point was "some people making a living on x" isn't really a great argument for "x is harmful and we might be better without it".
op: "I don't care about thing X disappearing"
re: "While you may not care about it because of Y, X also provides benefit Z to other people"
op: "But would there be any drawbacks?"
yeah, there would be drawbacks, other people would lose Z, which may matter a lot of them even if it doesn't matter to you. Someone just told you about Z, and you just responded as if you weren't just told about Z"These days I find it incredibly frustrating to deal with people who have conclusively decided they don't like something and that renders them incapable of acknowledging other benefits that said thing provides even if those benefits aren't relevant to them or are less relevant than the things they vocalize caring about.
I'd like to explicitly note that the parent post did not say "X also provides benefit Z to other people" - it asserted "Facebook is an unparalleled titan in the realm of advertising" which is a substantially different thing; it's not something that some people simply don't care about and a benefit to some other people and considering those statements as equivalent is a (very large) intentional blindspot. The current way of how advertising is done (driven, in part, by FB) is also a harm to many people and society at large, so publicly making an implicit assumption that "advertising" is at most neutral is not okay, it's something that should be called out.
This very "unparalleled titan in the realm of advertising" aspect is a major cost on society, a net harm that perhaps should be tolerated if it's outweighed by some other benefits FB provides (such as the "utility-level communication system for a big chunk of the globe"), but as itself it's definitely not something that should be treated as benign just because some people get paid for it.
If FB advertising disappeared with no other drawbacks, that would be a great thing. Of course, there are some actual drawbacks, but even so it's quite reasonable to motivate people to ask about the actual drawbacks of FB being down, because "oh but ads" (with which the grandparent post started) is not one.
For the better.
The world would only improve if it disappeared.
Uh, Google? It's definitely paralleled, and also preceded
"I felt a great disturbance in the DNS. As if millions of influencers suddenly cried out in terror and were suddenly silenced. I feel something terrible has happened. But you better get on with your content curation."
There I fixed it for you.
Not unparalleled - Google exists.
And we need less advertising, not more.
> and WhatsApp is basically a utility-level communication system for a big chunk of the globe.
Many other such systems exist - Telegram, Signal, Google Chat.
> Instagram is a key cultural driver of the Western world
Western culture will get along just fine without Instagram.
> the overall impact would transform the world around you.
For the better.
Messaging: people have been switching on hordes to every new free messaging system in the 90s and early 2000s, we will adapt to something else.
Netflix and video in general: same thing without the 90s/early 2000s.
Amazon: very convenient store, we'll spend a little less and somebody will fill the void.
Apple: can't say, never bought anything from them.
By the way, when I couldn't message on WA today I thought day they finally cut me off because I still didn't accept their new privacy policy from months ago :-) I resolved to wait and see for a couple of days.
Back then the IM population was a lot smaller. Also with "Free Basics" and other things in some regions of the world Facebook plays a game which makes it impossible to switch. (Using Whatsapp is free, for others one ahs to buy mobile data credits)
The group chat functions don't really exist in SMS (maybe in MMS but they never work properly), photos (same), whatsapp desktop, you can text when you have WiFi but no 4G (or using a different sim card when travelling), etc.
> "Ok." - Zuckerburg
shuts down $100B company
I hate the stupid strategy tax that makes me have an FB account to use their headset, and has it go down when they have an outage. I hope they can learn from MSFT that "Facebook Everywhere" is ultimately a self defeating strategy.
Zuck trying to give an example of what a world without FB would look like, kinda saying to detractors what would happen if they had it their way.
Anyway...
Mission accomplished, I'd say. For now at least.
I'm going to parrot the other comment here and say nothing of value was lost.
> ...there is no limit to the scandals, leaks, whistleblowers, lawsuits or penalties that will bring the Facebook mafia down.
Fine. 'Literally' bringing the Facebook mafia down like that would do.
But only for now.