- YouTube
- Google Play
- T-Mobile
- X (Twitter)
- Discord
- TikTok
- Pokemon Go
- Snapchat
It looks like they all have the same failure point.
- YouTube
- Google Play
- T-Mobile
- X (Twitter)
- Discord
- TikTok
- Pokemon Go
- Snapchat
It looks like they all have the same failure point.
It's as good as asking a neighbor what happened with a loud noise down the street. Sometimes you'll get something good, sometimes it'll be completely wrong.
> It's as good as asking a neighbor what happened with a loud noise down the street. Sometimes you'll get something good, sometimes it'll be completely wrong.
Asking my neighbors if they know what some loud noise was or about some local disturbance has been extremely reliable in my experience. The one time someone gave me an explanation about something which wasn't mostly right they qualified it with something like "So-and-so said it might be such-and-such but I don't know if it's true".
Car exhaust :: gunshot Appliance delivery truck liftgate :: gunshot Transformer explosion :: gunshot Garbage truck :: gunshot 787 at 25000ft :: complete ruining of peace and quiet Any police activity :: probably someone robbed a bank
For the record, my city has (statistically indistinguishable from 0) homicides and bank robberies and, by American standards (I know, I know) no particular issues with gun crime.
One time I heard a loud boom. A few hours later I saw a neighbor outside and asked if he'd heard it and if he knew what it was. He told me a house a few neighborhoods over had exploded. I was a bit skeptical of it but he turned out to be right.
Sources:
https://www.statista.com/statistics/942043/laboratory-incide... - meth lab incidents are down to about 900/yr and have been far higher in the past (presumably because the labs have moved to things besides meth)
https://rpgaspiping.com/blog/critical-safety-tips/gas-safety... (286 natural gas incidents per year) - I've tried to find a more credible source for this number but keep seeing it cited in various places and have no better source, higher or lower.
I can’t tell if you’re trying to demonstrate the problem you have with your neighbors ;)
It was a gas leak, not a drug lab. The utility failed to fix things in a neighborhood where there’d been reports of leaks for years: https://www.wbur.org/news/2023/08/17/eversource-fine-gas-exp...
I'm confused. Isn't listening for spikes in complaints about outages a great way to detect them? I know for a fact some service companies monitor social media channels for this purpose (among others). I'd be surprised if that wasn't more or less standard practice.
I've checked Down Detector for ISP outages in my area many times now. It's always confirmed them before my ISP did.
Maybe they should have a backup password (if they site allows it, f-ing Spotify doesn’t), but it’s still effectively down for them!
When there's a major ISP outage, people report problems with all the major sites. When Facebook's down, people report problems with any site that has "Login with Facebook" as an option.
It's almost never actually an outage impacting all of FAANG at once.
If users log into your site with Facebook, then the login functionality of your site effectively is down when "Login with Facebook" is down.
From the user's perspective, your subcontractors, including authentication subcontractors, are a problem for you to deal with and never show them. From your perspective, you could have architected your site in a way that logging in doesn't "go down" when Facebook login is down.
If the user chooses "Login with Facebook" over other authentication options available, and they don't want to use other options, educating them with a good error message might help. Or you could remove the Facebook login option, if you (totally reasonably) don't want Facebook's failures to reflect poorly on you.
There are plenty of sites where "Login with Facebook" is a convenience but hardly the only way to log in. Reddit, for example, has "Login with Google" and "Login with Apple"; it would be highly misleading to claim "Reddit is down" if Google's OAuth flow was having an outage.
> educating them with a good error message might help
Nothing in the API or OAuth flow would make that doable in an automatic fashion with this outage. It'd have to be something you put up manually as a banner after hearing of the outage.
> Or you could remove the Facebook login option, if you (totally reasonably) don't want Facebook's failures to reflect poorly on you.
I don't particualrly care; we're talking about why DownDetector isn't necessarily ideal for assessing. It can be a useful signal, in some scenarios, but I've seen plenty of spurious signals come from it.
That is fair: if I choose to architect my site such that a user-critical feature goes down when a 3rd party service goes down, it behooves me to monitor the 3rd party service and do whatever necessary to properly inform users what's going on.
I edited my post unfortunately after you replied, but another option is removing the parts of your site that rely on 3rd parties, if you don't want the failures of those 3rd parties to reflect poorly on you (which they reasonably would).
>we're talking about why DownDetector isn't necessarily ideal for assessing. It can be a useful signal, in some scenarios, but I've seen plenty of spurious signals come from it.
Indeed, and if a bunch of users say that a feature of your site is down, even if it's a result of a 3rd party failure: chances are, that part of your site is down, and it's partially your fault for relying on a 3rd party for that feature. The users correctly don't care what the root cause is, they expect you to either mitigate it or don't have a feature they rely upon be unreliable.
Take a look at https://downdetector.com/status/aws-amazon-web-services/ ; scroll down to the comments.
"SSH and Dbconnect stopped on all of my EC2 instances. Anyone else?"
"I can't add a payment method"
The chart shows a big spike this morning, but there was no AWS outage, nor does Amazon use Facebook login.
Again, DownDetector can be a useful "is something unusual happening right now" signal, but it'd be a mistake to take its attribution at face value.
>The chart shows a big spike this morning, but there was no AWS outage
Are you sure? If hundreds of users simultaneously reported there was some sort of outage, particularly a huge spike like we saw, chances are there was an outage.
>Again, DownDetector can be a useful "is something unusual happening right now" signal
Exactly! Specifically, "is something unusual happening right now with my site, in the eyes of my users?" Every site owner should know when that condition is true. What you think about your site "up-ness" isn't as important as what your users think about your site "up-ness". What you attribute your downtime to, isn't as important as what your users attribute your downtime to (you.)
But that's not the case. It's a false positive.
Pick a DownDetector service and open the page every day for a few days. You'll see it most of the time just reflects people waking up in the US timezones.
In other words, we have hundreds of people saying there was an outage, and 1 person saying there wasn't.
That's a problem AWS needs to resolve, regardless of what they think might be the root cause. If the users weren't experiencing any issues with AWS, I doubt they'd be reporting it.
Your comment about timing is a good point: if people are working with AWS early in the day, and AWS is giving them problems, then they will probably report problems with AWS early in the day. I wouldn't expect them to report problems while they're sleeping.
Yes. AWS was not down this morning.
> In other words, we have hundreds of people saying there was an outage, and 1 person saying there wasn't.
We have hundreds of millions using AWS and AWS-backed services successfully this morning.
I'm out.
The fact that some people accessed AWS without reporting issues does not mean that all people did. For those who had issues, AWS is responsible for dealing with those perceptions.
Indeed, it could have been a fault that affected a subset of users, for example 1 service in 1 availability zone. That's still an outage in the eyes of users, which AWS is responsible for managing. It could have been an issue with a route from 1 ISP. That's still an outage in the eyes of users, which AWS is responsible for managing.
An even better example is the DownDetector page for Facebook, with hundreds of thousands of reports. Do we really think there's no correlation between what DownDetector reports and what users experience?
tl;dr: what users think about your site is more important than both what you think about your site and the reality of your site, and you should be tracking it.
Exactly. If you click through down detector when things are _up_ you'll see people still complaining that $site is down. Could be a local power outage or even a flaky connection in their own home.
Down Detector is one of many signal sources and should have a "credibly" score associated with it that's proportional to the number of people complaining that something's down.
Its attribution of what/who is often incorrect. You'll see "maybe it's more than Big Site X!" comments come up on every HN thread like this citing DownDetector; it's almost never the case, and folks on HN should know better.
Yes? That's how all top-level reporting is going to work. It's not going to tell you which part of your service is inaccessible. It's just telling you that people can't access it. You obviously have to do additional investigation to figure out why people are having trouble.
Scroll up the thread a bit; https://news.ycombinator.com/item?id=39605354
Even here on HN, where people should know better, people take its incorrect attribution as useful info. TikTok isn't down. X isn't down. Google isn't down.
Pull the page up tomorrow and you’ll see the same morning spike there as people wake up.
I'd trust Down Detector a lot more if it was filled with Hacker News community -- people who are able to understand that there's "DNS" and "Routing".. and that your phone can have internet access at home while your home PC does not.
I personally hate Down Detector's graphing because it can make it 'look' like there's an issue when there isn't really... Facebook with 500,000 reports looked as down as Google with 1,000 reports... For equally sized / used entities, I would not trust that "Google" is down with 1,000 reports. I had a coworker ask me what was going on with the internet because "everything is down.. Facebook, google, gmail, microsoft!" (when seeing the Down Detector home page)
DD should normalize the graphs against the service history in some way. A service shouldn't spike because it had 30 reports / hour for a day, then suddenly has 100... when it has a history of being out with 100,000+ reports. The 100 reports are probably mis-reporting, but you can't tell until you dig into each service, one by one, with separate page loads.
hell, i'm surprised Down Detector hasnt been outright sued due to the graphs being an actual honest representation of availability that shitty companies cannot hide
Gmail is also on the list. You can't use FB auth to login to Gmail, can you?
A couple hours ago after watching a video I went to my home page, which usually shows recommendations based on what I've recently watched plus a few videos labeled as sponsored that have nothing to do with any of my interests.
Instead everything on the home page was either a sponsored video, or a movie that was free to view with ads, or something from one of their music products.
I tried from an incognito window to see if it had something to do with being logged in. Normally going incognito loses the history-based recommendations but at least recommends user uploaded content. But now it has just like my logged in home page. No user content. Just ads and videos from Google's movie and music services.
Refreshing gave an error that said something went wrong. I then logged in on that page and again got something went wrong. Another refresh got a page with some user content. Another refresh was the ads and Google stuff page.
A little later it seemed to clear up and now my home page is back to normal.
“Yeah so it turns out when Facebook and Instagram goes down so does Google”
I do not envy the SREs at either company. I'm pretty sure all those other ones use Facebook or Google as their OAuth provider which is why they are all being reported as down.
Edit: I found the following, I wonder if it's still the case.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
I looked up the CPU mentioned in the link from your other comment. It looks like HN handles enormous traffic on about 2x the power of the last Celeron chip ever made.
https://www.cpubenchmark.net/compare/2383vs5793/Intel-Xeon-E...
dang, linked in one of the ancestor comments. But I still suspect you are correct.
> Sorry, we're not able to serve your requests this quickly. reload
Note that this only seems to happen for actions. Doesn't seem to be the case if I am just loading a page quickly.
I saw it just a few minutes ago, but I don't remember the exact wording...
I wasn't overestimating anything, but with how easy it is to write concurrently software today, why limit your site to a single core.
It even looks like Arc, the lisp HN is written in has threads now, but Arc is built on top of Racket and uses Racket's green threads, so it only takes advantage of one CPU core. Racket does have OS threads, but Arc does not use them.
These are good for actual business needs, but bad for resume-driven development.
Couldn't use it two nights ago, IDK why.
I wouldn't trust it as a single source, but in a case like this where our internal monitoring shows a spike of issues with the Google APIs and we can see a huge spike in reported issues for Google on Downdetector starting at the same time, it's useful to confirm that the issues have an external source.
If I can't login to tiktok because FB is down, then tiktok is effectively down for me. When it comes to technology most people don't care about the trip, they care about the destination.
So yea, tiktok isn't "down" but for a lot of people it might as well be, hence coupling your infrastructure/auth on other providers has side effects like this you must take into account.
Its mention should honestly be banned from this site.
Both are above their baselines, but I bet some is just mis-reports, or increases in awareness due to more people checking in.
Meta seems to be the only one really affected from what I can tell.
https://twitter.com/elonmusk/status/1765048551023734801
Downdetector is user reports, not automated monitoring. It's... semi-trustworthy.
Even more so when the tweet in question isn't even a direct claim about Twitter, but just a meme making fun of a competitor.
That's an assertion, not a substantiation. A single tweet does not corroborate that, even if you ignore the fact that most outages of large global services (including some of the outages of these Facebook properties mentioned above) are actually partial degradations.
Edit: actually a more attractive theory, given the very short timelines and near simultaneity of all those failures, is that downdetector itself had a failure, possibly a Meta-dependence, that they noticed and corrected quickly.
Interesting to see that all static content was still working during the outage (at least for Instagram). It was still possible to swipe through all reels (I assume the list was cached).
For process of elimination, do all of these services do multi-platform logins? Or do some not connect to anyone else?
This outage will result in absolutely no ridiculous conspiracy theories.
Actually, if all of them including Xitter went down, maybe things would get better? All the sunlight photons might get sucked in by too many eyeballs though, and there could be grass trampling.
I agree. While I don't think it likely that Facebook or YouTube would enter into it, I'd pretty much bet that DNS being down would cause problems.
And yes, there are bigger issues with that. Much.