It is a shame that we have to find out about outages on HN.
It is a shame that we have to find out about outages on HN.
Marketing and sales. Not /s
We are in dire need of something crowdsourced, or where someone like DataDog or other telemetry systems offer you the ability to share non sensitive metrics publicly for various cloud or SaaS systems that they publish.
Edit: y’all are amazing with these monitoring tools!
At least for Steam, I've found it pretty useful.
Somewhat but not entirely similar to BGP looking glass systems.
This is literally what I'm building right now. See reply above: https://news.ycombinator.com/item?id=31825239
Shoot me an email if anyone is interested in getting beta access.
Looks like there’s been some errors too
Disclaimer - I work at New Relic but not on this.
When I got hired at Amazon in 2001 we had a "gonefishin" page that was a static page that would be served in the event of an outage (this was before status pages, but it was kind of the same thing -- public acknowledgement of a major incident). The standard protocol was within minutes of a sev 1 to make a decision to display the GF page once it was confirmed that the whole site was down and then work to fix the issue.
By the time I left in 2006 that was no longer policy since reporters had setup monitoring for that page to detect outages and report on service availability so they just let it crash and return 500s or whatever the failure mode was. Optimize for making the job of external agencies doing reporting on their availability harder instead of easier.
But I do agree they should be able to monitor things better and show some sort of update on their status page as soon as possible.
What would have happened if Zoom had worked fine on their end, but I was randomly unable to connect? Perhaps it would have been fine—they would have been understanding, and we would have rescheduled for another day. Perhaps if they hadn't been understanding, I shouldn't have wanted to work for them anyway.
But, I don't know. I wanted to work for them, and I was competing with other candidates who presumably interviewed on different days. Hiring processes are inherently imperfect, and lots of things can be consciously or unconsciously treated as a red flag.
(And yes, lots of other things could have happened on the day of the interview. But I still find this scenario particularly scary to think about.)
Exactly, so it’s weird to worry about a Zoom problem in particular. If anything it’s a little better now since most people are conditioned to think of technical problems as less likely the affected persons fault (that’s why I referred to the alternative as “morons”) - even if you left yourself plenty of time and did everything right and public transit fucked you over it was never a good look.
We use Skype for Business, which is so flaky at times that the default assumption if somebody is not joining is that the system conked on her.
If an electric company serving a million people leaves 100 of them in the dark, it's still an "outage."
Why give a free pass to Zoom? Because it's a tech company, and we've been trained to accept failures as the cost of admission?
But over the last 4 years specifically I not only understand it I can't imagine not having a service like it.
Disclaimer: I don't work for pingdom and my current company doesn't use their services, I have in the past, they're pretty good, but I'm just using them as an example here