Zoom.us is down
us02web.zoom.us
us02web.zoom.us
It is a shame that we have to find out about outages on HN.
Marketing and sales. Not /s
We are in dire need of something crowdsourced, or where someone like DataDog or other telemetry systems offer you the ability to share non sensitive metrics publicly for various cloud or SaaS systems that they publish.
Edit: y’all are amazing with these monitoring tools!
At least for Steam, I've found it pretty useful.
Somewhat but not entirely similar to BGP looking glass systems.
This is literally what I'm building right now. See reply above: https://news.ycombinator.com/item?id=31825239
Shoot me an email if anyone is interested in getting beta access.
Looks like there’s been some errors too
Disclaimer - I work at New Relic but not on this.
When I got hired at Amazon in 2001 we had a "gonefishin" page that was a static page that would be served in the event of an outage (this was before status pages, but it was kind of the same thing -- public acknowledgement of a major incident). The standard protocol was within minutes of a sev 1 to make a decision to display the GF page once it was confirmed that the whole site was down and then work to fix the issue.
By the time I left in 2006 that was no longer policy since reporters had setup monitoring for that page to detect outages and report on service availability so they just let it crash and return 500s or whatever the failure mode was. Optimize for making the job of external agencies doing reporting on their availability harder instead of easier.
But I do agree they should be able to monitor things better and show some sort of update on their status page as soon as possible.
What would have happened if Zoom had worked fine on their end, but I was randomly unable to connect? Perhaps it would have been fine—they would have been understanding, and we would have rescheduled for another day. Perhaps if they hadn't been understanding, I shouldn't have wanted to work for them anyway.
But, I don't know. I wanted to work for them, and I was competing with other candidates who presumably interviewed on different days. Hiring processes are inherently imperfect, and lots of things can be consciously or unconsciously treated as a red flag.
(And yes, lots of other things could have happened on the day of the interview. But I still find this scenario particularly scary to think about.)
Exactly, so it’s weird to worry about a Zoom problem in particular. If anything it’s a little better now since most people are conditioned to think of technical problems as less likely the affected persons fault (that’s why I referred to the alternative as “morons”) - even if you left yourself plenty of time and did everything right and public transit fucked you over it was never a good look.
We use Skype for Business, which is so flaky at times that the default assumption if somebody is not joining is that the system conked on her.
If an electric company serving a million people leaves 100 of them in the dark, it's still an "outage."
Why give a free pass to Zoom? Because it's a tech company, and we've been trained to accept failures as the cost of admission?
But over the last 4 years specifically I not only understand it I can't imagine not having a service like it.
Disclaimer: I don't work for pingdom and my current company doesn't use their services, I have in the past, they're pretty good, but I'm just using them as an example here
The pinging service and status page, at our scale, is free and our status page is actually useful and automated for 90% of our stuff.
Status pages that raise customers confidence in your service are good from a marketing perspective.
Automatically publishing uptime data without human review might be bad from a marketing perspective, if you don't trust the engineering department to actually deliver or if your service depends on too many external dependencies.
If acknowledging downtime causes you to violate an SLA and pay a bunch of penalties then you don't want some automated script to trigger it.
The SLA would not depend on what the status page reported, but the actual downtime. If the script malfunctioned, you wouldn't need to pay out because it wasn't actual downtime contractually - and if it was actual downtime, I guess it makes it harder to squeeze around but that's only if someone is carefully taking legally admissible evidence from the first minute the status page reports (which screenshots alone don't always meet).
Note that as another commenter said this only tells if the server is up, not if the service is working properly or not. In this case it won't work since the Zoom website is loading fine but the meetings don't work properly.
Alternatively use a centralized service (YouTube, Twitch).
Or, there's also the P2P Jami (but this requires the participants to install something):
What a weird coincidence.
like, yes I get it multi regions makes you safer, but given vendor lock if aws fucks up across the board with a bug, we all fuck up across the board.
Other features like recording a much more reliable if they are implemented server side.
Maintaining productivity when those tools cease to function encourages the employer to underinvest in tools and burden the employee.
Unless you don't want there to be a meeting(s)... just say'n.
I have not used it, myself, so I can't report on its effectiveness.
I was in a Zoom call with a big Chinese company some months ago, and they used something that looked the same as zoom but had a Chinese name.
Edit:
I think it was this one:
https://www.cnbc.com/amp/2020/08/03/zoom-to-halt-direct-sale...