The status page says all is well, though: https://www.githubstatus.com/. Hilarious.
The status page says all is well, though: https://www.githubstatus.com/. Hilarious.
Good reason why companies shouldn't be using Twitter/X for status updates anymore!
An all around stupid decision. That said, if management is that shitty, the platform probably won't be attractive for long anyway.
Facebook/Instagram were successful despite that to a degree, but this decision probably still did a lot of damage to their relevancy and user numbers.
Instagram is closer to broadcast, but it was always closely tied to the mobile app experience and the "follower" mentality. People didn't really share links to Instagram posts in other online venues in the beginning.
Twitter was always unique. It existed before smartphones, and there was a good chunk of years where people without smartphones would read twitter posts on desktops. Its producer/consumer distribution is much more skewed, many twitter users never post. Tweets were always getting posted to places like HN, reddit, discussed in news articles, etc.
I think Twitter's (former) position as a broadcast medium à la TV, radio, and newspapers is unique among social networks. There's a reason why Twitter was the place for journalists, politicians, academics, fire departments, web service status alerts, etc.
FB/IG/Whatsapp have half of humanity logging into their services once per month, so I'm not sure how much better they could be doing if they didn't have a login wall.
Meanwhile, Twitter (with no login wall) never broke 500mn. Like, personally I totally take your point about status updates but I'd have used my Twitter account a lot more if I'd needed to log in to see the content.
Also probably a class action suit lurking somewhere in there eventually.
Now 4 out of 10 services are marked as "Incident", yet most of the others are also completely dead.
This defeats the purpose of a status dashboard and is effectively useless in practice most of the time from a consumers point of view.
If your reliability metrics have lots of false positives, that's on you and you'll have to write down some reason why those false positives exist every time.
Then that company could decide for itself whether to update manually with "not a reliability issue because X".
This lets consumers avoid being gaslighted and businesses don't technically have to call it downtime.
Declare an incident first, investigate later.
Cheating SLAs by delaying the incident is a good way to erode trust within and without.
If that would be the best way to deal with it- why is literally no one doing it this way and what does that tell you?
Migration to a new host takes another 15 seconds thanks to both zfs and containers.
I don't know how many GitHub downtime reports I've seen during that time, we're probably into high dozens by now.
I've been moving most of my projects off of GitHub and into Gitea, and will continue to do so.
We are experiencing interruptions in multiple public GitHub services. We suspect the impact is due to a database infrastructure related change that we are working on rolling back.
So long as I can fetch/commit to my repos, pretty much everything else is of secondary, tertiary, or no real importance to me.
(At work, I do indeed have systems running that monitor 200 statuses from client project homepages, almost all of which show better that 99.999% uptimes. And are practically useless. Most of them also monitor "canary" API requests which I strive to keep at 99.99% but don't always manage to achieve 99.9% - which is the very best and most expensive SLA we'll commit to.)
Then, you would show the status based on the continent.
If a sensor that's basically in the same datacenter says you're up, but the route into the datacenter is down, then what? multiply this by the complexity of the whole site, and monitoring it all with 100% fidelity is impossible. Not that it's not worth it to try, there's a team at GitHub that works on monitoring, but beyond motivation about keeping the SLA up, as a customer, unless you notice it's down, is it really down? In a globally distributed system, downtime, except for catastrophic downtime like this, is hard to define on a whole-site basis for all customers.
I don't think anybody asked for 100% fidelity. We are talking about a complete outage that affected at least North America and Europe. If the status page shows green in such a case, its fidelity is around 50%. People expect better from GitHub.
Total outages are rare enough, and there's enough other work, that spending time building a system for that, just doesn't seem like the best use of their time. though I'm biased, having faced that exact question from the inside, at different company.
This is impossible regardless of how godlike the design is... Nobody is asking for 100% fidelity.
All GitHub Pages say
> We're having a really bad day.
> The Unicorns have taken over. We're doing our best to get them under control and get GitHub back up and running.
I don't think GitHub has recovered from the monthly incidents that keeps occurring. Quite frankly it is the expectation that something will go down every month at GitHub which shows how unreliable the service is and this has happened for years.
I guess this 4 year old prediction post really aged well after all about self-hosting and not going all in on Github [0]
I remember a time when systems would boast about their "five nines" uptime. It was before anything "cloud" appeared.
People use this page for guidance. I guess now we know how much it can be trusted.