GitHub was down
github.com
github.com
The status page says all is well, though: https://www.githubstatus.com/. Hilarious.
All GitHub Pages say
> We're having a really bad day.
> The Unicorns have taken over. We're doing our best to get them under control and get GitHub back up and running.
Then, you would show the status based on the continent.
If a sensor that's basically in the same datacenter says you're up, but the route into the datacenter is down, then what? multiply this by the complexity of the whole site, and monitoring it all with 100% fidelity is impossible. Not that it's not worth it to try, there's a team at GitHub that works on monitoring, but beyond motivation about keeping the SLA up, as a customer, unless you notice it's down, is it really down? In a globally distributed system, downtime, except for catastrophic downtime like this, is hard to define on a whole-site basis for all customers.
I don't think anybody asked for 100% fidelity. We are talking about a complete outage that affected at least North America and Europe. If the status page shows green in such a case, its fidelity is around 50%. People expect better from GitHub.
Total outages are rare enough, and there's enough other work, that spending time building a system for that, just doesn't seem like the best use of their time. though I'm biased, having faced that exact question from the inside, at different company.
This is impossible regardless of how godlike the design is... Nobody is asking for 100% fidelity.
So long as I can fetch/commit to my repos, pretty much everything else is of secondary, tertiary, or no real importance to me.
(At work, I do indeed have systems running that monitor 200 statuses from client project homepages, almost all of which show better that 99.999% uptimes. And are practically useless. Most of them also monitor "canary" API requests which I strive to keep at 99.99% but don't always manage to achieve 99.9% - which is the very best and most expensive SLA we'll commit to.)
Good reason why companies shouldn't be using Twitter/X for status updates anymore!
An all around stupid decision. That said, if management is that shitty, the platform probably won't be attractive for long anyway.
Facebook/Instagram were successful despite that to a degree, but this decision probably still did a lot of damage to their relevancy and user numbers.
FB/IG/Whatsapp have half of humanity logging into their services once per month, so I'm not sure how much better they could be doing if they didn't have a login wall.
Meanwhile, Twitter (with no login wall) never broke 500mn. Like, personally I totally take your point about status updates but I'd have used my Twitter account a lot more if I'd needed to log in to see the content.
Instagram is closer to broadcast, but it was always closely tied to the mobile app experience and the "follower" mentality. People didn't really share links to Instagram posts in other online venues in the beginning.
Twitter was always unique. It existed before smartphones, and there was a good chunk of years where people without smartphones would read twitter posts on desktops. Its producer/consumer distribution is much more skewed, many twitter users never post. Tweets were always getting posted to places like HN, reddit, discussed in news articles, etc.
I think Twitter's (former) position as a broadcast medium à la TV, radio, and newspapers is unique among social networks. There's a reason why Twitter was the place for journalists, politicians, academics, fire departments, web service status alerts, etc.
People use this page for guidance. I guess now we know how much it can be trusted.
Now 4 out of 10 services are marked as "Incident", yet most of the others are also completely dead.
Declare an incident first, investigate later.
Cheating SLAs by delaying the incident is a good way to erode trust within and without.
If that would be the best way to deal with it- why is literally no one doing it this way and what does that tell you?
This defeats the purpose of a status dashboard and is effectively useless in practice most of the time from a consumers point of view.
If your reliability metrics have lots of false positives, that's on you and you'll have to write down some reason why those false positives exist every time.
Then that company could decide for itself whether to update manually with "not a reliability issue because X".
This lets consumers avoid being gaslighted and businesses don't technically have to call it downtime.
I don't think GitHub has recovered from the monthly incidents that keeps occurring. Quite frankly it is the expectation that something will go down every month at GitHub which shows how unreliable the service is and this has happened for years.
I guess this 4 year old prediction post really aged well after all about self-hosting and not going all in on Github [0]
We are experiencing interruptions in multiple public GitHub services. We suspect the impact is due to a database infrastructure related change that we are working on rolling back.
Also probably a class action suit lurking somewhere in there eventually.
Migration to a new host takes another 15 seconds thanks to both zfs and containers.
I don't know how many GitHub downtime reports I've seen during that time, we're probably into high dozens by now.
I've been moving most of my projects off of GitHub and into Gitea, and will continue to do so.
I remember a time when systems would boast about their "five nines" uptime. It was before anything "cloud" appeared.
Hope I don't appear in the incident report.
I had a github page that was public, but it was made private and the DNS config was removed. Fast forward to today. I made the private repo public again and forced a deploy of the page without making a new commit. It said the DNS config was incomplete, so I tweaked it and hit "check again" and github went down.
Probably unrelated, but the timing was spooky.
So, you're both responsible and not responsible at the same time :)
> Hope I don't appear in the incident report.
Appearing in an incident report with your HN username could be pretty funny...
For the record, the status page eventually got updated - around 7 minutes after this submission was created.
For a service like GH, anything more than 30 secs is unacceptable
And simple HTTP monitoring would be too flappy for a public status page.
- an incident be declared internally to github
- support / incident team submits a new status page entry (with details on service(s) impact(ed))
- incident is worked on internally
- incident fixed
- page updated
- retro posted
Even aws now seems to have some automation for their various services per region. But it doesn't automatically show issues because it could be at the customer level or subset of customers, or subset of customers if they are in region foo in AZ bar, on service version zed vs zed - 1. So they chose not to display issues for subsets.
I do agree it would be nice to have logins for the status page and then get detailed metrics based on customerid or userid. Someone start a company to compete with statuspage.
```
Received a 503 error. Data returned as a String was: <!DOCTYPE html> <!- -
Hello future GitHubber! I bet you're here to remove those nasty inline styles, DRY up these templates and make 'em nice and re-usable, right?
Please, don't. https://github.co...
```
That's where it's cut off on my screen.
Curious what the link is :)
I like to think, someone did.
https://www.bleepingcomputer.com/news/security/github-action...
give the poor github ops folks a second to get things moving.
You ideally do not want to be making a decision on whether to update a status page or not during the first few minutes of an incident, bean counters inevitably tend to get involved to delay/not declare downtime if there is a manual process.
It is more likely the threshold is kept a bit higher than couple of minutes to reducing false positives rates, not because of manual updates.
[1] https://www.atlassian.com/software/statuspage/integrations
That happened to me a while back with an app listing that was almost 10 years old because the server I was hosting the policy on went down. Ironically, I switched it to Github pages so it wouldn't happen again.
Another reason is that MS may be in phase when it will ask to pay for using GitHub just for reads (rate limiter).
When you would usually create a PR, you use `git format-patch` to create a patch file and send that to whoever is going to merge it.
They create a branch and use `git am` to apply the patch to it, review the changes, and merge it to main.
It is nice that git supports multiple remotes, though. It feels good to know that `git push` might not work for my project right now, but I know `git push srht` will get the code off of my laptop.
I also had to setup a bidirectional mirror back when bandwidth to some countries was restricted. We would push and pull as normal, and a job would keep our mainline in sync.
It is sad that most organizations forget that git is distributed by nature. We often get requests to setup VPNs and all sorts of craziness, when a simple push to a bare mirror would suffice. You don't even need anything running, other than SSH.
Well, that's how it was designed to work! The whole point of Git is that it's a distributed version control system, and doesn't need to rely on a centralized source of truth.
The real reason not to use github anyway though is that it's terrible (the basic "github model" for doing code review was basically made up on the back of a napkin IMO)
What made me laugh though was when the "X is functioning normally" immediately followed by "X is degraded, continuing to monitor" messages that kept popping up then right back to "normal" again, all in the same 30 second timespan... made me giggle
This is a pretty good place to check. The lag is pretty minimal traditionally.
At the time of posting everything is broken.
Here is a good article on how to prepare for the situations like that, when GitHub is down: https://gitprotect.io/blog/github-restore-and-github-disaste...
Nix: barfs voluminous errors I've never seen before
Me: whaaaat the farrrrk
* nixos updates are pulled from a github repo
Also they had IPFS attempts, but not finished.
Update - Issues is experiencing degraded availability. We are continuing to investigate.
Aug 14, 2024 - 23:19 UTC
Update - Git Operations is experiencing degraded availability. We are continuing to investigate.
Aug 14, 2024 - 23:19 UTC
Update - Packages is experiencing degraded availability. We are continuing to investigate.
Aug 14, 2024 - 23:18 UTC
Update - Copilot is experiencing degraded availability. We are continuing to investigate.
Aug 14, 2024 - 23:13 UTC
Update - Pages is experiencing degraded availability. We are continuing to investigate.
Aug 14, 2024 - 23:12 UTCSeeing it all kind of went sideways at the same time, my money is on the typical load balancer config rollout snafu.
"As part of a routine configuration deploym..." [splat]Time to consider self-hosting like the old days instead of this weekly chaos at GitHub.
> We are experiencing interruptions in multiple public GitHub services. We suspect the impact is due to a database infrastructure related change that we are working on rolling back. Aug 14, 2024 - 23:29 UTC
Seems like they’re back up though. Or at least the Rust blog is back up.
"Why isn't this project done yet?"
"Didn't you hear? GitHub is down!"
and I get to go out for a long lunch
EDIT: The reply link is no longer available.
Update - Packages is experiencing degraded availability. We are continuing to investigate. Aug 14, 2024 - 23:18 UTC
Update - Issues is experiencing degraded availability. We are continuing to investigate. Aug 14, 2024 - 23:19 UTC
Update - Git Operations is experiencing degraded availability. We are continuing to investigate. Aug 14, 2024 - 23:19 UTC
EDIT: The reply link is no longer available again.
Everything is red now. Nearly lunch time in New Zealand.
Good luck to the devs and dev-analogues involved in getting the ship righted.
[error] [auth] Response content-type is text/html; charset=utf-8 (status=503)
The static content on the error page might also be on akami or cloudflare side.
the images on the page are all just base64 encoded right into the html
Unicorn has a slightly different architecture.
Instead of the nginx => haproxy => mongrel cluster setup
you end up with something like: nginx => shared socket => unicorn worker pools
When the Unicorn master starts, it loads our app into memory. As soon as it’s ready to serve requests it forks 16 workers. Those workers then select() on the socket, only serving requests they’re capable of handling. In this way the kernel handles the load balancing for us.fatal: unable to access: 502
cli, web, and iOS app :-/
It stated this but with Twitter; it will monitor latest tweets searching for a custom word combo and raise a server alert when found. I found it hilarious. Will post the source once GitHub is back on.
Both seem to be doing too much all at once. But really it is worse with Github if this is what Microsoft stewardship is incidents every, single week and each month guaranteed for years.
Anyways. #hugops for the GitHub team.
What’s the point of bringing up twitter? It is strange to seek victimhood for a petulant billionaire. Of course, it is worse with GitHub because GitHub actually provides useful functionality.
The same folks complaining about something at GitHub going down are the same people that stay and are willing to tolerate the regular incidents and chaos on the site.
It is the fact that not only the Github incidents have been happening for years, it has gotten worse as there is an incident every month.
> Of course, it is worse with GitHub because GitHub actually provides useful functionality.
That isn’t an excuse for tolerating regular downtime for a site with over 100m+ active users, especially with it running under Microsoft stewardship who should know better.
Any other site with that many users and with a horrendous record of downtime like Github would be rightfully branded as unreliable. No excuses.
Good opportunity to think about mirroring your repos somewhere else like Gitea or Gitlab.
This is not to mention all the projects that rely on Github for project management, devops, community and support desk.
GitHub is also very international, I doubt isolated netziens like those from China are shielded from this outage. I imagine very, very few software shops are unscathed by this. The whole affair is very on brand for 21st century software which is to say pitiful.
Another bonus is that we don't pay Microsoft.
- Make everyone dependent on all the centralized features of your software to use Git[1][2]
- Now you have a de facto centralized, hard to use VCS with thousands of SO questions like “my code won’t commit to the GitHub”
- Every time you go down a hundreds-of-comments post is posted on HN
How to get bought for a ton of cash by a tech mega corporation.
[1] Of course an exaggeration. Everyone can use it in a distributed way or mirror. The problem occurs when you’re on a team and everyone else doesn’t know how to.
[2] I’m pretty sure that even the contributors to the Git project rely on the GitHub CI since they can’t run all tests locally.
Senior: Ah found it! Let's just rollback one revision on the db. Newguy: let me fix this! `kubectl rollout undo ... --to-revision=1` Newguy: Ok, Started rollback to revision one! Senior: Uh-oh..