GitHub Incident
githubstatus.com
githubstatus.com
Compared to GitHub, how often does your Gitea instance go through code changes? How many major features have been switched on in that instance during the past year without any hiccups?
https://blog.gitea.io/2022/02/gitea-1.16.0-and-1.16.1-releas...
https://blog.gitea.io/2021/08/gitea-1.15.0-is-released/
But I'm sure Github changes a lot more. However, the question is if those changes are worth the instability. For my workflow I just want my repos to be accessible, have my source, issues, and wiki accessible, and of course my basic git operations working. If that core functionality is actually not working at times, I'd definitely consider changing the git provider.
The problem is that ultimately the downtime itself matters and not the reason, and if you don't need any of the features that GitHub offers, then the self-hosted route is a better option.
In particular, all of the new features whose addition causes instability.
New features breaking is a lot more understandable - even expected - than regressions and refactoring failures.
Here is a early version of their landing page from 2008 (the year they launched): https://web.archive.org/web/20081111061111/http://github.com...
Notice the logo says "Social code hosting" and the messaging of the page is mostly around popular repositories, collaborative features and other social elements.
And when they do go down, how long will they be unavailable? When this happens (and it will happen) please post here so we can say "why are you running infrastructure yourself?"
Now I’ll tell you the real problem with it: it worked great for people who were located in the U.S., but it was terrible for certain remote employees due to increased latency, and the fix for that would be to set up a complicated replication setup that puts data closer to people not near the main instance.
There’s obviously other issues, but that is one that is probably commonly overlooked that most SaaS solutions can do better on.
I currently manage an internal gitlab instance with some remote coworkers, but I haven't heard any complaints about latency. Although tbf, I feel like we've barely scratched the surface of what gitlab supports
When I moved to a new company, the GitLab instance was still up, and last I was aware, many years down the road, it is still in use. Feature-wise, we were very much satisfied, and the GitLab CI model with Docker executors worked very well for us.
Worst case is losing commits between going down and backups, but even then it'll be minimal if you set your backup to simply push to a another origin.
I pay for someone else to worry about most services, but git doesn't have to be one of them.
Very simple to set up, and as long as your data volume is not lost (which can happen, but is very rare), such a system would recover from a host failure by itself faster than most people could react to it going down. In the unlikely case that you lose the EBS volume, you'd just recover from snapshots manually.
If something like Git went down — which it never — it was a few clicks away from coming back
Working with competent people is a blast
Have we as an industry forgot how to make systems? that's pretty telling if so.
Why are people so utterly afraid of running their own systems.
If a SaaS platform goes down it's followed by "#hugops; running large systems is hard!", when someone says "Well, if it's hard, why not break it into smaller pieces and run it myself" it's met directly with "Do you think you can do better!?!".
Honestly, probably, maybe not. Why does it matter?
Those two things cannot be true simultaneously. You cannot say "running big systems is hard" and "they can run a big system better than you can run a small system".
I'm not going to fall into the trap of arguing points like the fact that if you own the data you can have your own D/R & backup strategies and the fact you can run maintenance's on your own time (hell: you can run maintenance's to migrate things around, which is more than github can do, which is not an indictment of them, just the nature of being a hosted service).
Honestly, I miss sysadmins. People who were not afraid of servers, even if they said "no" sometimes. Because, you know, being afraid of servers and hosting things is just pathetic.
When you run your own servers, not only do you need to maintain them, you are also responsible for their certification and getting them audited.
As a manager in a medium sized company, I think it is essential to identify what "needs building" and what could be bought to ensure the needs of stakeholders are met in a timely manner.
Back when I was working in eCommerce (PCI-DSS Tier 1, Cardholder data on premises) it was basically impossible to use hosted services.
As then, it is now: "Compliance" is just enough nebulous of a word to prevent critical thinking.
Sure you can build your own infrastructure but don’t pretend that’s superior. Each has its pros and cons.
The last line probably is making fun of their reaction — any person as dogmatic as the parent commenter would comment someone with their own infrastructure going down with that comment.
They can absolutely be true simultaneously. In fact, I can't name a logic system where those two statements comprise a contradiction. They are logically completely unrelated to each other.
DNS goes down for a particular dot com? BGP hijinks? Backhoe through the major fibre serving your workplace?
All of these are going to happen and will lead to downtime out of your control or the control of the SaaS provider.
Thankfully git is designed to be decentralised and developers can continue to work even if they have to use paper books as reference material and not Stack Overflow.
Rule of thumb is that 10^100 of anything doesn't fit in the observable universe.
> Yes, I know, GitHub is 10⁹⁹⁹ times larger than our puny Gitea instance, and that's why they're having issues
Exactly. That is also my overall point. [0] It makes no sense to go 'all in' on GitHub and something goes down and everyone is stuck once again.
No pressure, but I'm seriously considering standing one up but it would need to be production ready.
Actually not "vendoring" dependencies from github is very much not production ready in many cases.
So the big issue for any non-trivial or highly available setup is going to be how you get high availability for the local storage volume. There are tons of options with various tradeoffs and levels of complexity here--simple local disk RAID, distributed filesystems like gluster or ceph, etc. I think this is the real crux of getting a good gitea instance going.
No matter what figure out and test a good backup and recovery solution!
Thanks that's quite helpful. If we do it in prod we'll have to accept single-instance.
No k8s or anything like that. An Ubuntu LTS virtual machine on top of I don't know what hypervisor (probably Hyper-V). Gitea requires very few resources.
Data is stored in PostgreSQL 11. Probably should be upgraded at some point, but it's working fine for now.
A Drone CI instance runs on that same machine. It controls CI workers on a dozen other VMs.
Everything is in Docker containers under docker-compose for two reasons: easy upgrades, and the ability to shove all data into a single directory.
caddy for HTTPS.
Sonatype Nexus for package repositories (caching upstream npm, nuget, maven, composer, and our own internal repos). This one is pretty heavy and will probably be moved to a separate machine at some point.
Never had any issues with any of these in about three years we've been running this setup.
I was a bit facetious with the "zero downtime" statement. It has about 30 seconds of downtime per month, but it is always planned and outside of working hours deep into the night.
There's been zero seconds of unplanned downtime.
Basically you do:
$ docker-compose pull
$ docker-compose up -d
and half a minute later it is running new versions. If anything is broken (it hasn't been yet), call the sysadmin guy, and a few minutes later he restores a VM snapshot.Backups are being done on the hypervisor's level. I don't know about that much.
That’s frightening, we’ve been using it for multiple years now. It’s running fine and the only short downtime is every few weeks for the update.
What happened to your instance?
Tried restoring it to the older version to run through the upgrade path and the process didn't want to work. Wish I could be more specific here but I gave up on it and decided it'd just be quicker to whip up a script to recreate all the groups/projects/CICD in a new install and push repos from backups
Don't let your GitLab server get outdated and don't consider it a valid backup unless it was taken with the latest version to avoid the same
Also it wasn't relevant here but make sure you're backing up the etc files too, while that wasn't an issue for me it could easily trip you over in a worst case scenario ( https://docs.gitlab.com/charts/backup-restore/backup.html#ba... )
Not a complaint, it's just one of the hats that needs wearing sometimes, would have been easier on me if someone else was handling it though :P
We’ve been updating about as soon as an update is released for a while now, with very good results. (I’d rather have to restore a backup or a snapshot than to be running without the latest security fixes. If gitlab is down for an hour (didn’t happen in ~3 years) at least all developers have their local repo.)
We've been using exec and ssh runners when containers are not enough (for example, when you need to juggle a few VMs or build some complicated Docker image, and using docker-in-docker is not fun).
https://docs.drone.io/pipeline/exec/overview/
https://docs.drone.io/pipeline/ssh/overview/
GitHub Actions are nice… when they're actually working.
I can understand the problem for larger development teams, as a lot of communication and workflows can happen via GitHub Pull Requests or similar.
When you first setup CI, you don't actually just develop for the CI environment first, that's the second step after you know how to perform the same things locally. So not sure why there would be "independent CI for everything".
I mean, if you're working on a project that has tests, code coverage and binary builds setup, you usually have a `Makefile` or `package.json` or whatever to run your scripts, and your CI setup just calls those very same scripts but in their environment (sometimes with different arguments/environment variables).
Not sure why it would be different for GitHub Actions. It's certainly how I use it day-to-day.
> Not sure why it would be different for GitHub Actions.
because vendor lock-in. GitHub doesn't want to make it easy for you to switch.
I don't think github is trying to create lock-in, I think rather they were trying to make a way to easily share actions (not sure what other CIs systems are designed to have an ecosystem of publicly shared actions). The actions are public and therefore easy to make something that interprets them.
I can only guess at some point there will be a push for CIs to converge on some "actions" standard, maybe?
Gitlab's open source runner supports parsing the ci yml and running a job locally, presumably using the same code as the platform, like:
gitlab-runner exec docker some-job-name
Though they would both suffer on any dependencies on platform hosted environment vars, secrets managers, etc.Sure, have all the heavyweight stuff in separate scripts that are just called, but platform specification/multiple platform builds/specifics of caching/secret handling/deployment handling are always different. Some tools (e.g. codecov) do abstract over some platforms, but not all, and the GitHub Actions model of "here is a literally pre-prepared step in your pipeline" can be pretty appealing.
It's literally, pick your poison, and resign yourself to reimplementation if you ever need to switch platforms.
Both my personal projects and my $dayjob repositories have every test, etc triggered via `make test` or `test.sh`, then the GitHub Actions workflow YAML just `run`s it. Secrets also work fine - the makefile / shell script expects them to be defined as env vars, so the developer running them locally just needs to define those env vars regardless of how they obtained the secrets.
Parallelising your build/deployment will likely also be harder to do.
Sure? On your dev machines the dependencies are already installed. On GHA VMs the network is fast enough that installing deps is not slow. What's the problem?
And if you really have a problem, presumably you have some master tarball / container image that your devs use to set up their dev machines because installing your deps is so complicated, so scp / pull that in your script?
>can't utilize cache
A cache is necessary for CI VMs that are cleaned for every run. When building locally your dev machine already has everything the cache would have.
>and won't have a dependency graph as well. > >Parallelising your build/deployment will likely also be harder to do.
You realize the two things shell scripts are good at is running commands either series or parallel, exactly how you want them? Instead of learning a brand new DSL to be able to do `if` and `&`, you just write `if` and `&`. This is already covered in the comment I linked.
SourceHut is one source code hosting with the right idea here. Its CI only has the equivalent of the `run` command for running shell scripts - https://man.sr.ht/builds.sr.ht/manifest.md#tasks
> it's because you chose to do it to yourself.
I didn't choose shit. My company did. Why are you putting this on me?
> Both my personal projects and my $dayjob repositories have every test
Congrats on not actually using GitHub actions? I guess?
So many people here sucking Microsoft cock. And there is yet another incident today! What's that make, three days in a row now? Four, if we're actually counting. They aren't even hitting 2 nines uptime. Two. Fucking. Nines. Going on many years now. But apparently that is just fine because everyone is running self-hosted infra in parallel to their cloud shit.
>But apparently that is just fine because everyone is running self-hosted infra in parallel to their cloud shit.
In your haste to complain about downvotes and accuse other people of "sucking Microsoft cock", you forgot to actually read the comment you replied to.
Things in the comment you replied to:
1. An assertion that one can write their CI in a script that is not tied to one CI vendor, so that it's easy to run the same steps as what the CI does locally or in another CI. ie, no lock-in to the current CI.
Things not in the comment you replied to:
1. An assertion that I run self-hosted CI in addition to GitHub Actions.
2. An assertion that GitHub has good uptime.
There's also just the number of things it checks. jest runs, lint/build, e2e and acceptance tests, 2 docker builds pushed into ghcr, and then ansible to deploy. It's mildly error-prone to do myself, especially the docker and ansible steps because that's where the secrets come in.
So sure, it CAN be done manually, but the entire point of CI/CD is to do everything consistently, repeatedly, and without the risk of manual error. It took me hours to figure things out the first time. Why would I want to risk doing things manually now?
Solo? Yes. In a team? Less often just because you might have credentials stored in secrets that are not shared with the entire team for security reasons.
In larger corps it's easier to wait it out that go through the trouble of escalating permission requests.
But I do get your point. Personally for deployments I have Ansible playbooks that are invoked via straighforward calls which can be done locally.
Every single time I start doing development on Actions, GitHub goes down or actions start failing. I'm not sure if this is saying something about me, or GitHub.
That's not really the point though.. If you have services in github etc, why would you replicate all of them locally too for unplanned downtime ?
I just don't feel the need to defend GitHub here for the reason of : "Well you should be able to manage locally anyway"
It can't be "remove all SPOF or you aren't doing your job". How many offices are resilient to continue working if the building power goes out?
> Diffing Adobe Photoshop documents is no longer supported
GitLab is a large and complicated piece of software, which is a double edged sword: you do get an integrated issue tracking solution, Docker Registry with integrated access controls, as well as a pretty decent CI solution, which is great. At the same time, however, there's a chance for one of those subsystems to start misbehaving and bring down the whole thing (e.g. Omnibus install with default configuration for where to store container images --> running out of space; especially relevant if you don't find a good way to clean up old images), which may or may not be harder to solve than having 3 separate integrated systems.
Also, GitLab is pretty resource hungry and migrating to another solution actually improved the performance of certain systems/parts of functionality bunches: Gitea is really fast, Drone CI is about on par with GitLab CI and Nexus is a bit faster than GitLab Registry, however also eats too much RAM (e.g. the same problem within both Ruby and Java, Nexus seems to fail to start up below having approx. 1 GB of RAM, GitLab needed at least 2 - 3 GB of RAM).
However, running GitLab at work is still a pretty good idea, as long as keeping the instance up and running, as well as updated is someone else's responsibility and you can focus on just being the end user. At the same time, however, the update paths can be problematic sometimes, as was the case with my container based Omnibus installation that was multiple major versions out of date.
But for personal use? The Gitea, Nexus and Drone setup is perfectly sufficient, alongside maybe some issue tracking software (be it Kanboard, OpenProject or something else). In short, to the folks who are considering self-hosting in the comments: you can definitely try it out and have it be a learning experience, maybe on an affordable VPS or a device you have laying around!
Furthermore, there's nothing really preventing you from mirroring your repos over and keeping them in multiple systems, alongside any backups that you might be making!
You shouldn't ever trust anything to never go down again. It's just inevitable that services go down from time to time, you can't change that unless we have a major breakthrough in software engineering where somehow people stop making mistakes.
So instead of figuring out if you should "trust it to go down again or not", assume it will at one point in the future and plan accordingly. This is how you build resilient systems.
welp.
I'm thankful that at least I'm not on security investigating the severity of the okta hack.
For this one, I found myself rage-screaming at my monitor as their issue markdown preview API was timing out ~80% of the time.
Github enterprise is probably happening for us this year. I don't give a shit about the new features in the public build. I just want it stable and working at this point.
> Until the next time GitHub goes down again (hopefully that won't be in another month's time).
Oh dear. [0] Never expected it to be this soon or that bad. Expected it to be down in a few weeks at least. Last time this happened was just 5 days ago. [1]
Yet I see complaints here of GitHub being unreliable as predicted and even some considering to eventually self-host despite paying $4 a month even.
I don't think it make sense to go 'all in' on GitHub, does it?
There are a couple of strategies one could rely on to mitigate the risk of github downtime:
- create a map of features between github and competitors, find the common denominator and risk accept anything that is unique to github but viral for your business/product;
- maintain a primary self hosted instance of gitea/gitlab/sourcehut and use github as a mirror; - use github as a primary platform, maintain mirror with public gitlab and switch to gitlab in case of outage;
- if github is used beyond its primary purpose(git hosting) put some efforts to maintain the same features in gitlab/gitea/sourcehut, i.e. use CI features of both platoforms and push releases and artifacts to both of them so that end users have the choice.
- separate concerns and not go all in with github as a platform for the whole SDLC and project management and instead use different tools/platforms for different purposes.
Yes, it takes more time for the initial setup but it quickly pays back.
The above applies to those cases where your products need to be hosted publicly.
This gets tougher with CI/CD though. How urgent is it to run your tests or deploy right now as opposed to when the outage is resolved later, or tomorrow? In emergencies I think you need ways to deploy without the CD pipeline (for those that can't currently).
In my experience there's always more work to do and downtime is a temporary inconvenience. Of course... my customers would be blowing up our phones and emails if we were down, threatening to quit, etc. I assume Github doesn't have quite the same level of pain when they have downtime.
...but if you said anything you were just interrupting people playing with their new toys.
The shrug whatever coffee break while another department fixes things is the same now as it was back then.
And the VMs GitHub provides are far too low spec to run Cypress in. There's no option to pay for more resources.
I was going nuts. For my own products I use gitea and host repos myself. Some clients of mine however use github so in this case I have no choice
Strategically cyber attacks are the most effective when they damage your adversary heavily, and yet at the same time make it so that politically retaliation looks like aggression since there is no unquestionable link between the actor and the action.
Consider Stuxnet as an example. Widely understood to be a joint US/Israeli cyber attack on Iran, and at the same time difficult to retaliate against.
While annoying, these outages we've been seeing are hardly doing any serious harm. Very little money is ultimately lost because of them, and it's not remotely clear that this is a cyber attack. So if they were cyber attacks it wouldn't make much sense. Not saying that it's impossible they are, but it seems there are many other more plausible explainations.
> Git Operations is experiencing degraded performance. We are continuing to investigate.
Post mortems - TOTALLY fine. But just linking to a specific incident is IMHO not valuable to HN and just takes unwarranted frontpage space.
-Sometimes for conversations about alternatives
-Sometimes for conversations about the product's stability history
-Sometimes just for solidarity in frustration
So then submit a post mortem and drive the discussion there. Or alternatively post a "Ask HN" if you're legitimately considering alternatives.
> -Sometimes for conversations about the product's stability history
Which are readily available on those provider's status pages.
> -Sometimes just for solidarity in frustration
Isn't that what Twitter is for? Why should the frontpage be littered with <insert provider here>
My point is this - let's say GH goes down 8x over 2 weeks for 10 minutes each. Should there be 8 posts on the frontpage about GH going down each time? No, there should be ONE, otherwise the valuable discussions you're referring to get lost.
That's what post mortems are for. Happy for people to submit those.
Also, why are cause and remediation efforts important to you? If you rely on the service and it goes out, what more are you going to learn here that would help you? Genuine question.
> I always check HN for major outage news when it comes to github/AWS/GCP/Azure
Why not just check those status pages? Or check Twitter? HN seems like a weird place to check for it...
(I recognize not in this case, but I certainly noticed this outage and checked HN before the status page was updated)
Because I have to provide information to others in my company about why X is not doing Y. These bits of information help.
> If you rely on the service and it goes out, what more are you going to learn here that would help you?
Roughly the same answer - though I'm also just actively curious what's going on as well in the moment.
>Why not just check those status pages?
I do. HN often has more information.
>Or check Twitter? HN seems like a weird place to check for it...
I don't want to.
I also see these as an excellent learning opportunity. Typically, it is a lesson in managing complexity.