Gitlab.com is down
status.gitlab.com
status.gitlab.com
All that I can find left online is this [1], which is still informative, but not nearly as interesting as I remember the chat transcript being
[1] https://about.gitlab.com/blog/2017/02/10/postmortem-of-datab...
I can't even imagine the sinking feeling..
Then in the post-mortem about lack of backups:
> LVM snapshots are by default only taken once every 24 hours. YP happened to run one manually about 6 hours prior to the outage > Regular backups seem to also only be taken once per 24 hours, though YP has not yet been able to figure out where they are stored. According to JN these don’t appear to be working, producing files only a few bytes in size.
I have had (and inevitability will have again) bad days like poor YP. All I can count on is to maintain good habits, like making backups before undergoing production work like YP did.
The specific part you mention also brings up a really vital part of a backup system, testing that the backups generated actually can restored.
I've seen so many companies with untested recovery procedures where most of the time they just state something like "Of course the built-in backup mechanism work, if it didn't, it wouldn't be much of a backup, would it? Haha" while never actually tried to recover from it.
Although, to be fair, I've only seen one time out of the untested 10s where it had an actual impact and the backups actually didn't work, but the morale hit that the company ended up having made my brain really remember the fact to test your backups.
Things were “mehhh” for developer interest for the first month or so, as a I was working off of an internal todo list. I didn’t think anyone cared about my 300 line long todo.txt file, but then I started to wonder if I should find a way to put that doc out in the open for developers to follow, and possibly jump in and contribute.
I had a hinkling that I could use GitHub issues to help with this, but I believed the title “issues” would hurt my project. I was under the impression that a new open source project with a single contributor, and a ton of open “issues” would look bad to developers.
I started to inquire with devs on IH and hear about using issues for feature tracking. Much to my surprise, I got an overwhelming “yes, you need to use issues”. I was also told not to worry about the misleading “issues” title, and that enough developers were knowledgeable enough to know they weren’t just bug reports.
As I started to open issues, and ask for help; surprisingly I started getting traffic and interest. The more issues I opened, and the more open I was online about my code and plans for it; the more followers and contributors I’ve gotten.
My plan at this point is to just follow Gitlabs model, and go full open transparency with everything.
My side project mentioned above can be downloaded here: https://github.com/elegantframework/elegant-cli
It's also pretty worrying that a single developer was doing work directly on production databases without a second person there to say "yeah, looks good". This is a big operational mistake, no matter how good you think you are.
Github by comparison has had more than 30 outages this year alone. https://www.githubstatus.com/
"Dev Deletes Entire Production Database, Chaos Ensues" https://youtu.be/tLdRBsuvVKc
> Locations: Google Compute Engine, Digital Ocean, Zendesk, AWS
that's some _impressive_ blast radius and I actually struggle to think what would cause that much damage
$ host cdn.artifacts.gitlab-static.net
Host cdn.artifacts.gitlab-static.net not found: 3(NXDOMAIN)
They must have used that same healthcheck for every single load-balancer or something, though, for it to nuke AWS, DO, and GCPMost of the entries look like batch job processing, either by gitlab or managing jobs in runners.
Perhaps authentication, which is mandatory when managing batch jobs including in third party nodes.
If auth goes down on any service, you'll see similar blast radius.
The incident review is updated in https://gitlab.com/gitlab-com/gl-infra/production/-/issues/1... I have added your feedback into https://gitlab.com/gitlab-com/gl-infra/production/-/issues/1...
(They start to ruin that, but lets say, the idea is still visible.)
And they also said they were going to remove that ability too. Tbf to them after the inevitable response they did say they will come up with a proper solution, though that was some time ago.
Obviously the real solution is to use Bazel and not to use Gitlab CI as a crap build system.
I spent a good few months learning it and it's not the tool I would reach for in almost any circumstance unfortunately.
Docs are also lacking, which is certainly not a problem with GitlabCI
It would be nice if it had better integration with existing package managers like pip, cargo, npm and so on though. Vendoring your dependencies is fine for a big project like Chrome or Android or your company monorepo; it's a bit annoying for a small CLI tool or whatever.
Another option is something like Landlock Make, but I haven't tried it.
Granted, those build instance pull from the central git repo, and are triggered from GitLab... so also, no.
If you plan on a GitLab mirror, including the Git repository and groups/projects, the direct transfer migration can be helpful https://about.gitlab.com/blog/2023/01/18/try-out-new-way-to-...
GitLab Geo can be an alternative for larger, high-availability setup requirements. https://docs.gitlab.com/ee/administration/reference_architec...
http://Xkcd.com/303 but for GitHub is down.
For example on my team we have a few devs who push patches from laptop directly to their work desktop, as a part of day-to-day synchronization. No central ssh server involved.
Incidents can have different types, i.e. when an application bug or performance regression is discovered, this can involve reverting MRs and rolling back releases. The Platform, Delivery group has a top-level responsibility for ensuring continuous delivery of the GitLab application software to GitLab SaaS, https://about.gitlab.com/handbook/engineering/infrastructure...
Other incidents may involve hardware or infrastructure failures, or a combination of both, infrastructure failure that renders GitLab application services unavailable. This requires cross-functional collaboration from infrastructure, product, engineering, etc. teams in the incident.
To get a better understanding here, it is helpful to review the incident management handbook https://about.gitlab.com/handbook/engineering/infrastructure...
Additional helpful information:
- The GitLab.com SaaS production architecture is documented in https://about.gitlab.com/handbook/engineering/infrastructure...
- The Monitoring of GitLab.com handbook provides insights into monitoring workflows, incident management, SLAs, etc. https://about.gitlab.com/handbook/engineering/monitoring/
- Runbooks https://about.gitlab.com/handbook/engineering/infrastructure...
For the current incident discussed in this HN thread, the review issue can be followed in https://gitlab.com/gitlab-com/gl-infra/production/-/issues/1... to learn more.
But even for just version control for most use cases you need some common remote for people to push and pull to. Even a bare git setup has the potential of going down and putting people in the exact same boat that they are in now (they can commit locally, but can't push).
All these startups who experience a 100% dev outage when github or gitlab go down are fundamentally broken.
On this I totally agree. Although if you have deploys fully automated and gated behind CI/CD checks (which in many cases is required for compliance), how do you structure things such that a total outage of the CI/CD platform doesn't knock the dev cycle?
Fetching changes...
Initialized empty Git repository in /builds/jjg/cptutils/.git/
Created fresh repository.
fatal: unable to access 'https://git-us-east1-d.ci gateway.int.gprd.gitlab.net:8989/jjg/cptutils.git/':
Failed to connect to git-us-east1-d.ci-gateway.int.gprd.gitlab.net
port 8989 after 132297 ms: Couldn't connect to server remote: ERROR: The git server, Gitaly, is not available at this time.
What a drag... I guess I'll just commit locally... bah! Friday anyway!https://github.com/tylerjgarland/git2git
lol. I used it to migrate my gitlab to github on the last outage.
...and more importantly who gets to work so that there isn't a failure.
Next to all-in-one packages provided with the Omnibus package on Linux https://docs.gitlab.com/omnibus/ and in containers https://docs.gitlab.com/ee/install/docker.html you can also install GitLab into Kubernetes clusters using the Operator https://docs.gitlab.com/operator/ or Helm Chart. https://docs.gitlab.com/charts/
Most of GitLab.com SaaS is deployed on Kubernetes using the GitLab cloud-native Helm chart (with a few exceptions). More details in https://about.gitlab.com/handbook/engineering/infrastructure...
It's way more fun to debug than some single-monster code base with a complete stack-trace from the single-process. That basically tells you exactly where to look and how it got there! What fun is that, where's the adventure!?