Incident with GitHub Packages and GitHub Pages
githubstatus.com
githubstatus.com
I’ve used GitHub for 10+ years and don’t remember these outages historically. I’m sure there’s more usage now, but usage has been high for a long time.
My initial guess is Microsoft monkeying with the tech stack to use inferior solutions from their own stack rather than proper architecture design and evaluation.
I at least have found the recent improvements and additions pretty great — GitHub’s newish PR review UX is way ahead of BitBucket (where my employer currently hosts our codebases) for example.
- In Bitbucket, you can reply to _any_ comment as a thread, rather than a quote-reply, making conversations easier to follow. Github only lets you do this if you're replying to a comment on a file.
- Though this has been added very recently as a beta addition, the lack of a file tree made large reviews in our monorepo hard to follow.
- If someone force-pushes, the UI doesn't tell me what commits were removed; just the new branch-head.
- Our devops team has had to make it so we can re-trigger our CI (or add additional pipelines) via comment-commands, while in BitBucket we were able to add custom buttons to the PR menu. This means there's a lot of noise of comments that just say "/retrigger-ci" or some such.
- This one is not Github's fault per-say, but because we used to add our users via LDAP, we used to just be able to type <first-initial + last name> to add a user to a PR, but now because everyone has bespoke usernames, it's a lot more annoying to add people because I now have to also remember their username.
- In our repo, we work with lots of release branches for maintenance (committing to `develop` is not rare, but quite frequently we're adding features to old branches), and Github's PR view does not show the target branch at a glance, which can be annoying if someone has multiple PRs to cherry-pick a commit across branches and you don't know which PR you're clicking on.
- Being a cloud solution, we can't add pre-receive hooks. We previously used to reject commits at push-time that didn't follow the format we use for time/issue tracking. Now, we have to enforce this in CI which means rebasing/force pushing to fix it. While pre-commit hooks can help, this depends on every developer keeping their hooks up to date, which is a challenge.
That said, Github's "review" system, rather than limiting you to only individual comments, is great for email noise.
Also you’re not doing trunk? I mean I get it it’s hard but sounds like a no brainer thing to stop doing as soon as possible..
The PR workflow certainly isn't, but we're far from the only company using them (Dropbox comes to mind). It's surprising to see it not fully supported.
> Also you’re not doing trunk? I mean I get it it’s hard but sounds like a no brainer thing to stop doing as soon as possible..
No, unfortunately not. Trust me, you're preaching to the choir :) But unfortunately, long-lived branches is the workflow we're stuck with at the moment (we constantly merge commits forward, but it's a manual process). That said, even if we were using short-lived feature-branches, you wouldn't be able to know from the PR view if a commit was to a feature branch or not, so the criticism still applies.
Job re-runs are more than one click away unfortunately. But I'd be curious, outside a failing job due to flakiness or GitHub instability; for what other reason do you use this functionality?
Anyway, hope that answers your question :)
Lots of talk about moving stuff on-prem.
That can be complicated in itself, but GitLab's operational overhead is minimal ( both their omnibus and Helm deployment methods are highly automated) and IMHO it would be hard to have availability lower than Github's is at the moment ( last few months).
Of course we did make sure we can still deploy our code if GitHub was down but that sounds like a good redundancy practice anyway so why not?
Like I said, I expect them to go down once a month guaranteed. But this is like 6 times in one month. Now I expect something in GitHub to go down twice in one month.
Not looking good at 'going all in' and 'centralizing everything to GitHub' as I predicted [0].
Centralising on this infra was a stupid technical decision.
GitHub is still a no brainer for me.
GitHub downtime is something I can just tell managment and prepare things to put it out of high critical path.
Like I would not pull images from GitHub on my k8s cluster.
Gitlab works absolutely but you know I don't want to do gitlab updates, do backup, do regular backup/restore test, deleting backups etc.
I prefer to create new things.
Their releases often have security vulnerabilities though. But other than that, even the community edition is packed with features, it's a no brainer for me ;)
You also need to do backups, try restore, you might need to announce a maintenance window.
And for what? Getting stress from mgmnt when it breaks? With a GitHub downtime that's just it. No one cares, everyone accepts it as given.