Incident with GitHub Actions and Codespaces
githubstatus.com
githubstatus.com
Dedicated servers are just so cheap, esp with sth like CI/CD that's perfect to host at cheaper providers.
Seeing constant issues with Gitlab and Github products just reinforces that decision.
Upgrade migrations didn’t always succeed so then you need to go back and remember what you did last time to fix it.
It’s a cognitive burden and distraction. In theory it’s cheaper to pay Gitlab to handle for us. Except that now we lose more time due to Gitlab.com outages than we spent maintaining on prem.
There's so many companies out there whose infrastructure is run by this One Guy. Any company needs to assess whether they can risk losing their One Guy (see also: Bus Factor), or if they'll just have to swallow the cost and (some) risk of (some) downtime by using a service provider.
There's operational risk, but that doesn't go away if you shift to Gitlab.com infra. Just in this case, you can control all of it.
The problem with Github and Gitlab is that it's actually turning out that those companies aren't that good at running their own product, which is tilting the economics of the above.
Keep in mind that Gitlab.com sees much more load & activity than your on-prem equivalent, and their maintenance schedules don't line up with yours. Ideally, you don't want someone messing with your VCS/CI "every day".
Your on-prem is not only very unlikely to need any of the moving parts that a world-scalable system needs, but any maintenance can also be planned out-of-hours to make sure eventual operator errors don't cause downtime during business hours.
By this measure, a lot of companies have zero decent developers.
Every company has their own way of doing cloud-based infra, CI build pipelines (in their favorite CI provider's dialect, since those aren't portable either), etc.
Every time I start at a new place I have to reverse-engineer what the previous guys did; there is no standard. By that benchmark, one extra thing to document/keep in mind isn't that big of a deal, especially considering GitLab (installed using their standard "omnibus" install on one of their supported distros) is likely to be more standard than all those custom Terraform/CI pipelines/etc.
The SaaS will have more moving parts & operator activity which in turn makes outages more likely, and that "operator activity" can't be planned for unlike on-prem where you can schedule maintenance out-of-hours to make sure any accidents don't affect the business.
That's how you do it (and also how many open-source organizations have done it) as I have said and predicted years ago here going against 'centralizing everything to GitHub' [0]
Fast forward precisely 2 years ago since that prediction was made in (April 14th, 2020), GitHub has become (unsurprisingly) unreliable, especially for those that went 'all in'. Self-hosting for your own organization seems to have been more reliable in the long run.
> Seeing constant issues with Gitlab and Github products just reinforces that decision.
This also. Especially the SaSS ones.
The previous F500 company I worked at had enough engineers to self-host their things but thought SaaS was also the future and migrated a lot of their key tools and services onto these cloud platforms.
I wonder if both my start-up and the F500 company will suffer in the long run because of these decisions...
Most of the overhead is managing users, permissions etc, that won't go away with SaaS.
(Also, when upgrading GitLab, make sure to upgrade version-by-version and always wait for background migrations to finish before upgrading to the next version. You can check background migrations at /admin/background_migrations.)
To be fair though I'm using docker and watchtowerr to keep it updated, there was 1 time where I had to use an older docker image because the latest wasn't working.
Not unrecoverable though, so saying you need to worry about it may be a bit too alarmist.
Upgrades and backups are way more problematic. A while ago I discovered that backups broke and I had to fix it. I setup automatic APT upgrades for Gitlab, but every few major versions or so a manual upgrade process is required. In an organization where people are less vigilant, and/or haven't setup monitoring of backups, this can easily lead to a disaster down the road.
None of this is particularly hard — for me. But for some others this may as well be black magic. Furthermore all this work is a nuisance, a distraction from what I usually do.
We said the same. Now we lose twice as much time due to Gitlab.com outages
Please, please don't ever do this. Or if you do, make sure you're taking a snapshot of the box every hour or something..
A few days later, drop the snapshots.
Besides, we already make weekly automatic backups, stored remotely and encrypted. And Gitlab automatically makes a locally-stored backup before each upgrade, which is good enough unless the disk fails right after we discover that the upgrade is broken.
Upgrades aren't automated but as the server is only exposed to the internal network (and also not fully exposed there), the attack surface is much lower. Keeping servers up to date isn't hard, and we don't always insist on running the newest Gitlab version. We make sure to upgrade if there's a very critical bug or if we lag too far behind.
It's true, upgrades for Gitlab could be smoother, but we've never had significant down time from it. And while others might say it destracts from day to day work, I (and some others) find it a welcome break from coding.
Taking a actual break works too, I think. I mean, it is a virtue to see good in everything ..
GP's approach does have a bigger learning curve for those that have never worked that way, but after that the ongoing maintenance is simpler, less time consuming and much lower risk. That probably explains most of the difference in your experiences.
Anyone with some knowledge of Linux/Programming shouldn't struggle setting this up within a few days. And it adds knowledge to the company that's very valuable.
With Gitlab, as with most software, restarting it regularly outside of office hours is the best way to ensure a stable service. Also something SaaS can't do, as they never know when their customers need the software.
Ideally you want to have multiple runners, and for them to scale so that developers aren't endlessly waiting in job queues. On top of that you might want proper network segregation, and for it to be on a separate virtual network to your workloads to reduce your blast radius. You might have problems with connecting your server to third party SaaS outside your network, etc etc etc
Virtual network is also very easy to set up with most providers (even with dedicated servers), there's pretty much zero config.
So far we didn't have any issues connecting to outside SaaS, but we host most stuff ourselves anyway, so that's not a common use case.
Why would decisionmakers want to add outsource-able ops overhead to that very full plate?
Before we were losing productivity for an entire day due to GitHub outages, now we only notice because there's a post on HN :)
As for the cost of self-hosting, we already host our own services, Gitea is no more complex than any of those. I wouldn't recommend self-hosting to a company with no backend/ops capability in-house, but for any other company it's fairly obviously a better choice IMO.
We still have a $7m dollar balance on our account, so the issue clearly is not resolved.
Providing incorrect information on status pages makes it hard to trust in the future.
2 hours later they marked the incident as resolved, but they clearly have not completed the "cleanup" yet.
status page :
Resolved This incident has been resolved. Posted 7 hours ago. Apr 14, 2022 - 01:28 UTC
Update Actions and Codespaces are operating normally except for incorrect billing reports. We have started a cleanup of incorrect billing information and expect that to be completed in a few hours. Posted 9 hours ago. Apr 13, 2022 - 23:36 UTC
Update We've applied a mitigation to unblock running Actions jobs and using Codespaces. We are currently working on addressing the inconsistencies in billing information. Posted 11 hours ago. Apr 13, 2022 - 21:07 UTC
Anyway, it's already fixed for our account.
The risk of surprise bills, even with the advantage of higher scalability, is one of the major factors to put me off many cloud services. I wouldn't want to run a business where you always have the risk of some bug or unforeseen behavior destroying your budget.
It would be a pro-innovation law because people would then have the confidence to innovate.
In my case my cards don't have $12 million of a limit so they could not actually charge me that amount of money. It makes sense if you use invoice type billing i.e. like a company but as an individual these insane bills can't actually run.
they can still attempt to come after you for that $12 million
Well, when you owe $10,000 it's your problem, when you owe $10,000,000 it's their problem :D
That's why I use a cloud provider that doesn't charge for bandwidth usage and I use a fixed set of "stuff" such that I always know what my fee's should be.
I understand for big companies with millions of users they can't operate in that way but for me risk vs reward is too high.
They absolutely can. If anything, big companies have more resources to build their own cloud based on on-prem or rented bare-metal than smaller companies. It's just all cargo-culting.
The "cloud" is not some kind of unobtanium hardware you absolutely can't get anywhere. With rare exceptions, it's standard hardware you can buy directly from any server & network equipment manufacturer.
Wait until we have self driving cars that are somehow cloud dependent and in the middle of rush hour AWS or GCP decides to burp.
Our usage report said that the egregious minutes were run on deleted repositories so we were very concerned.
These outages and incidents are getting very tiring, and costly.
Not looking good for going 'all in' on GitHub products or 'centralizing everything' to GitHub as I predicted. [1]
I expect them to go down at least once every month guaranteed. Not 5 times in two weeks. I don't think they can handle a month without a serious incident at all.
In our case first sign was Jenkins not kicking off builds on push.
We've had to purchase additional data packs to unblock CI builds and folks pulling from the remote and assume we will need to purchase more. 11 days left in our billing cycle...