If anything, I hope that Microsoft's acquisiton of GitHub means that GitHub is going to keep growing in features for varied enterprise uses, and that we're going to see even more competition in this area.
If anything, I hope that Microsoft's acquisiton of GitHub means that GitHub is going to keep growing in features for varied enterprise uses, and that we're going to see even more competition in this area.
However, I re-valuated and did migrate about 2 years ago and it has been fine during that time. There have been a few hiccups, but not for more than an hour or so. I've had a team of 4-7 devs working in it all day for the last two years and we have not had performance problems. We run our own CI runners as well, and while the cloud runners do often have delays, I've never had issues with delays to my own runners unless they were all busy.
So while they still have improvements to make it would be a lie to say they haven’t improved at all.
Also, I don't think GitLab has had a long downtime recently. At least not for any of my projects.
That mostly depends on whether you're using CI/CD I'd think, that's had some day-long outages/problems lately. Of course, GitHub doesn't even have it's own CI/CD, and GitLab's is amazingly flexible, so it's still the better product. But it'd be nice if it were more stable.
(Note: all this is on GitLab.com. If you self-host, it's presumably much better.)
(I am using the free tier though, so this is more informative than that I'm complaining.)
Other comments further down are showing other's are too. Hacker News Hug of death?
It's not a great look.
Thanks to Comcast for creating Trickster.
Hope my post didn't come across as snarky as some others have... HN are like the Spanish Inquisition. No one expects.
But you know, when the internet decides it's time for everyone to look at your site, some random new stuff might be better than serving 5xx all day. :-D
HTTP 512 - Social Media induced DDOS due to media related frenzy :)
Good luck!
- I can't push or something in general goes wrong with one of my repos (but not others).
- Gitlab's status page is green
- Other people are having issues on Twitter and tweeting @gitlabstatus about it but there is not general across-the-board outage
This seems to indicate that Gitlab tolerates (and very often has) a reasonable amount of instability and error rates across its platform, but just takes the average of these as a baseline of performance: i.e. it's a very spikey graph with a reasonably high average line fit.
This tweet supports this impression:
https://twitter.com/gitlabstatus/status/1000001988183158785
"Errors should be down to normal" - the idea that there is an non-zero error rate that is openly described as "normal" is worrying. Not that I'd expect a constant zero error rate, but at least aiming for it should be a consideration.
Services at this scale will have errors for all sorts of strange reasons, it doesn't mean the service is poorly engineered. In fact, if users don't notice these problems it usually means the service is resilient and robust when it encounters strange situations.
Consider a really simply example such as making a breaking API change to your service API. Now what happens when a user doesn't refresh their web browser and continues running javascript that doesn't work against new API. This can happen with smaller services but the odds of this happening are much higher when you are a global scale.
There are other strange problems that come with large services which means all components should be fault tolerant if possible.
Also, please don’t make disparaging comments about other people’s experience unless it’s highky relevant. It doesn’t add anything to the conversation and will likely derail the conversation.
As per the really simple example: generally you'd be better off rolling out a second endpoint for the new api and then stop serving responses that use the old one. First this doesn't break everyone who had your page up, and second you can stop rollout safely if you find a problem with the new api.
Of course, and as I said, zero errors is not a practicably achievable in this type of context. The issue is with metrics though: the idea of taking averages instead of looking at troughs is still problematic.
> In fact, if users don't notice these problems it usually means the service is resilient and robust when it encounters strange situations.
True. But in the case of Gitlab, users are noticing these problems. Constantly. It's just Gitlab's own metrics that could be (I've not done more than browsed their Grafana instance a bit, so my comment is generally a bit speculative) ignoring the problems because they're focused on averages instead of specifics or thresholds.
> Consider a really simply example ...
lallysingh has already pointed this out, but I'll reiterate that this is a very apt bad example. You're right that ideally components should be fault tolerant if possible, but frankly that's a big ask. Especially for highly-scaled services supporting many many components of various types - ensuring that all of those components are completely fault tolerant is much more difficult than simply ensuring the old API continues to operate for a grace period while the new one is served from elsewhere.
I think your example is apt, because it's indicative of a common excuse for bad engineering: the assumption that downtime or disruption is necessary because of necessary software upgrades/improvements and poorly planned orchestration.
I love GitLab and it’s UI, but recently the performance of the hosted version is awful (not sure why - just being overloaded?).
In fact, even their own status page reflects it: https://status.gitlab.com/ - the current “project HTTP response time” is around one second which makes me cry when using the UI.
I wish them the best but would be moving to a competitor (or maybe a self-hosted GitLab) in the meantime until they sort it out.
I find it boggling that a commercial team chooses to accept this kind of external dependency. What do they offer which makes it worth the extra risk?
Then again, I come from a largely non-web background where external dependencies aren't just accepted as inescapable. I guess if your entire business is producing an add-on for some other company's web service (not saying yours is but many out there seem to be) then what's one more on the pile?
That’s a myth: https://en.wikipedia.org/wiki/Boiling_frog
This link is for the monitoring page. The imports are going through.
I ask because it's not a particularly good luck that the viewers from HN are capable of bad gateway hug of deaths to the site?
I do hope they're using Prometheus federation to expose this instance to the fickle internet and that they have one or more internal Prometheus instances that aren't directly queried by this instance. After all, that stuff is responsible for paging if something goes wrong in prod.
We used to use Federation, but now we just have the public server scrape the same targets as our private one.
I'm adding a caching proxy (https://github.com/Comcast/trickster) to the public server now to improve performance. :-)
Not to mention, it’s visually customizable.
https://docs.gitlab.com/ee/administration/operations/sidekiq...
https://github.com/gitlabhq/gitlabhq/blob/master/lib/gitlab/...
https://github.com/gitlabhq/gitlabhq/blob/master/config/init...
Many other reasons to be concerned about performance, but there's no evidence that they're withholding essential features like this from their free version.
Reminds me of the classic "The main Rails application that DHH created required restarting ~400 times/day. That’s a production application that can’t stay up for more than 4 minutes on average".