Gitlab.com was experiencing elevated error rates for Git, Web, and API
status.gitlab.com
status.gitlab.com
There are a set of high frequency (very frequently run queries) that are quite sensible to plan flipping (they normally run with a given execution plan; if the plan changes to a worse one, effects can be dramatic, given their frequency).
That leads to a lot of query timeouts (GitLab's database limit query execution time to prevent further damage), which are visualized as database errors. This in turn leads to partial or complete downtime (effectively, if the database is running few effective --non timed out-- traffic).
Some of these queries depend on a particular Postgres planner way of working that is less than ideal (in PG11, it is improved in PG12), and may lead to plan flipping. Which in turn, happens when table statistics degrade. In GitLab's case, they are normally close to the threshold, and slight statistics degradation caused a plan change, which in turn lead to many queries time out and excessive load on the fleet.
At the end of the day, the fix is quite simple, however: update the table statistics (running ANALYZE). For more information: https://gitlab.com/gitlab-com/gl-infra/production/-/issues/3...
Would this be a place where a plan hint would help? It seems like having to re-analyze would mean that this is just a matter of time before it becomes problematic again.
Thanks :)
> Would this be a place where a plan hint would help?
It is one option to consider (I wrote about it last week: https://gitlab.com/gitlab-com/gl-infra/production/-/issues/3...), but it is not an ideal solution, for several reasons.
Re-analyze is not the solution either, but as a short-term measure, cron-ed ANALYZEs will be run. Longer term, apart from refactoring some queries which are quite prone to trigger this plan behavior, statistics gathering process is going to be fully reviewed.
time to add a new backup remote to all our repos.
Australia has very weird and wide reaching laws regarding software backdoors.
Edit: Nvm on (b); somehow I got GitLab and BitBucket mixed up as to which one belonged to Atlassian.
Still curious about what sort of backdoor laws they have.
EDIT: I've done some more reading up on this, and The Assistance and Access Act of 2018 explicitly states that government cannot use backdoors: https://www.homeaffairs.gov.au/about-us/our-portfolios/natio...
https://about.gitlab.com/devops-tools/bitbucket-vs-gitlab/ ( marketing page, but mostly valid points )
And it seems to be related to scaling because it's always iffy problems like higher error rates or jobs not executing.
Either way, I love our on-prem Gitlab. It's free and it hosts over 200 projects, Gitops and the whole shebang. I'm just now beginning to use Kubernetes from it.
They will sort this out eventually.
https://about.gitlab.com/handbook/marketing/community-relati...
<td><a href="https://twitter.com/GitLabs">@GitLab</a></td>Please note that about.gitlab.com contains all our static content, including our 10,000 page handbook. This typo isn’t on our main about page but deep in the marketing handbook.
I merely wanted to verify the @ was authentic, and that page was the top "site:gitlab.com" result.
The page has formatting problems, I'll rework this in a separate MR.
Gitlab has been a fully remote proponent going back a few years and these persistent reliability issues add some doubt to that model - even if there might be altogether other issues in play.
My company is full remote and we don't have these issue.
Pushing commits often hangs completely, and we've had a number of smaller downtimes/degradations in recent weeks that have been super disruptive.
Having been Gitlab users since our inception, we're now seriously considering Github (I know - grass is always greener).
You can tag me on an issue on gitlab.com with @robotmay_gitlab if the details can be public, or email them to me at rmay@gitlab.com if not :)
That is correct - we also will be posting updates here: https://twitter.com/gitlabstatus/.
1: https://status.gitlab.com/pages/history/5b36dc6502d06804c083...
2: https://www.githubstatus.com/history (don't get fooled by them collapsing the incident list after three per month)
2: https://bitbucket.status.atlassian.com/history (also collapses after three per month)
At least with GitLab, you can set up a self-hosted solution for free as a backup, unlike GitHub (Unless you want to pay a lot for GH Enterprise) where some have 'gone all in on GitHub Actions' and then some couldn't push that critical change [2] before the start of the weekend. Oh dear.
Maybe its time to setup a self-hosted backup VCS and not depend entirely on GitHub/GitLab web.
[0] https://news.ycombinator.com/item?id=26439075
I've been wanting to do this for some time purely for git hosting (don't care about CI at all) but wasn't sure about how to make it failover...
Personally i'd be happy with a headless git server, but that's not fair on everyone else who wants the GUI to browse and organise stuff, so I want gitlab etc to deal with that. What would be nice is to have a headless backup server that allowed everyone to continue pulling/pushing from the CLI with their existing repos when gitlab is down. I can't see a smooth way of doing that since it would require messing with the git remotes, unless the solution is inverted and uses a single remote pointing at the "backup" server which then replicates to gitlab, but I don't think gitlab can be configured to change the default remote when people clone it from the UI.
I suppose this is why people just end up self hosting gitlab instead.
Gitlab has a repo mirroring feature [1]. But of course you'd also need to sync users public keys.
Downtimes are so infrequent and relatively short that this isn't worth the effort for me.
You can instead set up a read-only mirror so people can at least still pull and browse the code.. Gitea [2] might be a better choice than Gitlab, since it's much more lightweight and easier to host.
[1] https://docs.gitlab.com/ce/user/project/repository/repositor...
[2] https://gitea.io
> Downtimes are so infrequent and relatively short that this isn't worth the effort for me.
That's essentially the same conclusion I keep coming too, occasionally it has hit me when I go to push something but rarely has it blocked me or anyone else from continuing to work.
> You can instead set up a read-only mirror so people can at least still pull and browse the code
Yeah, this I need to do eventually just for peace of mind as a more automated backup solution. At least I don't have to care about failover.
I think the flakiness of jobs succeeding is also related whether you use their package registry. I can't say I had similar issues with Github.
I think it mostly comes down to "I didn't write it, so it's bad" :)
Though from the context it's 2.
"A fortnight is a unit of time equal to 14 days (2 weeks). The word derives from the Old English term fēowertyne niht, meaning "fourteen nights".
We actually had a programming interview question using biweekly at a previous job. It was about thorough requirements definition and understanding/agreeing to measurements.