Scaling the GitLab database
about.gitlab.com
about.gitlab.com
Nevertheless, the slowness is really annoying, especially because their product is so good on all other accounts. If scaling their database can help speed things up, I bet they will be glad to remove this embarrassing "missing feature".
In marketing terms, having fast page load would be called a "qualifier". For example: you expect a hotel to provide toilet paper. You won't pick any hotel because of it, but you will surely avoid one that doesn't.
wow that a lot higher that i would have expected. Linked bitrise page used " randomly selected 10,000 apps as a base, that are getting built regularly on Bitrise" , not "enterprises".
How did you comeup with 2/3 number? First line of your about page is a lie?
We got work to do in the 99% and merge request page load but the overall situation has improve dramatically.
We still got work to do in availability, so I changed the 'feature' to reflect this https://gitlab.com/gitlab-com/www-gitlab-com/commit/3b3bddf5...
We just kicked off a major effort to address the availability issue Sid mentions above. The highlights are that we're moving to GCP which should provide better underlying reliability. But interestingly we found that only about ~20% of our downtime minutes where from underlying infrastructure. Whereas ~70% came from features that didn't scale.
So the more exciting part of the project is to tighten the feedback loop between development and deployment with a continuous delivery pipeline. This may be obvious to some people, but it's harder to pull off when you've got an open source project, an on-prem product, and a large-scale SaaS sharing the same code base. I'm calling it "Open-core SaaS" and there are only a handful of companies that run a large, multi-tenant service based on an open source project.
The one on GitLab is considerably longer.
"Table stakes" is a common term for this.
By sharding by customer we're able to more effectively leverage Postgres caching and elastically scale. Since switching over, our database has performed and scaled much better than when we were on a single node.
There is some up front migration work. But that's limited some minor patching of activeRecord ORM, and changes to migration scripts. We considered casandra, elastic search, as well as dynamoDB, but the amount of migration changes in the application logic would take an unacceptable amount of time.
We've been working to make that upfront work easier as well with libraries that allow things to be more drop-in (ActiveRecord-multi-tenant: https://github.com/citusdata/activerecord-multi-tenant and Django-multitenant: https://github.com/citusdata/django-multitenant)
* Edit. I AM excited for this. Typo'd now = not previously
Edit: Thanks for the clarification, the typo part in particular through me off, makes sense now.
Secondaries cover this case for Gitlab, it seems, but that comes with a set of caveats as well (namely async availability of data).
Another caveat is that all secondaries will end up having more or less the same stuff in their cache. As your data set grows bigger that becomes unsustainable, because you're going to read from disk more and more.
When you shard across N nodes you can keep N times as much data in the cache. Combined with N times higher I/O and compute capacity, that can actually give you way more than N times higher read throughput than a single node (for data sets that don't fit in memory), and you can get much higher write throughput as well.
Is easy to overlook how much you can do with a decent RDBMS. And for the small data that gitlab use, I believe still exist a lot of big wins on performance.
Is just that the new generation not pay much attention to Sql databases...
P.D: I don't mean the gitlab developers, just on general
Edit: I would also add that most projects I've seen tend to undergo one or more large refactors / redesigns before growing to a size where complex DB scaling is needed. This is, of course, speaking from the perspective of a small startup or pet project. If you're writing a new feature for an existing product where you can assume that it'll have lots of users, then you would obviously build scaling right from the start.
And the cost of building the solution from the onset is hardly a killer. The pain usually comes from operational complexity.
You never actually know where your bottlenecks are going to be until they arrive and by designing everything with a "super scalable" architecture you will be making development ten times as painful and expensive as it needs to be while throwing away nice things that come "for free" and "just work" at the mid-low end like transactions.
Amazon and Netflix don't want to have to use their hideously complex and inefficient service architectures - they're forced to because of their scale.
Most people who engineer their systems for hyper-scale from the get-go never see a whiff of anything that looks remotely like high traffic. Often they go out of business before they get anywhere near that.
And once you get to serious scale, you really shouldn't still be running the code from back when you didn't really know what your business/product was.
2. Run into inevitable performance problems
3. Decide that your database is at fault, rewrite in NoSQL
4. Buy into NoSQL hype, push 5x resources on your NoSQL solution
5. Profit?
There are NoSQL options that are great for lots of things. It takes some significant expertise to know if your thing is one of those things. Expertise you probably don't have if you don't even have a moderately deep knowledge of RDBMS-es.
> Buy into NoSQL hype, push 5x resources on your NoSQL solution
NoSQL is so last year, NewSQL is the future!The software tries to do too much. CI/wiki/code review/git/container repository/analytics package.
CI sharding per team and carving up gitlab into departments and groups has been my only solution so far, and its just distributed risk at that. its the difference between the unix ethos of do one thing and do it well, versus the F35 approach of do everything forever.
CI/CD that is very easy to use. I can self host it - so no $600 a month contract like some of the other guys. Decent Issues that are not as heavyweight as Jira. Place for code / CI/CD / Issues to all be in one place.
> and the glaring problem of "ruby doesnt scale well" is everpresent
GitLab has many performance problems (and many have been solved over the years), but Ruby has thus far not been one of them.For anyone needing a small Git Web UI, I suggest https://github.com/go-gitea/gitea.
EDIT: Oh, you're that guy ^^;.
Anyway, my rant was probably a bit harsh. I know that GitLab isn't really written for hardware comparable to a Raspberry Pi and that it's used by small and large teams all over the world. But on my hardware Gitea renders pages in 20-100 ms, which is orders of magnitude faster than GitLab. I'm sure part of this is the fault of the language.
But be aware, gitea (like gogs) does NOT cache SQL queries, so you are heavily limited by that. I can’t get it to serve > 200 pageloads per second on my system, while even a normal Grails projects manages to serve 3000 (with more complicated queries).
It used to cache something, either SQL queries or rendered pages, but it might have stopped in the meanwhile.
You made me curious, so I just tested it: 102 rps on a project page, 365 rps on a user profile, as reported by wrk2.
Gitlab runs fine out of the box with 4gb memory and when screwing with some provided knobs it's fine with 2gb or less memory.
It's in the docs, it's in every FAQ, every Stackoverflow answer... so why the hate?
64mb memory? Use gitolite. 256mb memory? Use cgit + gitolite 1gb memory? use gitea or whatever else there is. 4gb memory? use gitlab if you want.
I'm not affiliated but I can't stand the "gitlab doesn't run on my home wifi router - it sucks" posts here... it works pretty well if you read & acknowledge the documentation and requirements.
My box has 16 GB of memory. The 128 MB value was the default value configured in unicorn. If you're not familiar with that (I wasn't), it has a feature to restart the web server workers when they end up using too much memory. That's what unicorn does and apparently it's what people do in Ruby web apps.
I've also used GitLab on a (real) server. It still feels too slow for my taste (pages load in 1-2s). No, it's not unusable, but I won't call that "fine".
I haven't gotten it to run reliably with any less than 8Gb, even though we don't have a whole lot of activity there.
It takes surprisingly long to start up too, even on a relatively beefy host.
I do like the product though, the all-in-one solution with code hosting, issues, code review (needs work though...) and CI is great. (Not using the deployment, orchestration stuff).
My biggest issue is that they seem to be too focused on pumping out new features.
There are 9k open issues just for CE. That's probably 12k+ for all products (runner, EE,...).
That either is a sign of very bad project management or a way to big agenda. Probably both.
The open issues for CE https://gitlab.com/gitlab-org/gitlab-ce/issues do include feature requests.
It's been very pain-less in setup and maintainance. Upgrades were mostly painless as well until now.
With 8GB for the docker container it's solid and pretty stable, but any less and behaviour would be very unstable (timeouts, very long request load times, ...).
Of course those 8GB are for the whole stack (Postgres, Redis, Nginx, Unicorn instances, ...), so I'll admit that 8GB for the Docker container is quite important context here. ;)
"Fault" is a pretty loaded word here. Pretty sure my old TI-86 also would struggle to run GitLab...
Both the "everything all in one" approach, and their insistence on "omnibus" packages that bundle fucking everything into a giant deb/rpm.
Its unfortunate to me that there isn't a simple (as in cgit/gitweb/hgweb) oss repo web viewer that handles multiple repo types.
I want issues, ci, project management separate.
If you're new to running your own postgres databases you should also check out Wall-e: https://github.com/wal-e/wal-e And the awesome pg stat statements https://www.postgresql.org/docs/10/static/pgstatstatements.h...
But it leaves me wondering: is there any sort of packaging of Postgres and these “friends” into a single opinionated virtual appliance, such that I could just stick up a couple of VMs with the same image with different role tags, and get a good cluster with automatic transparent N:M proxying, automatic transparent backup-and-restore, etc?
In other words, is there a product that is to Postgres as Gitlab CE is to Github (or like Dokku is to Heroku): a “host it yourself, but easily, and starting from a scale of 1 with no overhead” equivalent of a SaaS service?
I'm pretty sure there's some ansible/puppet/chef/saltstack code out there to build something similar as well.
[1] https://github.com/wal-g/wal-g
[2] https://www.citusdata.com/blog/2017/08/18/introducing-wal-g-...
- CockroachDB 1.1
- AWS RDS Aurora PostgreSQL-compatibility
[1] https://news.ycombinator.com/item?id=15458900 [2] https://news.ycombinator.com/item?id=13072861
Edit: Apparently it is GA as of 6 days ago, but still only available in 4 regions https://aws.amazon.com/blogs/aws/now-available-amazon-aurora...
Bonus points that rm -rf on a Cassandra server isn’t a problem, which seems like it could have been useful for gitlab.
The pipeline pages takes 7 seconds to load. It's really bad. We stopped using issues because just listing them was a chore. Pushing 1 file takes at least 30 seconds to 1 minute.
Anyways, the work falls on me to replace it. We're going with Phabricator and Jenkins.
- Self hosted Gitlab CE
- Gitlab runners in pre-emptibles nodes with autoscaling
Future (almost done, still in testing phase):
- Phabricator (for git, code review and issues)
- Jenkins ("lightweight" CI, more details at the bottom)
- Google Cloud Container Builder ("Heavyweight" CI)
Our main repository is in Phabricator, but it has mirroring to Google Cloud Repository (for backup and faster heavyweight builds)
On commits, we run the lightweight CI in Jenkins. What I mean by lightweight is that, there's only a few runners and what it does is check what has changed. We have a mono-repository with over 30 microservices, rebuilding all of them is a waste of time and resources, so we have scripts and easy to use manifests that defines what to do. We also have a lot of small CI checks that is just not worth it to run externally, like linting.
Here's what our homegrown manifests looks like
build_graphql:
when_changed:
- graphql
- node_libraries/player-graphdb
- node_libraries/player-auth
- node_libraries/player-models
script_build: graphql/cloudbuild.yaml
script_skip: graphql/cloudbuild-skip.yaml
The script skip takes the last successful build and retag the docker images to the current one.
The build script sends the instructions to Google Cloud Container Builder to build the docker image, run the tests, and push it. We use Cloud Container Builder because managing those CI servers really sucks.Let me know what else you want to know.
Comments, issues, pull requests, wiki, etc... should all be first class Git objects. I should be able to create pull requests while offline and push them up when I'm back online.
Git has native "pull requests" it's just not web based, it's email based.
A good wiki is markdown docs in a git/hg repo and something like Gollum to serve web viewers.
So - in a way those things can all be handled as regular content in git/hg.
There is a way to use prepared statements, it's called pre_prepare [0]. You still need to tweak your code though to not execute PREPARE directly but rather put statements in special table, but you do that only once, and after that you EXECUTE like before. Another minor (in my view) caveat is that you need to reset pgbouncer↔postgres connections each time you update that statements table.
BTW, I like how the banner on the bottom of [1] wants me to "Try GitLab Enterprise Edition risk-free for 30 days." currently.
[1] https://about.gitlab.com/2017/02/10/postmortem-of-database-o...
Why is this a problem? If you need to make a change to improve scalability, it doesn't seem unreasonable to ask contributors to follow code and performance guidelines.
Is that because it's already been done and not in scope?
Asking because for (at least) web applications, using Memcached (or similar) can significantly reduce the load on the backend database. :)
Regarding Spokes we're not planning something like that at this time. From https://gitlab.com/gitlab-org/gitaly/issues/650 "GitHub eventually moved to Spokes https://githubengineering.com/building-resilience-in-spokes/ that had multiple fileservers for the same repository. We will probably not need that. We run networked storage in the cloud where the public cloud provider is responsible for the redundancy of the files. To do cross availability zone failover we'll use GitLab Geo."
Re: misuse, I remember this incident https://news.ycombinator.com/item?id=11245652 made me chuckle and cry a little. If you build it, it will be misused.
Note, I'm not saying this is the case with GitLab at all and I assume Postgres will remain the primary choice for most of their DB uses in perpetuity, but for some use cases there is better tech to use and the only reason to say no shouldn't just be a dependency addition (though it is often a good reason among many to say no).
Chances are very good that preemptive scaling will just result in unnecessary development overhead and an architecture/set of abstractions that still fail to handle the expected load since they weren't battle tested against it.
Its just that Postgres may need a little more coding and is less automatic than MySQL that people shy away from it..
I do like the time and information spent about spreading load accross database servers and not nearly enough people are considering splitting reads and writes from a master/slave setup.
Its just that its frustrating to me that engineers view databases as some sort of mysterious black box that you need to solve by working around it or replacing it entirely.
I had my jaw dropped on that point -- they didn't use connection pooling? WTF? I'm seeing connection pools on pretty much any project I've worked on, even under 1/req/day load, it's pretty much "must have" since 90s. And it's really just a library (like dbcp2) and a few lines of config, with almost no drawbacks.
I cannot believe engineers at GitLab level are so bad at, well, engineering. I'm losing faith in their produce, please show me that I'm wrong.
Uhhh... WHAT? I'll simply sit by and watch for SQL injection vulnerabilities resulting from this change. Either you have a very good SQL query writer engine with rock solid escaping or you will get pwned by this.
People have bypassed this far too often for my taste, and there is a reason why SQL (and other injections) are Top 1 of OWASP (1st place actually, both in RC1 and 2 of 2017, but also in the 2013 edition).
Never, ever trust any form of escaping or it will bite you. Hard.
This is an absurd dogmatic statement. Everybody trusts a multitude of escapings everywhere every day all over the web without even knowing it.
https://www.postgresql.org/docs/9.1/static/libpq-exec.html#L...
I don't know what the data access is like, but I get the impression that it's like some "enterprise" patterns I've seen where data access is too "encapsulated", ie, each DAL opens it's own connection, pulls it's chunk of data and each request involves many of these instead of having each request open the connection and pass it down to the data access layers.