Both gitlab and bitbuckets source code browsing is slow and a little clunky. GitHub's source code browsing is definitely the best.
Both gitlab and bitbuckets source code browsing is slow and a little clunky. GitHub's source code browsing is definitely the best.
We solved most of the time waiting when pushing a new commit, see the API timings slide on https://www.scribd.com/doc/316471059/GitLab-Infrastructure-2...
The web interface is still slower than we like. We've doubled the team of performance engineers and we're making progress, see https://gitlab.com/gitlab-com/infrastructure/issues/59 and all issues labeled with performance https://gitlab.com/gitlab-org/gitlab-ce/issues?scope=all&sor...
Does the choice of RoR play a part in performance problems?
Can performance be improved by deploying more servers?
PS: I really enjoy using Gitlab, and would be ready to replace Github with Gitlab in my workflow if the performance improves.
We're working on switching to CephFS which looks very promising. See the relevant issue for more info: https://gitlab.com/gitlab-com/operations/issues/1
I can get someone from the backend performance team to comment too if you'd like :)
Sure!
> the biggest bottleneck at this point is with the file system.
Would something like AWS EFS [1] solve the problem?
It might not be a practical solution from a cost perspective though. Eg: for 50TB of data at $0.3/GB/month
0.3*1024*50 = $15360/month
[1] https://aws.amazon.com/efs/We build a product that you can host yourself. If we only solve the scaling issue by pushing it down to a specific vendor, then there is no actual solution.
The way we are facing the problem is first by enabling a really easy and simple form of sharding (that would solve most of the issues that a lot of big customers may face), and then by using an open source underlaying filesystem that can scale reasonably well.
> what are the major areas where you face performance issues? > Does the choice of RoR play a part in performance problems?
RoR does not play a part in the performance problems as much as any other language choice. Our performance problems come from at least 3 different fronts: lack of caching in some specific points making us call the same complex/slow operations many times, NFS (filesystem) performance as a whole, and algorithms that worked really well at small scale, but not anymore, both at app level and at DB level.
I think that RoR is a really good option for building a product fact, and eventually it is necessary to start specializing specific parts that do not perform anymore, and just replacing what cannot be specialized. The key element here is that we need to measure first to see where the problem is.
> Can performance be improved by deploying more servers?
Not really, more front end servers means more load on the NFS backend, so there is no easy solution here. The first step into fixing this issue right now is this one: https://gitlab.com/gitlab-com/infrastructure/issues/139 and then we will start playing with a distributed FS as a longer term solution: https://gitlab.com/gitlab-com/operations/issues/1
As a final note, we are using some specific issues to measure performance as a blackbox, and those issues are seeing some really good progress lately. So stay tuned :)
Could you also comment on this https://news.ycombinator.com/item?id=12056991
Regarding web performance, these last days we had really good progress: https://gitlab.com/gitlab-com/infrastructure/issues/193#note...
our github enterprise will not bother try show diffs on some pull requests if you changed more than a hundred lines or so. very worthless.
my workflow now is to always check the diffs locally because i don't trust theirs.
never used gitlab or bitbucket a lot so i don't know if this treachery is there too.
In Bitbucket Cloud - i.e. bitbucket.org - we cap the diff size for pull requests at 10000 lines.
In Bitbucket Server we also cap the number of lines in a particular diff at 10000 by default, but you can override this number[0], along with all other timeouts and thresholds with our various configuration properties. One of the engineering values of the Server team is "no hardcoded constants" - so you can configure basically any property that you like to suit your particular deployment.
Both Bitbucket Server and Cloud also generate a subtly different - and in our opinion more correct - diff than you'll find in GitLab and GitHub. Bitbucket actually creates a hypothetical merge commit between your two branches and shows the diff between it and the tip of the target branch. This means we can nicely render merge conflicts in the UI, and show how your target branch will actually be affected by the merge (rather than just the changes on the source branch). I wrote an article that discusses our merge algorithm in more depth[1] a little while ago.
[0]: https://confluence.atlassian.com/bitbucketserver/bitbucket-s... [1]: https://developer.atlassian.com/blog/2015/01/a-better-pull-r...
edit: correcting my previous statement on diff sizing in Bitbucket Cloud
In fact, it's a part of our code review/pull request functionality, therefore "it just works".