Supercharging the Git Commit Graph IV: Bloom Filters
blogs.msdn.microsoft.com
blogs.msdn.microsoft.com
I've never noticed the referenced git operation as being slow. In my experience,
git log -- <path/to/file/or/dir>
has always seemed instantaneous or nearly so. It's quick even on code bases that are quite large (years of work from large teams, 10's of thousands of commits).Maybe data and performance information on before vs. after bloom filters would help clarify the specific design goals.
I love seeing bloom filters in use in widely-used software in real life, in any case!
The issue tends to be mostly graph log, note that while this applies to all commit walks it's work mostly done in the context of the commit-graph feature as graphlogging can get fairly slow on big complex repos (mega-scale commit counts).
For example on $dayjob's main repository (~120k commits on master, ~160k total) `log --graph` has a time-to-first-byte of 2.5s, compared to 0.5s for regular log (time <command> | head -n1) and respectively 10 and 4 for complete data generation (time <command> > /dev/null). So there's a 5x difference between log and log --graph for TTFB compared to a 2.5 difference in total difference, log --graph cost is much more "front-loaded". And I expect that difference to grow as the repository grows.
On our monorepo, a simple 'git log -- random_file.cpp' can take anywhere from 10s to a couple of minutes.
Edit: This is on VMware instances with 16 Xeon cores and 32 gigs of memory with SSD backed storage. My puny laptop would probably struggle to even clone the repo.
Lol! The scale at which these companies are (ab)using git is multiple orders of magnitudes greater than that.
You see, they think it’s a good idea to put every piece of code ever written in the whole company in the same repo. They call it the “monorepo”, and it’s hundreds of gigabytes with many millions of commits.
Microsoft even created a virtual filesystem which they run git on top of: https://github.com/Microsoft/gvfs
So, before dissing this approach as 'non-SV-canon', read Google's [1] and Twitters [2] accounts on why they do what they do.
[1] Why Google Stores Billions of Lines of Code in a Single Repository (2016) https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...
[2] On monolithic repositories (2014) https://gregoryszorc.com/blog/2014/09/09/on-monolithic-repos...
Imagine having to torrent a starter pack of the repo, then trying to sync and failing multiple times. Then after it syncs, it could take minutes to do common operations such as change a branch, or check in a file. Not to mention many developer laptops had to have their SSD upgraded as they didn't have enough space to get the whole thing and do work.
I'm sure they've improved it since then, but this was one of many things Twitter simply cargo culted from Google without any real benefit.
My understanding is that they didn't cargo-cult anything. They made a conscious choice at Twitter to pursue the monorepo model.
One reason they mentioned is that while they have a lot of services running all over the place, those services tend to heavily use the same underlying base libraries.
Think the library that validates Twitter usernames, parses out the text of a tweet etc. By having a monorepo they can easily atomically change those libraries for all their consumers. This tends to be why companies go for the monorepo model in general.
This is all information from 2016-ish. Around that time users with *@twopensource.com E-Mail addresses stopped contributing to git.git. I have no idea why, presumably changed internal priorities or something like that. Maybe their internal repository structure changed so they didn't need to author their own performance patches anymore, maybe not...
As I recall the developers authoring the custom git patches left the company around that time. I don’t know what happened after that. There was talk of moving to mercurial with facebooks patches, but I had already left.
Switching branches per-se is really cheap in git, on linux.git it takes 200ms, around the time it takes to run a cold "git status". This is because just creating a new branch doesn't need to touch the tree at all.
I know "status" was a bottleneck at Twitter. They had the first inotify patches to git, but it never made it in. Eventually the patches Microsoft wrote to do the same thing made it in.
What can get expensive is if the tree you're switching to has drastically different content. On the latest linux.git (~60k files) switching to the 2.6.* era takes around 10 seconds or me (~10k files).
I'd expect with a monorepo model like what Twitter had (has?) that most working branches are relatively up-to-date with the master branch, so switching should be cheap.
Was it spending most of its time in the "Checking out files" phase, or before that?
Was this on e.g. OSX with some corporate virus scanner running where each I/O syscall was wrapped? That can drastically slow things down.
Or was it just that the repository truly had a ridiculous amount of files in the checkout (around 1 million?).
Edit: I remember now that I have an old copy of 2015-04-03-1M-git.git which David Turner of Twitter publicly shared a while back, it was meant to emulate the size and shape of Twitter's monorepo. It has around 230k files.
Your Twitter link defends monorepos in principle but is very clear about how Git is not designed to handle monorepos well. This seems to support the thesis that a monorepo is an abuse of Git.
what's that supposed to mean and how do you get it from the parent comment?
In my opinion it boils down to a single factor: if you don't have proper APIs between your components, then you will need to make cross-cutting changes across the whole codebase on a daily basis, thus you need everything in the same repo to work effectively. If you don't have the time to work out the proper boundaries/APIs between your components, then sure, by all means go for the monorepo! But that's like not fixing a blown tire just because you would need to stop to do it.
The APIs are there to ensure the behavior won't change and break anything. In a monorepo, that is easier to check by running all the tests that are affected by your code changes. The actual programming "syntax" is a detail though, and that can be changed (atomically everywhere) anytime with proper tooling.
It's only in the old traditional split model that APIs are supposed to be a guarantee for syntax + behavior (+ sometimes ABI).
If you can build your design philosophy on getting library code to be stable fairly quickly, then making a change to the system need not start an avalanche of rebuilds and redeployments, making monorepo less of a win.
But people like their kitchen sinks, their one stop shopping. And who has time to think about where the best place for a piece of code is. We are doing Agile! And that apparently means “no architecture” to way too many people.
Derrick Stolee, the author of this post, chimes in with some more details about how this works under the hood.
I know his work from generating graphs (eg https://arxiv.org/abs/1104.5261)
A blob is a file (ish, it can also be use for some other things since it's not intrinsically named), a tree is a directory.
However Git doesn't treat filesystem directories themselves as first-class e.g. you can't version an empty directory. A tree only exists in order to contain sub-items (ultimately blobs).
Finding the list of all filenames that were ever in a particular path is very expensive.
The commit-graph feature in general will make these faster by reducing time spent parsing commits. You can compute a commit-graph right now if you have Git 2.18 installed: https://blogs.msdn.microsoft.com/devops/2018/06/25/superchar...
Generation numbers will make these operations much faster in Git 2.19. Here are the related commits:
`git branch --contains`: https://github.com/git/git/commit/f9b8908b85247ef60001a683c2...
`git tag --contains`: https://github.com/git/git/commit/819807b33f820dc17d96f04374...
Those numbers not really useful interactively when very large, for example it doesn't help to print that one of my branches is "behind 132132". Maybe git could print "ahead 7, behind 1000+" for old stale branches stuff? This way it would limit the number of commits examined.
Here is my reply to the thread on-list that summarizes why I think this direction is futile: https://public-inbox.org/git/20180108154822.54829-1-git@jeff...
This ahead/behind calculation is in a lot of places, including 'git fetch' where it checks if each ref update was a forced update (checks if the new ref value has the old ref value in its history). For our version of Git that ships with GVFS, we added an option to skip this check, providing a significant speedup to users fetch times: https://github.com/Microsoft/git/commit/9616c7da3141f539a425...
I’m in a repo with 40GB in .git (hundreds of thousands of commits) and `git branch newbranch` is pretty much immediate.