Speeding up a Git monorepo
dropbox.tech
dropbox.tech
There are costs to having big repositories (e.g. TFA and needing to do scaling work), and there are costs to having lots of repositories involved in producing a unified working result (dependency resolution is NP-complete in many of its useful formulations). Big players have the muscle to optimize Mercurial and Git, so they get to do super slick trunk-only monorepo development at engineer-commit scale, but still often have auxiliary repositories that take e.g. machine generated commits. Smaller players probably aren’t hitting scaling limits on these tools. But every situation is different and if one approaches it as an engineering problem you can usually do something very workable.
Likewise with monolith/microservice: there is a happy medium where you introduce a network boundary for an engineering reason (maybe one part of my computation needs a lot of CPU but a different part needs a lot of RAM, so they run on different SKUs/instance types). My big giant web app that dates to the founding of the company? Probably don’t want to rewrite that so let’s spin services out of it incrementally when I need to write something in C++ or use a shitload of RAM or whatever. That’s bread and butter systems engineering.
But this “pick a side” mentality where it’s like one giant ball of PHP in one giant Subversion repository or every team has their own little service in their own little repo and I burn 40% of my cycles parsing JSON isn’t a set cover: you’re allowed to choose a happy medium.
Just do things for valid technical reasons and don’t have Conway’s law go apeshit on your architecture by shattering it into a zillion pieces. The human factor stuff can be addressed with engineering rigor and consensus. “This is too slow to be in Python now” is a good reason to make a service. “The iOS team shares no code with the web team and bisects will be faster/easier” is a good reason to make a repo.
“I want to have my own coding style and/or use some language no one else knows and/or learn k8s and/or not deal with that team I don’t like” are not engineering reasons to type ‘git init’ or make a network call.
How do you add engineering rigor and consensus? Let's say as an Individual Contributor.
To be a bit more practical, though, yes one should listen to others as well because in some cases they might actually have ideas that are good. Also, morally, it kind of is the golden rule that if one expects to be listened to one needs to listen oneself as well. In cases of large places with lots of people I would say that there should be some form of code ownership and hence people and/or smallish teams can decide what to do with the code that they own. One of the most important things that contributes to code quality is not too many changes of code stewardship.
In my experience engineers don't really think through this question carefully, and have to be prodded to do the right thing, or at least the thing that better serves others or their future selves.
There's also a hump to get over until network effects take over, where engineers have strong incentives to participate in the monorepo.
In your iOS vs web team example, we might imagine that they eventually want to run an integration test against the same API, whose team helpfully provides a hermetic version of the service for testing against. (Dropbox wrote some bazel rules for this sort of thing: https://www.youtube.com/watch?v=muvU1DYrY0w)
- Persistent process which watches workspace changes, or
- Workspace in virtual filesystem.
The other common factor seems to be trunk-based development with all commits rebased into a linear history. I’m not super hopeful that we’ll see an open-source solution in this space for a while, though—any company with a code base large enough to really need these solutions is also large enough to throw a few engineers at VCS, especially given that they’d already have engineers supporting VCS from the operations side of things.
I think Git upstream is trying to simplify configuration. They have a config option called `features.manyFiles` which enables most of the features we enabled for our developers (https://git-scm.com/docs/git-config#Documentation/git-config...).
We wanted to use this instead of deploying a wrapper, but it turns out that some of Git's features like fsmonitor do not interact well with repositories with submodules (there were Git crashes). And we have some developers that work on repositories with submodules. So we needed something more flexible, like enabling these features only on particular repositories.
And we made a few changes to Git to fix bugs (for example, `git stash` wasn't using fsmonitor data, so it was slow).
I don’t want to seem demanding, but it’s such a tantalizing article without this info :)
core.fsmonitor is set to our custom fsmonitor (this causes issues with submodules, at least on 2.24)
core.untrackedCache true
We use index version 4
And a slight hack: our wrapper sets GIT_FORCE_UNTRACKED_CACHE = 1. This forces `git status` to write the untracked cache if it notices a difference. I was too lazy to add a patch to configure that.
Since Scalar is open source, it can be a useful reference for git settings to try in large repositories. I know that if I need to remember how to configure git's sparse "cone" checkouts, Scalar is where I'd look rather than try to manually configure it directly in git's config files.
The last couple of places I've worked at are large non-tech companies, both orgs were internally using on-prem gitlab/github/bitbucket . These tools make it much easier for teams (or individuals) to create as many new repos as they want without coordinating with anyone else -- for better or worse.
I suspect what happens quite often these days is that people create many repos without consciously thinking about if that's a good idea or not -- because it is familiar and because there are relatively high quality tools/products to let you make more repos.
The small part of the org I currently work in probably has O(200) employees and O(200) git repos.
The last system I worked on in previous company had a single git repo containing all parts of a line of business application (db, API server, frontend, backend servers for batch jobs) but then there were about 40 other git repos containing deployment scripts etc used to deploy just this one system. It made it bloody hard to figure out exactly what version of what script or library was actually used to deploy ( to be fair, a lot of this was a consequence of using ansible modules which expects each module to be in its own git repo, and having a couple of people hack together a lot of ansible modules in a short amount of time without review)
I wrote a (fast?) fsmonitor hook in Rust...benchmarked against the reference Perl implementation it's quite a bit faster. On a repo of 130k files, my monitor is able to `git status` in 18 millseconds.
The core value behind monorepo (and monorepo-like approaches) explained in the book is that dependency management is harder than version control.
I wonder if anyone has tried attacking that end of the problem? Faster lstat on macOS would benefit all applications not just git.
But instead they had to build all this extra junk like an additional caching layer and daemon (wasting memory and CPU) atop the kernel's existing cache just to work around the kernel being slow.
sysctl kern.maxvnodesOne major blocker is that the code for this is largely closed source and not accepting patches.
People seem to like monorepos, what problem do you have in mind?
I worked at one place that kept everything in a mercurial mono repo and it was a real pain keeping branches in sync.
Another potential source of issues: the submodule remote URI is checked in as part of the .gitmodules file. If your CI system uses a different URI than your developers, you have to work around that. If you change where your source is hosted and you want to check out an old version, you have to figure that out too.
That said, I think the most common reason is the same reason some people prefer monorepos in the first place: they perceive the monorepo as simpler, and adding submodules is not simpler. They want to know about and manage 1 repository clone. They want to commit, push, review, and build out of 1 repository. They want to search for stuff in 1 directory tree. They'd also like to do that with 1 tool, ideally 'git'. Nothing really offers that except monorepos.
In a many-repo, dependency tracking set up, its an anti-pattern to pin to latest because that can change underneath you and its hard to know what version latest even meant at the time someone committed the dependency.
In a submodule setup, you can track exact dependencies but you pin each submodule reference. You can't easily update all downstream references. You can't use latest for reasons above. You also can't make single commits across many projects.
In a monorepo set up, you can point to latest (ie whatever is in any given HEAD) and trust that everyone will know exactly what that means now and into the future. Monorepo also opens up some ability for build systems across projects and possibly more insight into company wide dependency tracking.
git submodule deinit --all --force
git clean -dfx
git checkout newbranch
git submodule update --init --recursive
There might be cases where you can do a more lightweight switch, but they probably require iterating over the submodules.But I have no idea what came of that effort.
And both Facebook and Google have contributed to Mercurial on this front, though admittedly Facebook did more work than Google did. I know the tree manifest feature (https://www.mercurial-scm.org/wiki/TreeManifestPlan) is done by Google and upstreamed, and it benefits every repo with millions of files. (Just clone the hg repository and search for commits with a google.com author email and see what kind of commits they are.)
In particular, the recent Git v2.27.0 release does include an implementation of Bloom filters with speedups for `git log` and `git blame`. You need to manually run the command to make it work. The version I prefer is this:
> git commit-graph write --changed-paths --reachable
After that first write (which writes filters for every reachable commit) you can do a smaller write by adding `--split` to write incrementally [2]
[2] https://devblogs.microsoft.com/devops/updates-to-the-git-com...
By writing these filters, you will speed up most `git log` and `git blame` calls. There is an improvement coming in the next version that includes speedups for `git log -L`.
Caveat: The biggest reason these improvements have not been widely advertised is that the user experience has not been completely smoothed out. In particular, you can only write the changed-path Bloom filters using the command(s) above. If a commit-graph is written during GC (due to `gc.writeCommitGraph` config setting) then the filters will disappear. Similar for `fetch.writeCommitGraph`. We plan to have these resolved in time for v2.28.0, along with more performance improvements.
(Full disclosure: I am a contributor to Git, Scalar, and VFS for Git, which are referenced by the article.)
Even more importantly, when you have multiple builds using the same repo, CI systems usually work by setting up a local copy of the repository per build, so if you have a compile+unit test build, and another for slower integration tests running the same code, then the machine might end up having two git /objects directories somewhere, containing the same data. If you have a 100Gb repository and maybe 100 different builds against it, this quickly becomes unmanageable.
What I'd want is the ability to use a common objects directory for a machine, where common objects would be deduplicated. I don't know if this is achievable with GVFS or even by linking the directories?
From the article. That's a very big difference.
It would be better if the corporate developers just moved onto Linux Distros. The software and hardware quality is 99% there already.
TL;DR, the Linux codebase there (version 3.1) is about 15M LOC, versus e.g. Windows Vista at 50M, MacOS at 85M, and Google at 2B. Not sure where Dropbox fits in that, but it's probably a similar order of magnitude to Linux and co.
$ time git status
On branch master
Your branch and 'origin/master' have diverged,
and have 1 and 115 different commits each, respectively.
(use "git pull" to merge the remote branch into yours)
real 0m0.517s
user 0m0.240s
sys 0m0.472s
$ git ls-files | wc -l
130685
This is on Ubuntu Linux, no special additions done to git. I wonder why the author's experience is so different.Still highly recommended for large (or just old) repos.
I see why people dislike macOS but banning it at the workplace would be over the top.