Scaling monorepo maintenance
github.blog
github.blog
A git packfile is an aggregated and indexed collection of historical git objects which reduce the time it takes to serve requests to those objects, implemented as two files: .pack and .idx. GitHub was having issues maintaining packfiles for very large repos in particular because regular repacking always has to repack the entire history into a single new packfile every time -- which is an expensive quadratic algorithm. GitHub's engineering team ameliorated this problem in two steps: 1. Enable repos to be served from multiple packfiles at once, 2. Design a packfile maintenance strategy that uses multiple packfiles sustainably.
Multi-pack indexes are a new git feature, but the initial implementation was missing performance-critial reachability bitmaps for multi-pack indexes. In general, index files store object names in lexicographic order and point to the named object's position in the associated packfile. As a first step to implement reachability bitmaps for multi-pack indexes, they introduced a reverse index file (.rev) which maps packfile object positions back to index file name offsets. This alone had a big performance improvement, but it also filled in the missing piece in order to implement multi-pack bitmaps.
With the issues of serving repos from multiple packs solved, they need to efficiently utilize multiple packs to reduce maintenance overhead. They chose to maintain historical packfiles in a geometrically increasing size. I.e., during the maintenance job, consider the first N most recent packfiles, if the sum of the size of all packfiles from [1, N] is less than the size of packfile N+1, then packfiles [1, N] are repacked into a single packfile, done; however if their summed size is greater than the size of packfile N+1, then iterate and consider all the packfiles [1, N+1] compared to packfile N+2 etc. This results in a set of packfiles where each file is roughly double the size of the previous when ordered by age, which has a number of beneficial properties for both serving and the average case maintenance run. Funny enough, this selection procedure struck me as similar to the game "2048".
I also really like monorepos, but Git and GitHub really don't work well at all for them.
On the Git side there is no way to clone only parts of a repo or to limit access by user. All the Git tooling out there, from CLI to the various IDE integrations, are all very ill adjusted to a huge repo with lots of unrelated commits.
On the Github side there is no separation between the different parts of a monorepo in the UI (issues, prs, CI), the workflows, or the permission system. Sure, you can hack something together with labels and custom bots, but it always feels like a hack.
Using Git(hub) for monorepos is really painful in my experience.
There is a reason why Google, Facebook et all have heaps of custom tooling.
Working in environments where different people have partial access to different parts of the code never felt productive to me -- often, the time to figure out who can take on a task and how to grant all the access might take longer than the task itself.
The reason git doesn't have partial repo cloning is because it was written by people without regard to past experience of software development organizations. It is suited to the radically decentralized group of Linux maintainers. It is likely that your organization much more closely resembles Google or Facebook than Linux. Perforce has had partial checkout since ~always, because that's a pretty obvious requirement when you stop and think about what software development companies do all day.
I’ve been using Git/Hg for years and I still run into the occasional Gitastrophe where I have to Google how to unbreak myself.
From my experience git/vcs is not an issue for monorepo. Build, test, automation, deployments, CI/CD are way harder. You will end with a bunch of shell scripts, make files, grunt and a combination of ugly hacks. If you are smart you will adopt something like bazel and have a dedicated tooling team. If you see everything as nail, you will split monorepo into an unmaintainable mess of small repos that slowly rot away.
This means that every single commit every single test suite has to be run across the whole organization. For a small startup that's probably fine, but for a larger organization that adds a ton of issues and extra time. If your change to a shared project breaks another project that you may not be familiar with you end up not only with a delayed context switch (since you don't know about it until the CI has finished running all tests across the whole org), but now you have to figure out how to address it before someone else changes code in that shared project which causes a merge conflict.
Compare that to a split-repo setup where projects are sharing code indirectly (via package management systems). When you change code other projects will only consume the new code when they actually upgrade their packages, and usually a developer that is familiar with the code base will actively hit that breakage in real time.
Granted, this has its own trade offs such as all your systems may be running different versions of shared code, and the delay in breakage detection can cause some back and forth in breakage between isolated projects.
That’s my take anyway
And I agree with the general point that monorepos require a great build tooling as a match.
Note: I don't quite like Bazel; but my take is that Bazel supports too much configurations and options: my workflows are always much simpler. Which is why I am curious to hear what do your workflows require.
Overall, custom toolchains are common (especially, in the embedded world), and a good build system should make that very easy.
You could probably use jsonnet expressions and invoke them from a pre-commit hook.
That seems fairly elegant to me.
Nix bills itself as a fully reproducible package manager and Bazel as a fully reproducible build tool. In the abstract, both of these reproducibly build software, so that puts them in the same category in general.
> Can you re-run Make in a pipeline to verify that the actions in use match what would be generated based on the repo state?
Yes, and I’m doing something similar although not with Make. The problem is that this home-grown build system doesn’t do incremental builds, so everything will be fully built from source every time, and that can take a while for very deep dependency graphs.
Conceivably your “Actions generator” script could generate only the jobs that need to be run based on what changed since the previous commit, referencing artifacts from the previous commit for those build steps which weren’t invalidated. This is a neat concept, but we’re well on our way to reinventing Nix or Bazel.
That X is mostly a function of how open you are to the idea that maintaining + using a centralized, language-agnostic build graph is a Useful Thing.
But once that X is crossed, I don't think it's possible to happily go back to a non-monorepo environment. :)
FWIW, we've used Pants successfully for years, which has excellent Python support. However, we're in a long-haul migration to Bazel, so you might be better off revisiting Bazel if you want to invest in something that's extremely likely to issue dividends over the next 10 years.
Second, if they had decided to fork Git, then they'd have to maintain this fork forever.
Third, this fork could overtime become visibly or even worse subtly incompatible with stock Git which is still the Git running on GitHub users' machines, and both should interact with each other in 100% compatible manner.
So, in this case, not contributing upstream was literally no-go. The only rational choice would be to not fork Git.
I'd have appreciated a series of articles instead of one, for me it's way too much info to take in in one sitting.
If it was broken up, I don't think it would have been nearly as good. And I don't think I would have been able to keep all the context to understand smaller chunks.
I really enjoyed the whole thing.
I'm thinking about writing a blog post where I write a git commit with hexdump, zlib, and vim.
It works well for most repos but as you start to get out to the edges of lots of commits it can cause slowness. GL admins can repack reasonably safely at various times to get access speedups, but the solution that is presented in the blog would def speed packing up.
(I work as a Support Engineering Leader at GL but I'm reading HN for fun <3)
Sure it's simple but it makes it hard to find anything if you have a lot of stuff/people. submodules and package managers exist for a reason.
I think this is a bad analogy. Looking up a file or directory in a monorepo isn't harder than looking up a repository. In fact, I'd argue it's easier, as we've developed decades of tooling for searching through filesystems, while for searching through remotely hosted repositories you're dependent on the search function of the repository host, which is often worse.
The trade off is simplified management of dependencies. With a monorepo, I can control every version of a given dependency so they’re uniform across packages. If I update one package it is always going to be linked to the other in its latest version. I can simplify releases and managing my infrastructure in the long term, though there is a trade off in initial complexity for certain things if you want to do something like say, only run tests for packages that have changed in CI (useful in some cases).
It’s all trade offs, but the quality of code has been higher for our org in a monorepo on average
I'm reading between the lines here, but, I'm assuming you've setup your tooling to enforce this. As in: the various projects in the repo don't just optionally decide to have external references, i.e., Maven central, npm, etc.
This puts quite a lot of "stuff" in the repo, but with improvements like this article mentioned, makes monorepos in git much easier to use.
I'd have to think, you could generate a lot of automation and reports triggering out of commits pretty easily, too. I'd say that would make the monorepo even easier to observe with a modicum of the tooling required to maintain independent repositories.
I have tried in the past, trying to achieve the same goals, particularly around the dependency graph and not duplicating functionality found in shared libraries (though this concern goes in hand with solving another concern I have, which is documentation enforcement), were just not really possible in a way that I could automate with a high degree of accuracy and confidence, without even more complexity, like having to use some kind of CI integration to pull dependency files across packages and compare them, in a monorepo I have a single tool that does this for all dependencies whenever any package.json file is updated or the lock file is updated
If you care at all about your dependency graph, and in my not so humble opinion every developer should have some high-level awareness here in their given domain, I haven't found a better solution that is less complex than leveraging a monorepo
We can call it a multi-monrepo, that way our brainwashed managers will agree to it.
Also it's awful that a simple git pull doesn't actually pull updated submodules, you need to run git submodule update (or sync or whatever it is) as well.
I don't want to work with git submodules ever again. The idea is nice, but the user experience is really terrible.
I'm not at my computer to see if modern git prohibits that behavior, but it is indicative of the "watch out" that comes with advanced git usage: it is a very sharp knife
git subtree tries to emulate that, but it does not scale to huge repositories as it needs to change all commits in the subtree to use new nested paths.
Ideally “cached on the network” could be a sort of optional side effect, like with Nix, but you can still reproducibly build from source. That said, I can’t recommend Nix, not for philosophical reasons, but for lots of implementation details.
I've always had a bad experience using submodules, they're imo the poor developer's versioning tool. It's useful when you use a language without a good build/packaging tool, but otherwise, I'm better off leaving the language-specific tool fetch the depended code.
Also when I try to think about reasons to have atomic cross-project changes, my mind keeps drawing negative examples, such as another team changing the code on your project, is that a good practice? Not really. Well unless all projects are owned by the same team, it'll happen in a monorepo.
Atomic updates not scaling beyond certain technical level is often a good thing, because they also don't scale on human and organizational level.
2. You create a PR that fixes the code and the problematic call sites in a single commit. It gets merged and you’re done.
In the multi-repo world, you need to instead:
1. Add conditional branching in your library so that it supports both the old behavior and new behavior. This could be an experiment flag, a new method DoSomethingV2, a new constructor arg, etc. Depending on how you do this, you might dramatically increase the number of call sites that need to be modified.
2. Either wait for all the problematic clients to update to the new version of your library, or create PRs to manually bump their version. Whoops - turns out a couple of them were on a very old version, and the upgrade is non-trivial. Now that’s your problem to resolve before you proceed.
3. Create PRs to modify the calling code in every repo that includes problematic calls, and follow up with 10 different reviewers to get them merged.
4. If you still have the stamina, go through steps 1-3 again to clean up the conditional logic you added to your library in step 1.
Basically, if code calls libraries that exist in different repos, then making backwards-incompatible changes to those libraries becomes extremely expensive. This is bad, because sometimes backwards-incompatible changes would have very high value.
If the numbers from my example were higher (e.g. 1000 call sites across 100 teams), then the library maintainer in a monorepo would probably still want to use a feature flag or similar to avoid trying to merge a commit that affects 1000 files in one go. However, the library maintainer’s job is still dramatically easier, because they don’t have to deal with 100 individual repos, and they don’t need to do anything to ensure that everyone is using the latest version of their library.
1. A critical security/performance fix has no other recourse than breaking the interface compatibility of a library. Far more common scenario is this can be fixed in the implementation without BC breaks (otherwise systems like semver wouldn't make sense).
2. The person maintaining the library knows the codebases of 10 teams better than the those 10 teams, so that person can patch their projects better and faster than the actual teams.
As a library maintainer, you know the interface of your library. But that's merely the "how" on the other end of those 30 call sites. You don't know the "why". You can easily break their projects, despite your code compiles just fine. So that'd be reckless of an approach.
Also your multi-repo scenario is artificially contrived. No, you don't need conditional branching and all this nonsense.
In the common scenario, you just push a patch that maintains BC and tell the teams to update and that's it.
And if you do have BC breaks, then:
1. Push a major version with the BC breaks and the fix.
2. Push a patch version deprecating that release and telling developers to update.
That's it. You don't need all this nonsense you listed.
In multithreading this would be basically mutable shared state with no coordination. Every thread sees everything, and is free to mutate any of it at any point. Which as we all know is a best practice in multithreading /s
Semver provides just a few bits of information, not nearly enough to cover the whole gamut of shared and distributed responsibility.
The comparison with multithreading is not really valid, since monorepos typically linearize history.
I could have some comments on your "overlapping responsibilities" as well, but your description is too abstract and vague to address, so I'm pass on that. But you literally described the concept of library at one point. There's nothing overlapping about it.
Changing code under their nose risks breaking bunch of projects. We can also fix this by rather communicating, right? But if we CAN communicate... then we can go back to the previous option (telling them to update) as it becomes just as viable.
Communicating is always essential, and can't be avoided.
What happens when you roll this out and partway through the rollout an old version talks to a new version? I thought you still needed backwards compat? I'm a student and I've never worked on a project with no-downtime deploys, so I'm interested in how this can be possible.
If you’re changing the interface of an RPC service, then you can’t do that in a single commit, and need to fall back to something like the second approach, but with even more caution to make sure you properly account for releases and the possibility of rollbacks.
I leverage type systems and write tests to catch any mistakes they might make.
With git, you get a local stage/history which lets you rework/reorder your commits for clarity before pushing. It also allows for more options to resolve conflicts, although this increased ability has brought its own problems.
I assume that this is due to its design center around distributed repositories and patch files. But it’s still there even when using a central repo and a monorepo structure.
Good separation of concerns is like earning compound interest on your code.
Just keep the dependencies generic and tailor the higher level logic to the business domain. Then you rarely need to update the dependencies.
I've been doing this on commercial projects (to much success) for decades; before most of the down-voters on here even wrote their first hello world programs.
If you need to handle different versions talking to each other in production it doesn't seem any harder to also deal with different versions in source, and I'd worry atomic updates to source would give a false sense of security in deployment.
It's much more annoying to deal with multi-repo setups and it can be a real productivity killer. Additionally, if you have a shared dependency, now you have to juggle managing that shared dep. For example, repo A needs shared lib Foo@1.2.0 and repo B needs Foo@1.3.4, because developers on team A didn't update their dependencies often enough to keep up with version bumps from the Foo team. Now there's a really weird situation going on in your company where not all teams are on the same page. A naiive monorepo forces that shared dep change to be applied across the board at once.
Edit: In regards to your "old code talking to new version" problem, that's a culture problem IMO. At work we must always consider the fact that a deployment rollout takes time, so our changes in sensitive areas (controllers, jobs, etc) should be as backwards compatible as possible for that one deploy barring a rollback of some kind. We have linting rules and a very stupid bot that posts a message reminding us of that fact if we're trying to change something sensitive to version changes, but the main thing that keeps it all sane is we have it all collectively drilled in our heads from the first time that we deploy to production that we support N number of versions backwards. Since we're in a monorepo, the PR to rip out the backwards compat check is usually ripped out immediately after a deployment is verified as good. In a multi-repo setup, ripping that compat check out would require _another_ version bump and N number of PRs to make sure that everyone is on the same page. It really sucks.
Atomic deploys are not as important, because you can still decide to version your APIs or releases even if you're using a monorepo.
That being said, you can use multiple repos and still mostly avoid trouble by choosing how to cut your codebase (HR software is likely not going to depend heavily on presale, for instance). The metric to optimize is to minimize the required version bumps.