Improving large monorepo performance on GitHub
github.blog
github.blog
There's one thing I'd really like to see there: the ability to lock out the repository and perform a really aggressive repack. I'm talking `-AdF --window=500` or somesuch. On $dayjob's repository, the base checkout is several gigs. Aggressively repacking it reduces its size by 60%.
There's also a git-level thing which would greatly benefit large repositories: for packs to be easier to craft and more reliably kept, so it's easier to e.g. segregate assets into packs and not bother compressing that, or segregate l10n files separately from the code and run a more expensive compression scheme on that.
Why is it several gigs? Is that really necessary?
A lot of code written by a lot of engineers over a lot of years.
I'm not sure what other answer you're expecting?
I work with a compiler that has just tens of engineers working on it over just a decade or so and even that's a 6 GB repository. No binary assets. Just source code and configuration I think. I really don't think it's that unusual.
> Is that really necessary?
What would you do? Delete history every year or so? I regularly annotate files and see useful history from ten years ago that I need to do my work.
I've seen some things. Binaries, large files, isos, generally just any large file with no lines that somebody just didn't know better and committed, changed a few bytes a few times, and generated huge deltas quickly.
I'm far more surprised you don't know what he was expecting than I am that he asked, considering how many times I've run into this in the past.
https://github.com/oracle/graal
It's not even a mono-repo - this is just part of the project.
Maybe someone's got some tools that let them dig around in the history and find large things or explain why it's so large? I don't think they've been checking ISOs in.
Receiving objects: 100% (981372/981372), 187.08 MiB | 6.39 MiB/s, done.
% du -sh graal 328M graal
% du -sh .git 221M .git
And that includes many binary files that were added and removed over time, but not filter-branched out of existence, including jars, pngs, pdfs :)
(https://stackoverflow.com/questions/10622179/how-to-find-ide...)
The biggest issue is actually the l10n stuff. It’s all text with lines but there is a lot of it and historically the exports were not really ordered so every major update to the translation files scrambled everything.
It turns out folks can write an ungodly amount of code.
that or check-in binaries, logs, database backups or other things that don't belong in VC. TBH it's pretty much always this
That's a lot of code! I think that is at least a little unusual. The monorepos I see are usually large because people gratuitously commit large binary files (and sometimes change them, generating huge changesets).
Just wait until you have to store deep learning training data.
GitLab and Gitea come with a git-lfs server, but I don't know if there is any standalone server that you could use with git-shell.
It's just too much of a hassle to figure out the security implications of yet another server, and administrating it, setting up the connection between Git and the plugin, etc.
At one of my work places one of the monorepos was over 50GB... took an hour to clone and many minutes to pull/rebase if I was a few days behind...
I don’t like monorepo approach because of such monstrosity, but the upside is it’s easier to commit/sync changes that span multiple products/components, and easier to merge many commits a minute from thousands of engineers...
Most repos could be done in one operation, and If bet that as long as repos stay under 10 commits per second this strategy will always terminate.
This would be a huge rewrite of some internals but seems like it would be a lot easier to manage. It would also provide a some benefits as objects could be shared between repos (although some care would probably be necessary for hash collisions) and it would remove some of the oddness about forks (as IIUC they effectively share the repo with the "parent" repo).
I would love to know if something like this has been considered and why the decided against it.
I bet GitHub has much more read traffic than write traffic, so this trade-off does not make sense.
It seems like the sort of thing that would be an interesting open source research topic if you could build an object database for git that performs better than its packed in filesystem object store. But it's probably not something you want to do as a proprietary project with fewer eyeballs on its performance trade-offs and more engineering work every time git slightly changes its object storage behavior which would remain tuned for the filesystem object store because it was entirely unaware of your efforts.
If you want to get your feet wet, check out go-git[1]. It's a native golang implementation of git. They have a storage layer abstracted over a lean interface that you quickly create alternative drivers for in golang. You'll be effectively implementing poorly sharded file system on a database, then it becomes obvious why scaling the FS is just easier.
I think this is unfair - the author was not insinuating that the people who designed this system at Github are stupid in some way, but just asking if other architectures have been considered.
To me, asking an engineering org if they’ve considered alternative architectures for their main engineering problem is silly at best, overconfident at worst.
> I would love to know if something like this has been considered and why the decided against it.
... but that still sounds more like a grammatical hedge than an actual suggestion that github didn't think it through.
imo it's fair to lay out why you're surprised about some decision in the hopes that someone will enlighten you, even if it can be tricky to phrase that without coming off like a "why didn't you just..." comment.
That’s the polite way of calling them stupids ;)
He didn’t even open with a question, rather he opened with a statement. I don’t think there is even a question mark in the post as it stands right now.
I hate tiptoeing around to avoid offence, but it would have been better to frame the post as a question, it would likely avoided people reading negative intent behind the post.
And hey, at least that means a post GitHub FOSS world won't be leaving fundamental improvements behind!
An object store lacks an index which your typical FS will provide with a relatively high degree of efficiency. FS's can be distributed to arbitrary write velocity given an appropriately distributed block storage solution ( which will provide the k/v API of an object store that you're looking for ). Distributed FS's are conveniently compatible with most POSIX operations rather than requiring bespoke integration. Most object stores are optimized for largish objects and lack the ability to condense records into an individual write (via the block API) or pre-emptively pre-fetch the most likely next set of requested blocks.
In the GitHub's case the choice of diverging from GitCLI/FS based storage APIs could lead to long term support issues and an implicit "github" flavor of git rather than improving the core git toolchain.
Object Stores are great, but if you need some form of index they get slow and painful really fast.
Example: you can add erasure coding to the blob-data service for better efficiency. You can add fancy indexing to your metadata store. etc etc.
But somebody has to create it, that's the issue.
Systems such as HDFS use the NameNode for this task, but depending on the exact characteristics of the fileSystem a multi-master setup is often used. I know of at least one NFS implementation which uses postgres as its metadata layer.
A filesystem really is just an index over k/v storage (blocks). You can find similar diagrams for the implementation of EXT/XFS/GFS/HDFS etc. Like databases there are many tradeoffs and nuances in terms of how these concepts are implemented.
Maybe I'm a bit dense, but how did you get that from the article? I'm fairly certain that in other pieces of writing they showed that they are using an object store, and I'm guessing that's what the "file servers" in the article are.
> Perhaps it’s surprising that GitHub’s repository-storage tier, DGit, is built using the same technologies. Why not a SAN? A distributed file system? Some other magical cloud technology that abstracts away the problem of storing bits durably? The answer is simple: it’s fast and it’s robust.
1) Framing it as such with poor justification "a lot easier to manage"
2) "This would be a huge rewrite of some internals" Becoming a multi-year migration quagmire
3) The dawning realization that you have used a write-heavy architecture in a read-heavy system
Scalable receive, but having to send hundred of thousands of objects one-by-one is very inefficient. The fact that there are hundreds of thousands of servers to receive them does not help you.
(and same problem when you pull)
1) The filesystem is a surprisingly decent datastore,
and,
2) Unless your write-volume is unreasonably high for a GH repo, your main problem is making sure your locking (and unlocking) is rock-fucking-solid; Git has built-in locking, but I bet Github has either replaced it or has some serious state-machine stuff going on to manage it outside of Git itself,
also and,
3) Reads on files are ultra-cacheable, obviously, and reads on HTML pages are too, and serving a second-or-two-stale read isn't a problem ~100% of the time in GH's use case, so read volume isn't really a problem.
[EDIT] source: have designed and developed a git-backed, client-accessible product at a much smaller scale than GH, but enough to see where the problems & strengths could/would be, at greater scale.
[EDIT EDIT] oh and with a bunch of clients using a bunch of different repos that effectively never interact except through normal git-push/PR behavior, obviously horizontal scaling is a breeze and a half. All you need is a little routing to figure out which repo is where. Git itself, plus your locking solution (whatever that may be), give you the tools you need to migrate from one server to another if you need to move a repo, probably with an unnoticeable amount of downtime (=locked repo) if you're clever about it.
Something like the ability to have github.com/<org>/<repo>/<subproject>/issues be a shard of all the issues for a subproject.
You can do that with tagging, but that's a bit of a PITA because that's all fairly bad and unscalable of a UI.
On a sidenote, git itself can also get painfully slow with large monorepos. Hope GitHub can push some changes there as well.
I know FB moved off git to mercurial because of performance issues.
e.g. https://facebook.github.io/watchman/ - used as part of Facebook’s Mercurial solution, I think.
The Mercurial extensions are then an alternative client for Piper.
One very interesting part of that is the effort that has gone into the git commit-graph: https://git-scm.com/docs/commit-graph.
It's part of what makes scalar interesting compared to some of the projects you hear mentioned used inside the FB and Google gates: not only is scalar itself open source, but a lot of what scalar does is tune configuration flags to turn on optional git features such as the commit-graph, sparse checkout "cones", etc that are all themselves directly supported by the git client. Even if you aren't at the scale where it makes sense to use all of the tools that scalar provides, you can get some interesting baby steps by following scalar's "advice" on git configuration.
https://github.blog/2020-01-17-bring-your-monorepo-down-to-s...
Or at least that was the plan at 2016 when I left Google.
say you have a structure like: projectA projectB sharedUtils
Each time you push you might have a build for projectA and projectB but it builds both each time you push to master. Ideally you could use Git to see if anything in projectA or sharedUtils changed to trigger projectA's build and same for projectB, but I'm curious what others are doing.
I have a big monorepo at work but whenever anything changes I want to rebuild everything to generate a new firmware image. I have ccache setup to speedup the process given that obviously only a tiny fraction of the code actually needs to be rebuilt.
It's a bit wasteful, sure, but if I were to optimize it I'd be worried about ending up with buggy, non-reproducible builds. Easier to just recompile everything every time and make sure everything still works the way you expect.
So basically my approach is KISS, even if it means longer build times.
See this documentation page: https://docs.github.com/en/actions/reference/workflow-syntax...
Their example is:
on:
push:
paths:
- '\*.js'
but I believe you can also specify the subdirectory you care about.If you do a full build on every commit, it gets slow much sooner than you'd expect, and people are going to do less work while they context switch to posting to HN while waiting for their 15 minute build for a 1 line code change.
I worked at Google and we had a monorepo, and there were hundreds if not thousands of engineers working on build speed and developer productivity, and it was still significantly slower to "bazel run my-go-binary" versus "go run cmd/my-go-binary". In many cases, it was worth it, but in very isolated applications, it was definitely not worth it. (And people did work around it, by just setting up Git somewhere and using Makefiles or whatever, and that ended up being even worse. But it gets worse incrementally over time, and you're kind of the frog getting boiled alive.)
Where I'm going with is to advise you to be very careful. The tools to support real productivity in a monorepo are expensive in terms of your org's time. If you can get by with a repo per app and a common modules repo, and just update the app to refer to a version of the modules repo as though it's some random open source project you depend on, you're going to get much farther with much less tooling work than you would with a monorepo. But, the modules repo is going to break apps without knowing, and that's going to be a pain. Monorepos do exist for a good reason.
(The other thing I like about monorepos is that you do less per-project setup work. Want to make some new app? You can just start writing it, and you get the build, deploy, framework, etc. for free. It can be very productive if you're finding yourself starting new projects regularly. In my spare time, I write a software, and I really regret splitting it up into multiple projects. But, it's kind of necessary for open source stuff -- people don't want to download ekglue if they want to just run jlog. So I split them, but it costs me my valuable free time to do something I've already done ;)
My TL;DR is that you will be tempted to take shortcuts and the shortcuts will suck. If your project has the resources to have someone set up Bazel, distribute the right version of Bazel and the JRE to developer workstations, setup CI that is aware of Bazel artifact caching, and SREs to be around 24/7 to support your now-custom build environment, you will have a good experience. Be aware that a monorepo is that level of investment.
Meanwhile, if you just have a frontend and a backend in the same repo, you can probably get away with a full build for every commit. And you don't need that shadow team of tooling engineers to make it work, you just need a docker build, and a script that runs "go test ./... && npm test" or whatever ;)
This ordering is probably in order from "most" to "least" obvious, and I don't have an answer for what the correct solution is. Most probably "separate repo", making the interface not privileged to either side of the "interface barrier".
FWIW, building at Google when at a 8h offset from Pacific Time worked fine, indicating that part of the problem is "does your repo/build solution scale to tens of thousands of simultaneous users".
However... I would strongly advise not going for a monorepo. No, I don't mean something like tensorflow where you have a bunch of related tools and projects in a single repo. I mean one repo for the entire org where totally unrelated projects live.
Every company I've been at that used a monorepo found themselves struggling to make it work since you need a ton of full time engineers just to keep things working and scaling. Many of the problems that monorepos try to solve (simplifying dependency and version management) are traded for 10x as many problems and many of them are hard (incremental builds, dependency resolution).
Google has a huge team in charge of helping their monorepo scale and work efficiently. You are not google... don't be tempted.
My contrasting anecdotal experience is that whether at BigCo or on a small team monorepo is almost always the right answer until your requirements get exotic enough that you're in special-case land anyways (like a separate repo for machine-initiated commits, or something that's security-sensitive enough to wall off some contributors).
Both `git` and `hg` scale easily to to really big projects if you're storing text in them (at FB our C++ code was in a `git` monorepo on stock software until like 2014 or something before it started bogging down, I'll gloss over the numbers but, big): the monorepo-scaling argument is brought out a lot but rarely quantified.
The multi-repo problem that gets you is dependency management, which in the general case requires a SAT-solver (https://research.swtch.com/version-sat), but of course you don't have a SAT-solver in your build script for your small-to-medium organization, so you get some half-assed thing like what `pip` and `npm` do.
Again purely anecdotal, but in my personal experience multi-repo too often gets pushed by folks who want to make their own rules for part of the codebase ("the braces go here"), push an agenda around unnecessary RPCs, or both. That's not true of all cases of course, but it's a common enough antipattern to be memorable.
It’s funny, I’ve heard this exact same argument for why you should not use micro services.
At my job our cloud team uses multiple separate repositories which makes sense, but it also moves the burden of versioning to run-time. This is because they have to interface with multiple different versions of the device firmware. So they deploy different run-time versions of the APIs to support legacy and current production firmware versions. But our firmware repository is a monorepo in that the sources and build system builds the artifacts for multiple devices from the same source tree.
So it's not so cut and dried as "never use a monorepo" or "always use a monorepo." It involves engineering tradeoffs and decisions that are made in a context, and you can't extract your advice from the context in which it exists. What works for our cloud team would be a terrible mess on the embedded side simply because of how the software is deployed and managed.
Working effectively in a monorepo does require suitable tooling (full disclosure: I am a core contributor to Pants, a monorepo build system with, currently, a Python focus). But with that tooling in place a monorepo can make it a lot easier to collaborate, to manage changes and dependencies, and to avoid balkanization and fiefs. Granted there are no silver bullets, but we are working to make tooling good enough to avoid needing the large in-house support team you allude to.
Having a large number of small repos can work if they don't have a lot of interdependencies, and maybe this is what you mean by "totally unrelated projects". But if there are significant interdependencies (and repos usually end up this way, in my experience, which may or may not be representative) things become very brittle, and it's very hard to make changes and reason about how they affect downstream code, not to mention the challenge of version resolution, aka "dependency hell".
In the end, managing a large codebase, whether it follows a monorepo or multi-repo architecture, requires suitable tooling. On balance I think that monorepo is often a better choice, but YMMV for sure.
This is worth a read: https://yosefk.com/blog/dont-ask-if-a-monorepo-is-good-for-y...
The main point is the "better tooling than average" section near the end - you need tools that can understand relationships between different parts of the codebases, and test appropriately, which is where the integration process for monorepos starts to fall apart.
Building on that dependency graph and soundness in Yarn 2, I built yarn.build [0] which works in a similar way to all the monorepo build tools.
It builds your packages based on their dependency graph, skips whats already built, and has a bundle command for creating a zip ready for lambda/docker/etc
Isnt this changeset introducing a race condition? One of the replicas' checksum could change between the checksum is computed and the lock is taken. Otherwise, there is no need for the lock at all.
The key thing here is that prior to the lock if data changes you recompute the checksums. As long as any change outside the lock triggers a recompute of the corresponding checksums and no changes can occur during the lock, there is no race condition.
I imagine that this may result in data getting de-synced/failing the checksum comparisons more often however it's still a net performance increase as long as the aggregate time spent re-syncing the data is less than the extra time spent waiting for checksums in the lock.
In general, you cant assume your code will always read the new checksum before entering the critical section, unless you synchronize it. That is, when leaving the critical section you have to make sure that all threads waiting to enter it will recompute the hash.
But, from the entering thread's perspective, you never know when youll be forced to recompute. So you've got to wait on the lock. And recompute when you are woken up.
I'm not stating what they've written is wrong. Just for me it's a bit vague and looks like potential race condition.
1. All the zones start precomputing the checksums using a worker pool. Changes are fed into the pool with the change time stamped. The workers don't store the checksum if a newer checksum is already present.
2. The zones enter the lock/critical section. At this point new changes are no longer added to the work queue and must wait until the lock exits. The worker pool continues to process the queue.
3. The work queues are empty and the checksums are compared between zones. Those that match are "locked in" and those that don't are set aside for the next time the zones enter the lock/critical section.
4. The zones exit the lock, a timer starts, and step 1 starts again.
At no point here would there be a race condition. As long as non-matching changes are pruned and retried in the next cycle, progress will be made. In certain conditions the retries could degrade overall performance but consensus is eventually achieved and the system continues to make progress.
I don't recall the article mentioning it, but if you don't compute the actual checksum under the lock then you either cache it (under the lock) and update only successful writes, or you risk a race condition between time of the check and action.
That is, optimistic locking assumes that if the condition is true I immediately get exclusive lock. But here, condition check is completely separate from the action. When the checksums are equal, but another thread manage to enter and leave the critical section before you, then that invalidates the condition. But you need to synchronously let other threads know they must recompute.
Technically it works the mostly the same as multiple repos, but theoretically allows to have something like a bootstrap script with everything self contained in the same repos. Looks like an alternative tradeoff between a monorepos with shared history and multiple repositories.
It's offtopic because the second worktree has a detached HEAD, so that doesn't help in the case you mention.
As far as I know git worktree is just to have different branches of the same repo checked out in different locations. At least, that's the only way I use it (and it's great!). Are you suggesting to have different projects on different branches? So an empty "master", then "project1", "project2" etc. as branches?
I use worktree locally so that, for example, I can have my working copy that I am doing development in and then a separate working copy where I can do code review for someone else, without having to interrupt what I am doing in my own worktree.
My own experience is that if you are using branches with radically different content for different purposes in the same tree, it's going to end up a mess at some point. Worktrees, as far as I am aware, do not help with that in any special way.
I don't believe VFS for Git will ever be abandoned by Microsoft, but I'm doubtful it will ever get any more major improvements from them.
Scalar does use the VFS for Git client-server protocol, and both Scalar and VFS for Git rely on the same improvements to the git app itself, so I could imagine that GitHub would adopt the GVFS protocol and support Scalar without formally supporting GVFS itself.
[1] GitHub did announce future GVFS support in 2017 - https://venturebeat.com/2017/11/15/github-adopts-microsofts-... - but if anything came out of that I don't see it in GitHub help today.
When I left Microsoft about half a year ago, GVFS and Scalar were both in heavy use there.
Scalar also supports Windows and macOS, while VFS only supports Windows: https://github.com/microsoft/VFSForGit/blob/v1.0.21014.1/doc...
Right now we don't plan on supporting it; most of our work is focused on upstreamable changes and opinionated defaults. But that could change if we're missing some important use cases.
Feel free to email me - my HN alias @github.com - if you prefer to discuss privately.
Forest Mountains of Zion National Park, Utah
Here's a recent post: https://devblogs.microsoft.com/devops/introducing-scalar/
For me, its any repository where I would think "damnit im going to have to do a fresh clone" if the situation comes up. There isnt a hard line in the sand, but there is certainly some abstract sensation of "largeness" around git repos when things start to slow down a bit.