We put half a million files in one Git repository (2022)
canva.dev
canva.dev
Jokes aside, and coming from a place of ignorance, it's interesting to me that a file count that size is still a real performance issue for git. I'd have expected something that's so ubiquitous and core to most of the software world hasn't seen improvements there.
Genuine, non snarky question: Are there some fundamental aspects of git that would make it either very difficult to improve that, or that would sacrifice some important benefits if they were made? Or is this a case of it being a large effort and no one has particularly cared enough yet to take it on?
It’s hard to look at a million files on disk and figure out which ones have changed. Git, by default, examines the filesystem metadata. It takes a long time to examine the metadata for a million files.
The main alternative approaches are:
- Locking: Git makes all the files read-only, so you have to unlock them first before editing. This way, you only have to look at the unlocked files.
- Watching: Keep a process running in the background and listen to notifications that the files have changed.
- Virtual filesystem: Present a virtual filesystem to the user, so all file modifications go through some kind of Git daemon running in the background.
All three approaches have been used by various version control systems. They’re not easy approaches by any means, and they all have major impacts on the way you have to set up your Git repository.
People also want e.g. sparse checkouts, when you’re working with such large repos.
There are various versions of Git that use alternative object storage, like Microsoft’s VFS, if I remember correctly.
That's not, the last time I checked, how everything on Windows works.
Windows provides the ability to hook into FS system calls, so that things like virus scanners work.
Everything uses the hook to get notified of all changes, and uses those mods simply to update its index (which is faster than scanning a file for viruses, so it's imperceptible to users).
It's a great idea, and I don't think there is anything similar in Linux or BSD (inotify isn't the same thing, AFAIK, it uses up file descriptors).
>Findlargedir is a tool specifically written to help quickly identify "black hole" directories on an any filesystem having more than 100k entries in a single flat structure. When a directory has many entries (directories or files), getting directory listing gets slower and slower, impacting performance of all processes attempting to get a directory listing (for instance to delete some files and/or to find some specific files). Processes reading large directory inodes get frozen while doing so and end up in the uninterruptible sleep ("D" state) for longer and longer periods of time. Depending on the filesystem, this might start to become visible with 100k entries and starts being a very noticeable performance impact with 1M+ entries.
>Such directories mostly cannot shrink back even if content gets cleaned up due to the fact that most Linux and Un*x filesystems do not support directory inode shrinking (for instance very common ext3/ext4). This often happens with forgotten Web sessions directory (PHP sessions folder where GC interval was configured to several days), various cache folders (CMS compiled templates and caches), POSIX filesystem emulating object storage, etc.
Since git is essentially a filesystem with extensive version control features, it doesn't surprise me that it would have problems handing large amounts of files.
But there will be some trade-off.
And I don't think people generally put "a million files" in the requirements because it's fairly rare.
It's hard to get people past the demo phase "works for me" when they have played with one image, to realize they really need a reasonable container format to play nice with the systems world outside their one task.
Now I have a system that will find subsets in just a second or two (even when the whole set contains hundreds of millions and any given subset might contain hundreds of thousands of matches). Here is a short video of a demo: https://www.youtube.com/watch?v=dWIo6sia_hw
ext4 on an old system, feeble in comparison to yours, performs much better.
ext4, 8GB memory, 2 core Intel i7-4600U 2.1GHz, Toshiba THNSNJ25 SSD:
$ time ls -U | wc -l 555557
real 0m0.275s user 0m0.022s sys 0m0.258s
stat(2) slows it down, but sill this is not as poor as your results:
$ time ls -lU | wc -l 555557
real 0m2.514s user 0m1.126s sys 0m1.407s
Sorting is not prohibitively expensive:
$ time ls | wc -l 555556
real 0m1.438s user 0m1.249s sys 0m0.193s
Drop caches, sort, and stat:
# echo 3 > /proc/sys/vm/drop_caches
$ time ls -lU | wc -l 555557
real 0m6.431s user 0m1.249s sys 0m4.324s
I always marvel at it and think: "wow so git goes through its history, pulls out many small files and chunks and patches, updates the whole file tree and all of this after hitting enter and being done like immediately."
I can't speak to improving git, but I think some light on this area can be shed by Linus' tech talk at Google in 2007.
1. Linus says there's a specific focus on full history and content, not files ... so it's a deliberate, different axis of focus than file count:
https://youtu.be/4XpnKHJAok8?t=2586
... AND it's a specific pitfall to avoid when using Git:
https://youtu.be/4XpnKHJAok8?t=4047
2. As Linus tells it, Git appears to be designed specifically for project maintenance while not getting in the way of individual commits and collaboration. But the global history and more expensive operations on things like "who touched this line" are deliberate so lines of a function are tracked across all moves of the content itself.
Maintainer tool enablement: https://youtu.be/4XpnKHJAok8?t=3815
Content tracking slower than file-based "who touched this": https://youtu.be/4XpnKHJAok8?t=4071
===
I have no answer, but ...
Practically, I've used lazy filesystems both for Windows-on-Git via GVFS [1][2] and Google's monorepo jacked into a mercurial client (I think that's what it is?). Both companies have made this work, but as Linus says, a lot of the stuff just doesn't work well with either system.
Windows-on-Git still takes a lot of time overall, and stacking > 10 patches of an exploratory refactor with the monorepo on hg starts slowing WAY WAY down to the point where any source control operations just get in the way.
[1] https://devblogs.microsoft.com/devops/announcing-gvfs-git-vi...
I've seen production databases that had 40,000 tables for valid reasons.
I've personally deployed an app that needed 80,000 security groups in a single LDAP domain, just for it. I can't remember what the total group number of groups across everything was, but it was a decent chunk of a million.
Making something like Git, or a file system, or a package manager? Test what happens with millions of objects! Try billions and see where your app breaks. Fix the issues even if you never think anyone will trigger them.
It's not about scaling to some arbitrary number, it's about scaling, period.
All the effort spent on testing and optimizations for scaling purpose was a waste of time and resources, that could be better spent elsewhere.
I'm not telling one should not care, or code sloppy, but there is a balance where code is just good enough for the purpose. There's a lot of truth in this "don't do premature optimization".
A parent comment mentioned SQL databases for example; those are great because they can scale both horizontally and vertically these days, sometimes with the click of a button in AWS.
Other good practices are things like stateless back-end services so they can scale horizontally, thoroughly documenting (and maintaining documentation) on business processes handled by the software, monitoring, etc.
Disclaimer: I'm an armchair expert, I've never had to deal with back-end scaling.
We build some services from the ground up for very high traffic and the hoops you have to jump through and the tradeoffs just don’t make sense for a basic CRUD thing which can run on a boring ole machine and a little SQL instance
I work in enterprise software and one of the big problems I see is companies suck at software growth when it's obvious the software is in the upward curve.
Large companies will throw huge amounts of data at your app once you sell it to them.
But that chance isn't the sole reason why you're doing that load test. The reason is to improve the software. You're identifying defects by stressing the limits.
When you're doing a load test (or any test really) the possible outcomes are basically three: (1) it works! (2) it broke. (3) huh, that's interesting. If your tests are always coming up (1) then you're not obtaining any benefit from them. Don't you want to know where the limiting factors are in your app? If you're able to remove those limits, but not for production (at least not right now), wouldn't it be great to know what will break next month (or next year) when you do?
Think of the person who writes unit tests for every piece of code, but not as TDD. There's a school of thought that you should write the test first, then write the simplest code that passes, and that's fine but not what I'm talking about. Imagine a person who writes perfect code and perfect tests. Every code works, every test passes confirming that it worked. What is even the value of writing the test?
That's what load testing under only the expected conditions is like. We already know the software works under those conditions (likely) because it's already in production, handling that amount of load. So while there is value in a load test that runs prior to deployment, in order to check that nothing of the change is likely to induce a break under the expected/existing load, it's a different kind of testing and produces different value than a stress test that is designed to hopefully induce a failure and show you where there is a defect in code. Where it segfaults, for example.
And just because you've identified a limiting factor outside the bounds of what expected activity is likely to go through the system soon, doesn't mean you need to fix it now. Having one less "known unknown" on the table is a thing of value. Now that stress won't be able to surprise you later, when that parameter has drifted into the danger zone because of organic development, and now it's becoming a thing in the way.
The time you spend making it work for millions of users that won't be needed is time not spent making value to customers that do need it.
Someone experienced will know how much work a certain approach will take and its capacity.
Sometimes there are quick wins to give like 100x capacity to a system just by doing things slightly differently, but only with experience will you know that.
What if scaling to millions of objects forces real tradeoffs for the hundreds of objects case?
It feels like you're asking people to only create Postgres, but SQLite has a perfectly valid use case as well.
In this case, git checking the access time of 500k files is fundamentally slow. The only way around this is to change how git tracks files, which all come with other usability tradeoffs. Git itself supports a fsmonitor that makes handling more files faster, but very few people use it because the tradeoffs aren't worth it.
I don't mean this to be a "well actually" comment, but because I found it interesting when I learnt this a few weeks ago - some limits for SQLite [1] are actually higher than the limits for Postgres [2] (specifically the number of columns in a table and the maximum size of a single field).
Sometimes edge cases can quickly detect bugs that only happend rarely under normal circumstances and therefore is difficult to reproduce/debug.
E.g. when programming in C for little endian computers it can be a good idea to test code on big endian CPUs as the difference in endianess can reveal "out of bounds" writes for pointers.
One thing you should learn, and many don't, about perf analysis is that you start getting serious artifacting in the data for tiny functions that get called an awful lot. I've found a lot of tangible improvements from removing 50% of the calls to a function that the profiler claims takes barely any time. Profilers lie. You have to know what they lie about.
When I'm trying to optimize leaf- or near-leaf-node functions I've been known to wrap the call with a for loop that runs the same operation 10, 100, 1000 times in a row, just so I can see if some change has a barely-double-digit effect on performance. These predictions usually hold up in production.
Just be very, very sure not to commit that for loop.
Or use representative data that is ridiculously large compared to the average case.
Unlike with many user-space applications, filesystems have very well-defined range of conditions they have to work in. Eg. every filesystem worth its salt will come with a limit on number of everything in it, i.e. number of files, groups, links and so on. And these limits are tested, they aren't conjectures. Ask any filesystem developer how many metadata operations per second can their program do, and they will likely be able to answer you in their sleep. This might be surprising on the consumer end of the deal, but to the developers there's nothing new here.
The problem here is that we need a reasonably small interface for filesystem to enable competing implementations, so, for example, we don't have a filesystem interface for bulk metadata operations, because this is an unusual request (most user-space applications which consume filesystem services don't need it). So, we can only query individual files for metadata changes through "legal" means (i.e. through the documented interface). And now you end up in a situation where instead of fetching all the necessary information in a single query, the performance impact of your query scales linearly with the number of items queried.
Even if Git developers anticipated this performance bottleneck, there's not much they can do w/o doing some other undesirable stuff. Any solution created outside of the filesystem would risk de-synchronization with the filesystem (i.e. something that watches the state of the filesystem dies and needs to be restarted, either loosing old changes or changes done between the restarts). Another solution could try going behind the documented filesystem interface, and try to salvage this information directly from the known filesystems... which would be a lot of work compounded with the potential to screw up your filesystem.
Maybe if we'd have Git integrated with the kernel and be able to thus integrate better with at least the in-kernel filesystems. But this would still put people on anything but Linux at a disadvantage, and even on Linux, if you wanted some filesystem that's not in the kernel, you'd also have the same problem...
Test with, say, one of largest open-source projects in existence at the time? Like, for instance, the Linux kernel?
* Most simple things like a "Person" were multiple tables because you had to include audits and historical changes for each field
* A "Person" wasn't even all that useful because it included guests or other fairly transient entities like vendor contacts so you had an explosion of more tables as you classified roles into "Student", "Faculty", "Employee", etc... (many with histories as above).
* Addresses and other non-core demographic information were usually sharded into all sorts of categories like "primary", "parent's", "last known good", "good for mailing", etc... (more histories, etc...)
* All coded information like label types such as "STUDENT", or "MAILING" were always handled as separate validation tables with strict FK constraints and usually included extra meta information like descriptions and usage notes within parts of the system.
* Each functional sub-system (HR, Payroll, AR, AP, etc.) had its own dedicated schema.
* All external jobs, processes, and external integrations were configured separately.
* All enterprise integrations usually had a whole a dedicated schema for configuration.
* Most parts of the interactive web UI were database driven (Oracle's Apache mod PL/SQL) with many templates and other components stored in large collections of tables.
I'll stop there, but basically just imagine a very large application that tries to be 100% database-driven. That's how you get a lot of tables.It probably feels weird for devs to drive the UI off the db but it's just Wordpress by another name.
When I rewrote the application I just hardcoded the form fields, nobody should need to do a database migration to change an otherwise mostly static form.
Would it be "better" if they had one table with json/xml/whatever and handled schema in code?
They made a trade-off they found right. When they hit the limit with their approach, they even implemented their own DB (S4/Hana) to support their system.
The alternative would be we both have access to the same tables with a permission layer to grant access to row.
Both choices have trade offs but if company makes a mistake and I now have access to your rows? Seems easier to control access at the table layer rather than the column layer.
Or... you can just split things by tables. Or even shard by databases where I don't have access to your database and vice versa.
doing stuff in the application and leaving everything in one database/schema is an option... but don't think you aren't making trade offs and leaving open possible issues by not taking the more comprehensive option like sharding.
And that's just one question to ask. Another is what about upgrading the database and segregating customers. can't do that if everyone is on the same database/schema. What if a customer doesn't want to be updated or upgraded? Much like companies paying for Windows XP support because stuff they have relies on the older version of software?
"where user.id = 123" is a simple solution that quickly becomes more complicate to put it mildly.
edit: I tell a lie, I separated the forums and wordpress databases on a website I run.
Thankfully, our product has customer specific use patterns that we've been able to manage/plan/predict peak load for and what not. Of those 300, a random subset of 20-30 would be 'busy/critical' at any given time, and the others can easily tolerate a delay as migrations + schema changes lock and manipulate things.
I'd rephrase it: if expecting massive or longterm use - know how that thing is built/designed.
Picking a technology based on general popularity or vendor's marketing is not a way to solve _your_ problem.
Eventually, you may get a monstrosity like Android Repo [2] though. And an Android checkout and build is pushing 1TB these days.
But there, perhaps, the submodule idea wins again. Replace most of the submodules with prebuilt variants, and have full source + building only for the module of interest.
One fundamental way it differs is atomic commits. You can't change something in repo A and subsubrepo XYZ in a single pull.
A monorepo allows you to do things like atomic commits to arbitrary pairs of files in the repo, which among other things opens up the possibility of enforcing single-version of libraries, which in turn removes a whole class of diamond dependency issues.
There's other benefits, but imo it's probably not worth it for most companies because of the staggering number of things it breaks in the developer toolspace once it gets large enough. Eventually you need teams of people that do nothing but make tooling to support monorepo scaling, because everything off the shelf explodes (what do you do when even perforce can't handle your repo?)
For example, at Google we have a team of people who do nothing but, effectively, recreate the cross referencing and jump-to-def everyone else gets for "free" from IntelliJ / VS intellisence, etc. (We do other stuff too, but that's a fair paraphrase). And on top of that the team really only exists because Steve Yegge is a Force to be Reckoned With, otherwise we might still be flailing around without jump to def, idk.
> That's especially valuable if you need to refactor something and need to adjust it's dependencies.
And yes exactly, being able to change a library and all of its callers at the same time is pretty handy.
Things like moving a file from one git submodule to another is more cumbersome than just `mv foo dir/bar`. That means your directory structure is in practice tightly coupled to the tree of git projects.
Also, since any of the git sub-repos can be branched, the chaos of merging development branches seems like it gets even more complicated in a submodule architecture.
It may be possible to put a user interface that abstracts away the submodule architecture and forces everything to live on HEAD. But at that point it might be easier to just provide a git-like UI to a centralized VCS.
Submodules, like LFS, are a great idea that suck in practice because they're bolted on to git to avoid compromising the purity of the base project.
I'm very curious how it performs compared to EdenFS: https://github.com/facebook/sapling/blob/main/eden/fs/docs/O...
Dev Drive seems to be a special type of disk volume with higher reliability or something for dev-related workloads.
EdenFS is "Facebook's CITC", i.e. a virtual filesystem view into a remote version-control system.
"Monorepo" is a culture around having a single branch with a single lineage, and not developing anything in any isolation greater than a single developer's workstation.
I agree that Git is not very adequate for large monorepos, but I'd say that most open source projects are on Git, and most of them are trivial monorepos.
Microsoft's blog posts have indicated a move to use something as close to off-the-shelf git as possible, though. They say they've stopped using VFS much and are instead more often relying on sparse checkouts. They've upstreamed a lot of patches into git itself, and maintain their own git fork but the fork distance is generally shrinking as those patches upstream.
Submodules are essentially broken with no way to fix them. It was a good idea that never took off.
I remember it "only" being somewhere around 200Gb.
They upstreamed almost everything. The last version of "scalar" was mostly just a configuration tool for sparse checkout "cones" which needed a bit of hand-holding, and that is easier to configure in git itself now, or so I hear.
So why to store them under VCS at first place? I think they are doing it wrong.
Im always strict what is going into VCS. Anything that is generated or preprocessed should not belong there.
Analogous to something like node_modules in web dev, for example.
The files could be gitignored and then generated both on local machines but also on the CI/build system.
I feel like often people make problems much harder than they need to be by imposing arbitrary constraints on themselves that could be avoided if they approached the problem differently.
I mean I don't know if they ever needed history but, just saying. You get certain things for free by using a filesystem / git.
I dunno but there are folks who would put anything in git. I work with someone who manages to exceed disk space of company's Gitlab instance by git adding everything. The disk is full again once a month.
I can't imagine using anything else for my current project. In fact, the only domain within which I would even consider something different would be game development. Even then, only if the total asset set is ever expected to exceed a gigabyte or so. Git is awful with large blobs. LFS is an option, but I've always felt like it was a bandaid and not a fundamental solve.
However, since then we've migrated our engineering blog from medium to a self-hosted stack, so HN doesn't link it to the previous discussion automatically.
I think that submodules are better suited for separation of concerns and performance, even while achieving the same composite structure as an equivalent monorepo?
The reasons for those things being separate projects had a history (dating to a time before Git was popular, even) and can be explained, but ultimately it doesn't matter; by the time I was around, all of those reasons ceased to exist or were simply not important.
I will personally never, ever, ever, ever allow Git submodules in any project I manage unless they are both A) extremely low traffic, so updating them constantly doesn't suck, and B) a completely external dependency that is mostly outside of my control, that cannot be managed any other way.
Save yourself hair and time and at least use worktrees instead.
So, atomicity of changes can be guaranteed, but you need to write a few more commits. However this effort of small increases of commits is far outweighed by the modularity imo.
Not remotely, as the scale of the codebase increases, the benefit of modularity goes to zero and the benefit of atomic changes increases.
Also: it's not always feasible to break up a change into smaller commits. Sometimes atomic change is the only way to do it.
The advantage of this is, that work can be done by devs on the individual modules without much knowledge of the overarching architecture, nor strong code ties into it.
Right now our persistence is done with SQL, but we could swap it with anything else, e.g. mongo, and the parent codebase wouldn't notice a thing since the submodule only returns well defined python objects.
Of course, this comes at the cost of higher number of commits as you mentioned. But in my opinion these are still cheap because they only affect trivial quantity and not brain-demanding quality.
To take a concrete example, I'm working on a codebase that houses both a Node.js server-side application and an Electron app that communicates with it (using tRPC [0]). The Electron app can directly import the API router types from the Node app, thus gaining full type safety, and whenever the backend API is changed the Electron app can be updated at the same time (or type checks in CI will fail).
If this weren't in a monorepo, you would need to first update the Node app, then pick up those changes in the Electron app. This becomes risky in the presence of automated deployment, because, if the Node app's changes accidentally introduced a breaking API change, the Electron app is now broken until the changes are picked up. In a monorepo you'd spot this scenario right away. (Mind you, there is still the issue of updating the built Electron app on the users' machines, but the point remains - you can easily imagine a JS SPA or some other downstream dependency in its place.)
[0]: https://trpc.io/
So far, we could avoid it though, by strict encapsulation.
But I definitely see the point in your example and wouldn't follow through with submodules there probably too.
It's just that in OP's link, I'm quite sceptical as the monorepo approach requires quite some heavy tweaking.
>At Canva, we made the conscious decision to adopt the monorepo pattern with its benefits and drawbacks. Since the first commit in 2012, the repository has rapidly grown alongside the product in both size and traffic
while reading it i was having trouble keeping track of where I was in the recursion, it's sort of "Xzibit A" for "yo dawg, we know you use source repositories, so check out our source repository (we keep it in our source repository) while you check out your source repository!"
Thankfully, nowadays git has one built in for some OSes, and it's much, MUCH better than watchman ever was.
Anyone know observability software are they using to visualize the GIT_TRACE details? (Or is the assumption that the UI is Olly as well?)
There's a couple of other tools like https://github.com/derrickstolee/trace2-flamechart
Something being autogenerated, or binary, doesn't mean it shouldn't be in version control. If step one of your instructions to build something from version control involve downloading a specific version of something else, then your VCS isn't doing it's job, and you're likely skirting around it to avoid limitations in the tool itself. People still use tools like P4 because they want versioned binary content that belongs in version control, or because they want to handle half a million files, and git chokes.
In my last org, we vendored our entire toolchain, including SDKs. The project setup instructions were:
- Install p4 - Sync, get coffee - Run build, get more coffee.
A disruptive thing like a compiler upgrade just works out of the box in this scenario.
It's a shame that the mantra of "do one thing well" devolves into "only support a few hundred text files on linux" with git.
I think the SHA should be in version control. The file should be reproducibly built [1], then cached on a central server.
This means that a build target like a system image could be satisfied by downloading the complete image and no intermediate files. And a change to one file in one binary will result in only a small number of intermediate files being downloaded or reproducibly built to chain up to the new system image.
This is something that's really lacking in, for example, Git.
Requiring reproducible builds to handle translations or images is a bit much. Also, if it's cached on a central server, that now means you need to be connected to that central server. If you require a connection to said central server, why not just have your source code on said server in the first place, a la p4?
I do agree that NixOS is a great idea, but personally 99% of my problems would be solved if git scaled properly.
> why not just have your source code on said server in the first place, a la p4?
That would be great. A version of git where cloning is almost a no-op, and building is downloading the package assuming you haven't changed anything.
I'm not aware of how p4 allowing this. My recollection of perforce is that I still had most source files locally.
> then have your build toolchain unzip the archive before it runs
My build toolchain shouldn't have to work around the shortcomings of my environment, IMO.
> et voila you now have this large file versioned in Git.
No, it's on a separate http server that is fetched via git lfs. Subtle, but important difference.
This is a non-issue for images and autogenerated files, since you shouldn't ever be doing a merge on them.
> breaks the concept of D in the DVCS of git.
git-annex is distributed and works well for files that will never be merged (such as images, or autogenerated files)
(There are various good reasons why you might not! But "because binary files shouldn't go in version control" is not one of them)
You vendored all your compilers/language runtimes in the source control repo of each project? Including, like, gcc or clang? WTF?
> It's a shame that the mantra of "do one thing well" devolves into "only support a few hundred text files on linux" with git.
Because the Linux kernel source tree and its history can accurately be described as "a few hundred text files".
Yeah, right.
Yep. Along with paltform SDKs, third party dependencies, precompiled binaries, non-redistributable runtimes, you name it.
Giant PSD or FBX files? 4K Textures? all of it.
Client mappings are the bread and butter of P4 (or Stream views more recently which are not as nice to work with) - you say "I don't want the path containing MacOS" if you don't want it.
> Because the Linux kernel source tree and its history can accurately be described as "a few hundred text files".
I was off by a little bit, it's ~60k. But it's still "only" 60k text files, no matter how important those text files are.
Having local commits intermingled with an upstream code base can make for really hairy upgrades, but I guess every situation is slightly different.
Well we don't put them in git, we put them in perforce because git keels over if you try and stuff 10GB of binaries into it once every few months.
I think the real question is the other way around though, why _not_ use git for versioning when that's what it's supposed to be for? Why do I have to verison some things with git, and others with npm/go build/pip/vcpkg/cargo/whatever?
There are plenty of cases when including generated files is appropriate. It has many advantages over not doing that - probably the biggest are
* Code review is much easier because you can see the effect on the output.
* It's easier to find the generated files because they're next to the rest of your code. IDEs like it much more too.
In fact the upsides are so great and the downsides so minimal I would say it should be the default option as long as:
* The generated files are not huge.
* The generated files are always the same.
Even when they are huge it might still be a good idea, but you can put the files in a submodule or LFS. I do that for a project that has a really difficult to install generator so users don't need to install it.
That said, if the autogenerated output is stable, it's fine. After all, in a sense, compiling your code is also a kind of autogenerating and few people will advocate for keeping compiled code in git.
- Makes it easy to develop applications and libraries together in a single branch
- Similarly, makes it easy to make a breaking change to a library, then change all clients of said library, in a single branch
- And because of the above, makes it easy to keep all dependencies on internal libs at the latest version, which can greatly reduce all sorts of “dependency hell” issues
- Generally makes integration testing a bit easier
The downside is you have to invest a lot more time in tooling, keeping both local and CI builds fast. And even with that tooling, builds won’t be as fast as they trivially are with multi-repo. But if you do invest that time in tooling, you can generally get them fast enough, and then reap the other benefits for a very productive dev experience.
Have done both monorepo and multi-repo at different, decent sized companies. Both have their pros/cons.
Because if so, why not just use one repo per app/library/tool? Sure, if you have a cluster of things that all depend on each other, or a cluster of things that typically is needed in bulk, by all means, put those in a single repo.
But putting literally all your code in a single repo is not a very sane technical choice, is it?
But I like to look at the problem from another perspective. Why _not_ use a single repo. The only real reason would be to work around technical challenges with your source control of choice, not because having everything tracked together is inherently bad.
One of the biggest advantages is that there is no version chasing and dependency questions. At commit X, everything works consistently. No debating about whether this or that dependency is out of sync.
The article says. But if Google does it it must be good.