We Put Half a Million Files in One Git Repository, Here’s What We Learned
canvatechblog.com
canvatechblog.com
To be fair, that could mean a _lot_ of functionality and code providing a rich, SPA, JS heavy experience. Modern JS frameworks aren't exactly known for being concise.
But still, 60 million lines of code and half a million files is a sure sign that someone said at one point "sure, we can throw all these generated files into git!". Didn't someone on the team say "hey, it takes 10 seconds to run git status, can we move this junk out and do this another way??"
Given that 70% of their repo is generated files, that discussion and the tradeoffs involved don't get nearly enough attention from OP.
Why do you assume they didn't?
Just because they arrived at a different conclusion than you that doesn't mean they didn't thought about it. I might very well mean you did not considered the tradeoffs they had to take into account, mainly because you're out of the loop.
This is, I think, very common with stuff like browsers, where you have artifact checkouts that basically include built stuff since otherwise you're sitting there compiling forever all the time.
Stuff like Bazel in theory helps with this, but tools that help with this are either super idiosyncratic about how they work (meaning hard to adopt) or outright don't work.
I mean personally I would find that pretty annoying and unclean but I like DAGs.
The likelihood of a 2000 employee software company simply not considering that they could streamline their build process is pretty slim.
All of the problems they are having are basically due to their use of a monorepo: they do explain that they made the decision early, but I wonder what are the advantages over multiple repos they are seeing that it was worth it all this trouble?
I would encourage you to do some research and keep an open mind.
The question is specifically about their usecase, since not everybody would hit the same bottlenecks as they did with git monorepos.
Iow, have they stopped and thought whether it's still worth it (eg how often do their engineers make use of the monorepo benefits like cross-project refactorings)?
By the numbers you mention, 70% of the files make a ratio of code files to translation files 1-3, so unless you only support 3-5 languages, it's definitely not one XLIFF file per source file, so I wonder at what granularity it is?
(My experience is mostly with localizations using GNU gettext tools, and you usually do a small-finite-number of PO files per project per language, where that small-finite number is exactly one for like 99% of projects)
Good job managing all that regardless of the approach :)
The only exception is our generated OpenAPI spec, because we want people to be explicit about modifying the API, and have a CI task that verifies that the API and OpenAPI spec match.
Something like: "This change will result in the following new API endpoints: ... do you wish to continue?"
What does it compare against though? Need to add more state to the CI? We kinda like the interface be part of the version control and having an audit chain that's part of the code.
Though even without that, I'm not sure how it'll even mechanically work.
It is gone now.
Docker is a congregation of technologies held together with duct tape and glue.
Eg. permissions handling is completely different on Macs with Docker Desktop from the Linux dockerd stuff: on Macs, it automatically translates user ownership for any "mounted" local storage (like your code repository), whereas on Linux user IDs and host system permissions are preserved. Have some developers use Macs and others use Linux, and attempt to do proper permissions setup (otherwise known as "don't run as root"), and you are looking for some fun debugging sessions.
No, it's not. What a wild conclusion to reach from the example you gave.
The price of letting less experienced people "go crazy" in the repo
If a new project these days commits node_modules to git, it's likely a mistake, but for legacy projects started before 2017 it was the lesser of two evils.
Edit: spelling.
Prior to lock files (and potentially after, as checked-in files are beyond trivial to modify and review and that can be worthwhile) committing dependencies in some form was basically the only reasonable way to have reproducible builds, unless you wanted to build your own package manager / lock file implementation.
Which is what Yarn did.
> it takes 10 seconds to run git status
People coming from the SVN world do not think that this is unusual or problematic. And unfortunately even recently I've seen SVN still in use at large legacy companies.For many processes I think SVN is (and has been for many many years) been an absolutely fine method of version control
> didn't find it offered any features that would significantly improve our process, but found 1000 more ways to shoot ourselves in the foot
You're not wrong about this.I really like git for the cheap branching, which encourages branching and merging often. But SVN might have cheap branching now, as another commenter implies.
I do not remember if "stat" was particularly slow, but SVN in general is slow.
svn status, even for an entire repo checkout (which is not common) is also fast.
And yeah, it has virtue of simplicity as well doing very well at narrow and shallow even though I'd love to have mercurial's feature set.
It's also rather good in the "wiki" situation since people can operate on their single files without needing to update, sync and merge.
https://www.bitquabit.com/post/unorthodocs-abandon-your-dvcs...
A fun rant, even though git has gotten better-ish at large files.
> Creating branches is virtually instantaneous. It's just a copy which is a free operation (just a link).
Copy is not a "free" operation, but a symlink is close to "free" if you're measuring disk space.What version SVN are you using? I'm certain that older SVN versions would actually copy the entire project's files, not symlinks but real copies. That would take forever and running out of disk space was a real concern.
I can perhaps imagine a large repo plus a broken svn client requiring checking out unneeded portions of trees to do a copy, but no client I've used works like that.
Hm. Another theory. Perhaps someone who knew nothing about svn and was using TortoiseSVN's Windows file manager integration was doing a Windows file manager copy, then checking that in as a "branch" with the only link being the commit message instead of using svn's copy which is free and properly links content. That would indeed be an expensive operation, and the wrong thing to do.
I have this alias in my .hgrc file fad=fastannotate -u -n -wbB --deleted
It's by the Facebook engineer Jun Wu who also made the even more awesome "absorb"
It's perfectly fine to use Git to track things other than sourcecode. In fact, right on the manpage, Git calls itself "the stupid content tracker".
I've been using Git with git-annex to track archival files with their associated metadata. We keep our data separate from our sourcecode, and segment our data into individual Git repositories for each collection. Git gives us many features that we would have had to build into our app in other ways (data integrity, fixity, etc), though this came with costs.
To my eye it probably would have been better for Canva to use multiple separate repositories instead of a monorepo, but I'm not them and their use-case is not mine.
The content of these translation files are snapshot in time aligned with the text in our product so simply removing them we would lose all the changes made to translations each time texts are changed.
Sorry about that, hope this clears it up!
Edit: for more information on translation and xlf files, we have another blog post all about them https://canvatechblog.com/how-to-design-in-every-language-at...
Obviously, the problems they are solving (and admitting to solving) are due to their dedication to the monorepo.
With all the effort spent on working around the drawbacks, I really wonder what advantages they are seeing that make it worth their while?
If we are being pedantic, git was not designed to host multiple projects in a single repo (otherwise, git would have been a subdirectory in the kernel tree). But tools are made without knowing how they'll be used, and that's ok, so I wouldn't stress on what the purpose for monorepo was, but how it's used and what value it brings.
Translations do not look like code to me. Rather human generated artifacts.
What I understand is there is a strong tie/match needed between versions of these translations files and the code itself, so I believe this is where having all in the same repo would make sense, having the translators update those file when code has been modified...
> English source strings used in our frontend code live in Typescript files (.messages.ts). Source strings used in our backend code live in Java files (Message.java). Our internationalization (i18n) pipeline converts this into a series of XLIFF files (ending in .xlf), with one file for each locale. All these files live in the repository, but the translated .xlf files should never be modified by hand since they are updated automatically when strings get translated.
Isn't it cache?
I have this argument with people all the time and the conclusion is always like: "it is too hard to integrate the generator with the build system so we check them in".
The big problem with generated files is merge conflicts. How do you resolve a merge conflict on generated files. especially if they are binary.
For example, all the localisation files could live in a separate project (if we accept the need to commit them at all). Some tools would be needed to deal with the inevitable problem that developer working sets would not align with project boundaries, but that seems like an easier job than making git scale while maintaining response times.
Not that easy to get this to work on multiple repos.
Monorepos, like basically all solutions, solves some problems and introduces new ones which you didn't have before. It depends on each individual case which drawbacks are more worthwhile.
Only the code that is closely related – read/modified/built together frequently – should live in the same repo. If two pieces of code that don't have much to do with each other (that is, they are not read/modified/built by a developer in a single developer workflow frequently) live in the same repo, then they are just being a burden to the overall development lifecycle of devs who work on those disjoint sets of code.
The unrelated code in the same repo is a distraction to the developers who checkout that code as it costs storage space, iops, cpu cycles and network bandwidth to lug that code around, load/index in IDEs, track changes, build and discard dependency graphs by build/dependency systems etc. Then, to deal with these issues more complexities are incurred. Instead, it is better to optimise for the common case and deal with the complexity only for the rare cases.
No, it may not. Perhaps occasionally it does. That is a bug that you must fix - a pipeline that takes even 30 minutes is horrifically slow.
Multirepo management is extremely frustrating compared to "it's all in the same folder".
Surely if it is an advantage to rename once in a ginormous, single code base there must also be leaky abstractions, poorly defined interfaces, god objects, etc present at the same time?
Whenever I find I need to rename anything across domains, it's a matter of updating the "core" repository and then just pulling the newest version.
Which means that you've got to do independent backwards-compatible changes anyway, and that for anything remotely complex, you are better off with separate branches (and PR/MRs) anyway.
Monorepos mostly have a benefit for trivial changes across all repos (eg. we've decided to rename our "Shop" to "Shoppe"), where it doesn't really take much to explain with multiple repos, but is mostly a lot of tedious work to get all the PRs up and such.
Sometimes you have to ship a feature. Shipping that requires changing 3 parts of your app. A lot of times that _entire_ set of changes is less than 100 lines of code.
Having a full vision of what is being accomplished across your system in one go is very helpful for reviewing code! It justifies the need for changes, makes it easier to propose alternatives, and makes the go/no-go operation much more straightforward.
At a smaller scale, you often see the idea of splitting frontend and backend into separate repos. Of course you can ship an API and then ship the changes to the frontend. But for a lot of trivial stuff, just having both lets you actually see API usage in practice.
I think this is much more applicable for companies under 100 people though. When you get super large you're going to put into place a superstructure that will cause these splits anyways.
Most projects start out as monoliths (which is good) and splitting up on this axis is unfortunately very hard/costly.
Unfortunately it's hard for me to recommend Bazel, it's such an uphill climb to get things working within that system.
ISTM that the complexity of managing any repo will be bounded by the size of that repo; a monorepo, being unbounded in size, will, in time, become arbitrarily complex to manage.
While a multirepo might occasionally require developers to apply changes to more than one repo at a time, I’ve never found this to be much more than a minor inconvenience; one that could be solved readily with simple tooling, if we had ever felt that the “problem” was even worth solving.
Multirepo also allows you to roll out that change incrementally instead of big banging all the time.
Meanwhile reviewers don't have context about changes, so it's easier to get lost in the weeds.
It's not always this, of course. But I think that way too many tools are based on "repo" being the largest element, so things like cross-repo review are just miserable.
Then there's version management. Do all your repos use the same versioning scheme? "They should", but in the real world, they sometimes don't. Whereas if you only have 1 repo, you are guaranteed 1 versioning scheme, and 1 version for everything.
How do you know which version of what correlates to what else? With N repos, do you maintain a DAG which maps every version of every repo to every other repo, so when you revert a change from 1 repo, you can go back in history and revert all the other repos to their versions from the same time? Most people do not, so reverting a change leads to regressions. With a multirepo, there only is one version of everything and everything is in lock-step with everything else, so you can either revert a single change, or do an entire rollback of everything, with 1 command.
How do you deploy changes? If each repo has an independent deployment process (if your repos even have a deployment process that isn't just waiting for Phil to do something from his laptop), are you going to deploy each one at a time, or all at once? What if one of them fails? How do you find out when they've all passed and deployed successfully? Pull up 5 different CI results in your browser every couple hours, and when one fails, go ask that team to fix something? If you only have 1 repo, there is 1 deploy process (though different jobs) and merging triggers exactly what needs to happen in exactly the right order.
The reason people use multirepos is they don't want to build a fully automated CI/CD pipeline. They don't want to add tests and quality gates, they don't want to set up a deployment system that can handle all the code. They just want to keep their own snowflake repo of code and deal with everything via manual toil. But at scale (not "Google scale", but just "We have 6 different teams working on one product" scale) it becomes incredibly wasteful to have all these divergent processes and not enough automation. Multirepo wastes time, adds complexity, and introduces errors.
There's a couple of different views of this:
If the subsets are overlapping then the monorepo has had great value. Let's say you've got modules A, B, C and D. Dev 1 is interested in A and B, Dev 2 is interested in B and C, etc. In a multirepo world you have to draw a line somewhere and if someone has concerns overlapping that line then they're going to have to play the 'updating two projects with versioning' game.
The other way of looking at it is "data model" vs "presentation". Too often with git we confuse the two. Spare checkout is a way of presenting the relevant subset to each user. It is nice to be able to consider that separately from whether we want to store all that data together.
That's the wrong way to split files: it's as if you said let's split a monorepo so all the .sh files are in one repo, all the Makefiles are in another, all the .py files in yet another...
What you want is to split into "natural" repositories instead. Having 50 or 150 localisation files in an otherwise 40-file repo is not a big deal for anyone. Of course, how the split happens would have an outsized influence on the ergonomics.
Also note that localisation files are tightly linked to source code (the way they use them, similar to GNU gettext model, though they do use XLIFF): you put English strings in the source code, and when you change them (reword, fix typos, or outright change them), all translations need to get their English version updated and translations potentially needing updates marked as such. In short, they are managing their translations as source code (even if translators would be using translation tools akin to IDEs for development).
If submodules worked.
[0] https://www.plasticscm.com/documentation/xlinks/plastic-scm-...
There is a massive loss in dependency management if you move to multiple repos.
Do polyrepo build systems exist that give you the same capabilities as bazel? Particularly with regard to querying your dependency graph.
Also, I looked up .xlf files and I still don't understand. It's xml, that part makes sense, but it's basically a config file? To tell what process to read which files?
Also, I've heard of Canva, but had no idea they were this big/ubiquitous/whatever.... and learning about pseudo localization is interesting too. And the graph for lines of code looks pretty exponential, maybe it's common up to a point, but if it continues at that rate, it will be infinite by about 2026 (okay, I just made that number up, but you get the idea)
The comments in this thread from the devs have been informative as well.
I'm always surprised at the fact that almost every product has more lines of code than the entire Linux repo. The scale of these products is astounding.
There is also '-merge', which will cause git to not attempt to merge the contents, but just ask you to pick a side.
The challenge however is then verifying the contents of these files in things like merge requests.
Is there a reason why that type of file couldn't be better place into an artifact repository, or just generated and consumed in CI as part of generating a final build output?
This is not surprising at all. In fact, it's quite standard to commit string translations. Just because you can run the code generation/string replacement step as part of the build that does not mean it's a good idea to generate everything from scratch at every single build.
String translations hardly change once they are introduced, running the build step takes significant amounts of time, and if anything fails then your product can break in critical and hard to notice ways.
Additionally, to respond to your comment, if string translations don't change much then it may be possible to push them out as an internal 3rd-party library, and then they're even quicker to build.
You're missing the point. Storing translated files is caching things in git, and it is not a bad idea. It's a standard practice that saves your neck.
You either place faith on a build step working deterministically when it was not designed to work like that, or you track your generated files in your version control system.
If you decide to put faith on your ability to run deterministic builds with a potentially non-deterministic system, you waste minutes with each build regenerating files that you could very well have checked out and in the process risk sneaking in hard to track bugs. Then you need to have internationalization test steps for each localization running as part of your integration tests to verify if your build worked, which consume even more resources.
Or... you stash them in git?
You use git to track changes, regardless of where they came from. Just because you place faith in some build step to always work deterministically that does not mean you are following a good practice and everyone else around you is wrong.
I'm sorry, what? Why would a build not work deterministically?
> If you decide to put faith on your ability to run deterministic builds with a potentially non-deterministic system
If your build is non-deterministic, how can you have any faith in the binaries it produces? You would have much larger problems in that case.
> You use git to track changes, regardless of where they came from
You probably don't want to do that if it is 70% of your codebase and slows down all your developer's git.
> Then you need to have internationalization test steps for each localization running as part of your integration tests to verify if your build worked
I'm convinced you've never used a build system before. Your build should fail if required files are missing. Downloading translation files at build time from some artefact repository vs storing them in git is how a lot of companies do it.
Because they don't and never did?
Do you understand build systems and individual tools were not designed to ensure deterministic behavior?
https://reproducible-builds.org/docs/deterministic-build-sys...
Anyone with any professional experience developing software can tell you countless war stories involving bugs that popped up when building the exact same project separate times. What leads you to believe that translations are any different? In fact, more often than not we see unexpected changes during translation update steps.
> If your build is non-deterministic, how can you have any faith in the binaries it produces?
First of all, all builds are not deterministic by default.
To start to come close to get a deterministic build, you need to do all your own legwork after doing all your homework.
Did you ever did any sort of this work? You didn't, didn't you? You're not looking and are instead just placing blind faith on stuff continuing to work by coincidence, aren't you?
> You probably don't want to do that (...)
Yes, I do. Anyone with their head on their shoulders wants to do that. It's either that or waste time tracking bugs that you allowed to go to production. Do you want to waste your time hunting down easily avoidable and hard to track bugs? Most of the professional world doesn't.
If one has to be more granular than that, and have versioning and verification against the repository, they can still store the multiple versions on another service and store the hashes on git. Even though I'm not a fan of this for translation (especially if you have lots of languages/lots of strings), since there's an advantage of decoupling the translation process from the development process.
The problem with storing those files on git is that it can cause more problems, including developer experience issues.
It depends on how much you're storing on git. Some CSS files? Fine. 70% of files of the project, like in this case, slowing down everyone's workflow? Definitely not.
You're also doing that everywhere else. How do you think anything works? Why do you think Git is deterministic somehow? Why more so than including some files in a build?
No reason at all, but when you need the files during development, and testing, and CI, and in production, and you don't want those things to fail when your artefact repo or source of data is down, then putting the latest versions in git makes sense.
The cost of having them in the repo is a tiny bit more complexity in your git workflow and config. The benefit is being able to access those files everywhere you access the code. It seems like a no-brainer to me.
This adds yet another moving part to the system, and another place things can go wrong.
> generated and consumed in CI as part of generating a final build output
This can get quite slow, and on larger projects you have to expend a lot of effort to keep build times reasonable.
Also, if you're serving a library for public consumption, you generally don't want to add the burden of extra build steps for the user to follow before they can use it. If it can all be automated to the point of invisibility to the user that's fine, but often it can't.
I think that this is not a sensible definition of a generated file. A more sensible definition is that a generated file is created automatically from some source, which is not user input (i.e. an other file). This means generated files do not need to be kept under git, as long as their source is checked in.
Translations files, even if they are not created with a plain text editor but with some other tool that handles the XML layer, are clearly not generated, as long as the translation is done by a human.
This is a very typical workflow. Most people are not out there modifying xlf files by opening them in a text editor. For a start, translations usually aren't done by developers.
(Huge shoutout to Lokalise btw. I can highly recommend it. It makes building a multi-lingual app across different platforms so much easier.)
You opted to keep translation files out of version control; you could also keep images there, or source files. All this stuff is the (pretty direct) output of non-deterministic human intervention.
(BTW, how do you build an old version of your application? Is lokalise able to give you the appropriate translations for a specific git commit / app version?)
Or, in more words: The format of the files is just the representation on disk - it’s not directly connected to how the files are generated or edited. XML files can be written by hand with suitable editor support.
For these translation files, I’d imagine there may be occasional work to modify them even after they are initially generated.
You're using git as a cache. You don't need to version a cache.
If it was a physical product, you couldn't keep making it bigger and more complex ad infinitum, because making a physical thing bigger takes more material, and bounded physical resources would be consumed. With software, it's all just bits, and computers can hold a lot of bits.
This leads to bigger problems than just git running slowly.
> running these commands multiple times a day reduces the total productive time engineers have every day
I love the attention paid to this. Often opportunities to prioritise seemingly small efficiency gains are neglected.
At 10 seconds per command, an engineer that uses git status 50 times per day spends ~10 minutes per day waiting; an entire work week per year!! Well above the threshold warranting optimisation, and that doesn't even factor in distractions and context switching.
Then if I have a command that will take a while, like a stupidly long `git status`, I do `git status && beep`.
This is something that's really a degradation of the newer generations of engineers, since I clearly remember the time where these "somethings" would never take less than a couple minutes, and people did not immediately flee to their nearest distraction, but actually planned their time around it. In fact, if you go further back, these "somethings" would have taken hours, and the older generations still got work done.
First thought is why not to zip/tar away all of these "convenience" files per generation and add a line into build/install script to unpack them after checkout?
Additionally, add the .xlf into .gitignore to exclude them from untracked.
Noone cares to diff them as long as their contents is consistent with the checkout. Text compresses quite efficiently so this should not introduce any unreasonable build/install delays.
The mentioned .zip file is to be kept in the repo. Instead of a whatever number of individual .xlf files per generation, these would get zipped together (say, 'assets/xlf.zip') before the commit and the resulting .zip added to the commit.
Similarly, when reverting or on a checkout, it's the .zip that gets checked out and then the .xlf files are unpacked.
The packing/unpacking could be done by the same process that handles the .xlf generation (??build).
Also this may be automated by git-hooks, though it's more natural to handle the packing of assets during the build stage.
Or, at the very least, once your git monorepo reaches a certain size, you should either split it up, or switch to something that handles monorepos better. Even if that something is e.g. a virtual filesystem on top of git.
Perforce's branches are a disaster and should have been deprecated years ago and replaced with streams a decade ago. Branches are _incredibly_ slow, they're straight up copies, and they are pretty much isolated from each other. Streams are an improvement, but still are very primitive compared to git's branches - the change tracking across streams is poor, and the enforced hierarchy has too many escape hatches that can make a gigantic mess. Streams are loosely enforced with views which can't be customised per workspace, a major regression from branches. In practice, every team I've worked on has had a "convert merge to edit" style action to fix perforce's messed up idea of a merge. Stream switching is also dog slow (on my last project, it was quicker to delete the 150GB workspace, and re-sync than have perforce actually figure out what had changed).
Perforce is eye wateringly expensive, and very difficult to license - licenses are 4 figures per seat per year for medium sized businesses (and close to 4 figures for small companies), and maintaining a p4 server is genuine work. The hosted offerings of perforce (assembla only - https://get.assembla.com/pricing/) are very lackluster, and very limited (no triggers unless you pay the "contact us" pricing).
My experience in large teams was that running perforce at scale involves not using some/many of the features, and that actually keeping it running well pretty much requires an active support contract from perforce.
All that said, P4V is by far the best GUI client to any VCS, it _handles_ bianry files, and p4 sync's performance makes git clone feel like you're working on a 56k modem. It's just when you want to do anything other than submit or sync, the wheels fall off the track.
That's because P4's branches are a pale imitation of what branches can be (and to be fair to P4, branches long predate git and they can't exactly up and change the behaviour, however there's really no excuse for what they did with streams. They bolted a loosely enforced hierarchy onto the existing branch system, created a split in the tooling, and shipped something that has as many footguns as problems it solves.)
> People who want to branch in Perforce are often trying to bring a git mindset to a different tool. With a trunk-based edit/sync/submit workflow where you have different p4 clients for your different projects (what you would use various branches for in git) you need not branch.
"often" is a very nebulous adjective, and a naive view of what git branches do. Perforce and a task branch based workflow is a terrible idea, yes. If you want to do the PR based flow that github and gitlab encourage, you're going to have a bad time. P4's shelves are an excellent tool, but they encourage ad-hoc and self managed version control. Shelves to bring changes across streams (if you're using them), Shelves to share a WIP or a quick change with someone else, and iterating back and forth with shelves with v1 v2 v3 etc in them, shelves for temporary debugging code/non prod features/work in progress feature that's ticking along in the background.
Again, I'm not saying perforce has no place, but git's branches are a force multipler, even with trunk based development.
Eden's equivalent of 'git status' should run almost instantaneous, as checkouts are hosted by a virtual file system (FUSE) that tracks changes.
Does anyone have any good resources for why and how best to implement a monorepo?
With multiple repos it's harder for teams to share code and collaborate. Each team has a repo that becomes a little fiefdom where they are oblivious to who is using their code and how they're using it. Suddenly they'll push out what they think is an innocuous refactor and inadvertently break core functionality other teams took a dependency on for better or worse.
So what happens is the team with a dependency now copies the old code into their repo and take on all the extra burden of maintaining this old version, trying to backport fixes, etc. It becomes an enormous mess and time sink. No one ever has time to go back and fix things, and when they eventually are forced to do so it costs more time and effort than it would have taken to do it right from the start. You'll also run into horrible versioning problems where you're stuck on old version X but depend on widget foo which needs current version Y of that dependency.
You might say well bad on that team they should have engaged the product managers, made sure their dependencies and usage were well tracked with them, been looped in the process of changes, etc... but in the real world when your boss says X feature needs to be shipped in a few days all of that process goes out the window.
1) https://trunkbaseddevelopment.com/monorepos/
2) and (but I don't know it as well) https://monorepo.tools/
And if you can employ enough engineers to break git, you can probably afford a team to work on scaling git.
1. Standardized tooling
2. Fewer issues related to dependencies
3. Hermetic tests
4. Reduced code duplication and easier code sharing
disclaimer: canva staff working on source control
In a monorepo with the right tooling, I can make a branch where I delete the API and get pretty immediate feedback as to every module I have broken. From there I can update all the call sites and I also know which teams/engineers I should give a heads-up to.
In a multi-repo world, this is much more difficult. Even learning what all the reverse dependencies are could be a challenge. Most likely, other teams have to pull in my changes on their own schedule, and other teams have very different incentives than my own. The cost of mistakes is therefore higher because they are more difficult to undo.
Monorepos are very important if you want to empower engineers to achieve broad changes across the org. Some people fundamentally disagree with this methodology. And a fair number of people just don’t care enough - marking the old thing deprecated, calling it a day, and letting it rot for eternity is good enough for them. These are the same people who you’ll find saying “not my job” a lot, in my experience.
What I have seen as a real problem, time and time again, is having trouble locating all usages of an API in a multi-repo scenario.
Anyone who fucks this up in a monoreppo will probably fuck it up worse with multiple repos.
We have escape hatches, which I mainly use when deleting code.
Whether this is a pattern or anti-pattern depends if you want a single engineer being able to change the entire architecture to “just ship it” or you if you value conceptual integrity more.
A few things:
- your build system will need more code
to deal with a repo-mega-forest than
a megamonorepo
- code indexers may not be able to see
cross-repo dependencies
etc.You'll have pain no matter what. My preference would be for a megamonorepo approach to scale properly. That means partial/sparse cloning, as well as shallow cloning, and also all the hacks that are supposed to make git-status and git-log (and git-blame, and...) fast in partial clones.
... and they make up 70% of their repo
Why would you include generated file in a repo?
Do they take to long to remake?
[EDIT]: especially given the fact that they're using bazel which is supposed to be the bee's knees of build system?
Reason #1. Monorepos go against single-team ownership principles
Reason #2. Monorepos encourage bad practices involving massive refactoring
Reason #3: Small repositories are better than large ones
[1] https://www.infoworld.com/article/3638860/the-case-against-m...
> In Google’s case, more than 45,000 changes are made to its monorepo every day. This code management becomes an exponential problem in overhead as the number of developers of an application grows, and the number of components within the application expands.
Speaking from experience, the overhead of code management is far less at Google than any other place I've worked at, even on projects with just a couple hundred lines of code. If this author is imagining a monorepo means each engineer having to constantly check out terabytes of code to make any change and race with other developers to merge to HEAD, they don't really understand the landscape well enough to write an article criticizing monorepos (which certainly have real cons, particularly with the source control tools that are available to most companies).
Of course, you need actual tooling and infrastructure for that beyond what git offers.
One person's massive refactoring is another's tech debt reduction
One person's multirepo is another's inability to find that broken bit of upstream code
But in terms of the reasons in "The case against monorepos":
> Reason #1. Monorepos go against single-team ownership principles
I think it's up for debate about whether or not single-team ownership is desirable. But even if it is, I don't see the difference. Just have teams own their directories within the monorepo.
If you want to enforce that a team owns a particular part of the repo, just put some rules into the code review tool to ensure that a change to that component can't be merged without someone from that team reviewing/approving it.
> Reason #2. Monorepos encourage bad practices involving massive refactoring
Again, this just appears to be the author's opinion that massive refactoring is somehow problematic, without providing much of an argument against it. If you're changing service APIs, then sure, you need to be careful about the order in which you roll out changes. But if you're changing the APIs of shared libraries, then being able to do a single large refactoring is absolutely valuable.
> Reason #3: Small repositories are better than large ones
This is just a tooling issue. You can still separate projects by directories, and if your tools allow you to just check out a subdirectory, or if you have some sort of virtual filesystem on top of your source control, then you don't have to pay the penalty of pulling down the entire repository.
> This is just a tooling issue. You can still separate projects by directories, and if your tools allow you to just check out a subdirectory, ...
Even a forest of repos has tooling issues: your build, code review, and code indexing tools will need to support the forest, and that's a lot of work. It's probably comparable to the work needed to make a monorepo work.
The alternative to the monorepo isn't all rosy.
The sentiment on this thread, if it is indicative of the greater talent pool, suggests this blog post is having the complete opposite effect.
I remember when Uber was proud of their thousands of repos. Here it’s the 60 million lines of code. It’s not just red flags, but seems like stuff that might get leaked to Programming Horrors / WTFs.
They wrote a blog post on how clever they think all their workarounds are, at least one of which involves sparse-checkout -- which is perilously close to chopping up your monorepo into several, while still pretending monorepo is fine.
I feel like somebody's job and/or ego is heavily invested in keeping things as they are, even if it demonstrably does not scale to their needs, and the solution is blindingly obvious to even a casual observer.
That is institutional insanity.
Certainly I wouldn't have decided to put all that junk into a single git repo from the start, since I'm not an idiot. But, even if I did, or I inherited something like that, I would fix it, not double down as they're doing.
Depending on your use cases a zip file can work, e.g. Python packages can be imported from within a zip file, and the standard lib is distributed that way.
We have the same issue at $dayjob, the repo is quite large and 80% of it is translation data. Even if they compress ridiculously well (99% last I checked) the number of translation files and the number of exports makes them the vast majority of the cost.
However that structure remains a convenient nuisance, and more importantly removing them would really only be useful if we rewrote the entire repository, which breaks all working copies.
There’s been a task in a wishlist for years now, but the business incentive just isn’t there.
Exit: actually the translation files are 80% of the working copy, they’re closer to 90% of the repo, and on the far side.
Not to mention transgressions often have their own lifecycle e.g. is common to find translatable strings which are untranslated or incorrectly translated, and want to ship updates independently from the software’s.
As such keeping the translations outside the source is also perfectly defensible.
Does this data change as often as the code does? If not then get it out of the repo.
Windows 10 has about 60 million.
I doubt that any other software project has as many.
How would you even end up with 500K source files in the repo?
I'd be somewhat intimidated on my first day of work if I cloned a repo with that much sloc...
It's Thursday and you wanna release a hotfix - but nope - all builds are failing because git can't cope with junk and your CI can't proceed. You then have to call BB instead and ask them to clean up, hoping they'll do it within the same day. Ah, exciting times.
Technically it has modes which need to scan all commits in the current branch as well, like the ones that grep for certain changes.
Otherwise you are risking to end up in a situation when hundreds of your engineers have to spend tens of minutes every time they need to simply push their code changes.
https://www.perforce.com/sites/default/files/still-all-one-s...
So advice would be: start small but think big :)
>we found that .xlf files made up almost 70% of the total number of files
If one xlf file is used to keep the translation for 1 language, how the heck do you have so many then ? Makes no sense.
If you have many xlf files for 1 language, may God save your soul.
5 components, 5 languages, would be 25 xlf files.
This is a part of most monorepos. I figure having only a single build artifact is rare.
Runs forever even on small repos
I think it's a cool company, with a good product, but they're WAY too small to be having a "source control team" on staff. That's at least 1.2MM a year salary / benefits cost.
disclaimer: canva staff working on (for now) source control
So I assume 5 people working specifically on Git performance ;-P
That's the entire Source Control team (of 5 people) today. Among other things, they work on improving Git performance ;)
(Improved testing practices is often second, but that’s not a low-hanging fruit. It’s a big, juicy fruit that’s waaaaay up in the canopy. And then comes joint product decision-making within the team, which may as well be on another planet entirely.)
Because then they won't need a 5-man team to handle the fallout any more.