Edit: Why am I getting downvoted for asking a question?
We use multiple repos and we have a pretty large codebase consuming a TB a day.
Edit: Why am I getting downvoted for asking a question?
We use multiple repos and we have a pretty large codebase consuming a TB a day.
The alternatives:
1. Completely de-coupled code. I.e. teams ship libraries, like a 3rd part api. This slows down development considerably, and makes re-use of code harder.
2. Keep everyone on one repo, and allow to changes stream in.
1. Slows down dev. (it suddenly becomes more beourocratic), 2. requires scaling your code versioning tool (gir/mercurial, etc) and processes associated with it.
Also, single repos, make sense in one domain (e.g. server side, ios, android, etc...). You can have different repos for different 'domains' where code doesn't intersect with each other that much.
I've found this to be a feature of polyrepos because monorepos can easily become a rat's nest of dependencies. Polyrepos make you think harder about what should really be exposed and shared.
Allowing fine grained visibility at the level of a file, package, or artifact is better than at the granularity of a repo.
(And at least at Google visibility changes required the approval of the team who you want to depend on)
Having seen "a bunch of repos" in action, I think people don't talk enough about just how awful the experience can be, and how much more work it is to manage multiple repos. As you increase the number of repos, the pain gets worse at a rate which is faster than linear. There are plenty of articles talking about how wonderful a monorepo is, just not many articles about how bad multirepo is.
NPM hell is a close approximation of the multirepo experience. Try upgrading the dependencies in a large NPM project and you'll see all sorts of problems. You might find that upgrading X breaks Y, but you have to upgrade X in order to upgrade Z, and you need to upgrade Z for some reason.
With monorepo, all the versions march forward in sync. If you fix trunk, you will probably ship it, eventually. With multirepo, you need to fight the tooling just to show (from the example above) that a patch to Y will let you upgrade your project to use a newer version of Z.
Not to mention that Google internally avoids multiple repositories (and instead runs of a single monorepo) - it's just the android/chrome/public stuff that's mostly split up.
This thinking right here, sounds like a reason why Google retires a lot more services at much higher frequency compared to other companies?
In https://thehftguy.com/2019/12/10/why-products-are-shutdown-t..., the HFT guy brought up the "5 years upgrade pain" as the reason why a service would be retired, at the time when the pain from upkeeping outweights the gain from revenue.
By regularly making backward incompatible API change, as afforded by the monorepo, while also having engineers moving freely between projects, the upgrade pain cycle becomes way shorter for Google services.
By the way, this line of thought also leaks into public facing source code, with guava being the poster child of breaking backward compatibility, althought it has learned its lesson starting with version 21.
Of course Google, Facebook etc have given back a lot to OSS. Just not extemporaneously to when they received the value. It may be years until the internal rebuild of something some guy copied from somewhere is rereleased.
Git was designed for OSS. That includes its radically transparent form of development. Then again with submodules you can easily vendor your private stuff, rather than doing things the other way around.
They do vendor, but that's for security and efficiency reasons mostly, not to avoid giving back.