Google _is_ a monorepo
Google _is_ a monorepo
If you're small, everything will fit nicely in a monorepo.
If you're large, you'll want lots of repos. There aren't really any off the shelf monorepo options that scale super well, so using a bunch of small repos is a great way to deal with the problem. Plus, you probably don't have a full time staff babysitting the source repos, so you want some isolation. If someone in another org is breaking stuff left and right, you don't want the other orgs to be affected.
If you're GIGANTIC, monorepos are a pretty great option again. You'll probably have to build your own and then have a full time group of people maintain it, but that's not a huge problem for you because you're a gigantic tech company. You can set up an elaborate build system that takes advantage of the fact that the entire system is versioned together, which can let you almost completely eliminate version dependency hell. You can customize all of your tools to understand the rules for your new system. It's a huge undertaking, but it pays off because you've got a hundred thousand software engineers.
How can you say this when Perforce on a single machine took Google to absolutely terrifying scale? There is no chance that your mid-sized software company will even slightly tax the abilities of Perforce.
What I believe you meant was there aren't really any good options to make git tolerable for non-trivial projects, and with that I wholeheartedly agree. And that's why these threads are so tiresome: they always boil down to people talking about what git can and cannot do.
https://www.perforce.com/sites/default/files/still-all-one-s...
That's not an endorsement for a company as big as Google was then looking for an easy, off the shelf solution. It'd probably be just fine for a company of hundreds, but so would git.
On the other hand, if you play to its strengths, it's probably a great choice. Maybe a team of dozens of content developers checking in large assets for videogames. Perfectly great use case for Perforce.
It's just a boring semantic point that I'm making, that "non-trivial" was a hyperbolic word choice.
Sparse checkouts, shallow fetches/clones, partial clones, etc allow you to work with an egregiously large repository without needing to ever actually mess with the whole thing. Most existing build tooling can be made to work with these features pretty easily however some tools are easier than others.
Enforcing clean commits avoids the issues with keeping track of individual project histories and past that the existing git tooling largely already supports filtering commits to only expose commits relevant to specific directories/pathspecs.
---
The only time I really see an organisation outgrowing a monorepo is if the org is incapable of or unwilling to maintain strict development and integration policies.
Also worth noting because I don't see it mentioned enough but not everything has to be in the same monorepo. Putting all closely related products and libraries in the same monorepo is kosher but there's little reason for unrelated parts of an org's software to all be in the same monorepo. So what might be 50-200 independent projects/repos could be 3-20 monorepos with occassional dependencies on specific projects in the other monorepos.
All signs for the rest of the world point to the opposite conclusion: unless you're Google-scale, you don't have Google level resources. Google has more engineers working on developer experience than most companies will ever have period.
And monorepos work best when workflows are carefully thought out with clever application specific tooling.
-
I think the author is probably working on a 1-10 developer project (and I'm leaning towards 1) and has confused the convenience with having things in reach when the entire system fits in your mind with the general benefits of a monorepo.
I also wonder if they read any of the letters they linked too...
I feel like this was true ~5 years ago, but these days the tooling around scaling monorepos is safely supporting O(100) developers without a lot of overhead.
> And monorepos work best when workflows are carefully thought out with clever application specific tooling.
I don't see any meaningful distinction with how well thought out workflows need to be between mono and polyrepo.
You completely failed to parse the sentence. The point being made isn't "well thought workflows are only for monorepos", that applies to the "needing clever application specific tooling" part.
Vendoring/Versioning for discrete packages is a heavily invested in problem space for most tech stacks you'll come across. But if you build a monorepo and don't end up with a build system that takes on those responsibilities you end something that doesn't scale to even moderately large interconnected components.
OP has such a tiny project that I'm not convinced they're even dealing with dependencies in a traditional sense. But well before you get to Google scale, you'll run into situations where you just want to change one thing and don't want to change every single downstream dependency which normally would have been isolated via a discrete package that doesn't have to change in lockstep. And then that exact same pain starts to exist for deployments and needs to be worked around.
> I feel like this was true ~5 years ago, but these days the tooling around scaling monorepos is safely supporting O(100) developers without a lot of overhead.
Nothing about the above has changed in the last 5 years, it's kind of the ground truth of monorepos via multiple repos: You're the first person I've ever seen imply monorepos don't offload complexity to tooling, even amongst proponents.
Not updating dependencies is the equivalent of never brushing your teeth. Yes, you can ship code faster in the short term, but version skew will be a huge pain in the future. A little maintenance every day is preferable to ten root canals in a few years.
I feel obliged to point out that I work at a company that uses a monorepo, so this isn't a "never use monorepos" counter-post. Instead my points are borderline tautological:
There's a balancing of near-time sacrifice vs long-term sustainability. But you need good reasons to pick the side of the scale that historically got less resources invested into it and puts an impetus on your engineering team to adjust to the knock on effects of that disparity while still building a fledgling company.
That's a strawman: the choice is not between updating and not updating. The choice is between updating on my terms or not.
I recently updated stripe from 2.x.x to 5.x.x in one of the projects. That's several years without updates. Wouldn't it be fun if somebody was forced to update multiple projects every single time stripe ships a new minor version? And if we were to do the true monorepo, at what pace do you think stripe would be updated, if it was their responsibility to update all dependents?
Also, Amazon went through this whole thing. They have tons of tooling built up around managing different versions of external and internal dependencies and rolling them out in a distributed fashion. They are doing polyrepo at a scale that is unmatched by anyone else. And you know what they've settled on? Teams getting out of sync with the latest versions of dependencies is a Really Bad Thing, and you get barked at by a ton of systems if your software is stale on the order of days/weeks.
> Within your company you don't want every team to have to operate as a Library Vendor
But you want some teams to operate this way. And the best way to do it is by drawing boundaries at the repo level.
This is similar to monolith-services debate. Once monolith gets big enough there's benefit in breaking it down a bit. Technically nothing prevents you from keeping it modular. Except that humans just really suck at it.
> take advantage of the command economy you operate in to drive changes across the company rapidly
Driving changes across the company is a self-serving middle-manager goal. There's a reason why central planning fails at scale every single time it is attempted.
> Teams getting out of sync with the latest versions of dependencies is a Really Bad Thing
It definitely can be a bad thing. But you know what's even worse? Not having the option to get out of sync. If getting out if sync is a problem, polyrepo offers simple tooling to address it.
In practice internal teams don’t have this type of bandwidth. They need to make changes to their implementations to fix bugs, add optimizations, add critical features, and can’t afford backporting patches to the 4 versions floating around the codebase.
Repos work for open source precisely because open source libraries generally don’t have a strong coupling between implementers and users. That’s the exact opposite for internal libraries.
You don't need bandwidth to maintain backward compatibility in polyrepo. As you said yourself, you need loose coupling.
When you are breaking backward compatibility, the amount of bandwidth required to address it is the same in mono- and polyrepos (with some exceptions benefitting polyrepos).
The big difference though is whose bandwidth are we going to spend. Correct me if I'm wrong, my understanding is that at Google it's the responsibility of dependency to update dependents. E.g. if compiler team is breaking the compiler, they are also responsible for fixing all of the code that it compiles.
So you're not developing your package at your own pace, you are limited by company pace. The more popular a compiler is, the slower it is going to be developed. You're slowing down innovation for the sake of predictability. To some degree you can just throw money at the problem, which is why big companies are the only ones who can afford it.
> can’t afford backporting patches to the 4 versions floating around the codebase
Backporting happens in open-source because you don't control all your user's dependencies. Someone can be locked into a specific version of your package through another dependency, and you have no way of forcing them to upgrade. But if we're talking about internal teams, upgrading is always an option, you don't have to backport (but you still have the option, and in some cases it might make business sense).
> open source libraries generally don’t have a strong coupling between implementers and users. That’s the exact opposite for internal libraries.
I disagree. There's always plenty of opportunities for good boundaries in internal libraries.
Though I'll grant you, if you draw bad boundaries, polyrepo will have the problems you're describing. But that's the difference between those two: monorepo is slow and predictable, polyrepo is fast and risky. You can reduce polyrepo risks by hiring better engineers, you can speed up monorepo (to a certain degree) by hiring more engineers.
When there's competition, slow and predictable always loses. Partially that's why I believe Google can't develop any good products in-house: pretty much all their popular products (other than search) are acquisitions.
100% disagree. The problem of "How do I define my dependencies and have a package manager reify that into a concrete set of versioned dependencies" may be a solved problem, but the tools for tracking dependencies across many repos and driving upgrades is neolithic. About the only company I've seen that does this well is Amazon, and they have yet to sell us version sets as a service.
> OP has such a tiny project that I'm not convinced they're even dealing with dependencies in a traditional sense. But well before you get to Google scale, you'll run into situations where you just want to change one thing and don't want to change every single downstream dependency which normally would have been isolated via a discrete package that doesn't have to change in lockstep. And then that exact same pain starts to exist for deployments and needs to be worked around.
As I alluded to above, the dual of this is that getting everyone to update their dependencies is orders of magnitude more difficult when you have a polyrepo setup, even if we're working under the ideal situation where repo setup is standardize to such a degree that a person can parachute into a repo and become effective within minutes.
> Nothing about the above has changed in the last 5 years, it's kind of the ground truth of monorepos via multiple repos: You're the first person I've ever seen imply monorepos don't offload complexity to tooling, even amongst proponents.
Both polyrepo and monorepo have complexity in scaling that is handled by their tooling. The difference historically is that OSS polyrepo tooling has been better and better integrated because that's just how most things are built in any language, but that has been improving over the past ~half decade
* Bazel maintenance complexity has dropped precipitously, and many of the initial bottlenecks you hit with it have been solved in OSS.
* If you're anti-bazel, gradle and cargo monorepo support is decently good. I believe the same is true in js these days, but I don't have hands-on experience
* Services for managing monorepos like sourcegraph for codesearch or mergify for submit queue now exist that make it easy to adopt the patterns that work well at large companies.
* Microsoft and others have invested in git to improve scalability of developing against large repos.
* you have OSS tools like git-branchless that further improve the experience of working in a monorepo
There are a bunch of companies in the O(100) - O(1000) developer range that are using this stuff and it works very well.
For company's not at this tier, it isn't that hard to migrate to a monorepo and the benefits will be more immediate because the tools (eg. git) won't be screaming under the load.
(My personal 2c is that you can be well below Google-scale and still hit the limits of the common tooling when using monorepos. Canva, Stripe, and Twitter are examples)