A monorepo misconception – atomic cross-project commits
snellman.net
snellman.net
If the change affects the public interface of a service, then there's no option but to make your changes backward-compatible.
Not necessarily; you can accept downtime/breakage instead. That is always an option!
But, well, I also expect nobody's life to depend on it. There would be a short window between people getting into that situation and they not have any life to depend on anything.
This is the big downside of monorepos: they strongly encourage tight coupling and poor modularity.
No, that's not true. Why would you say that?
You should of course use good discipline to ensure that doesn't happen. Compared to mutli-repo it is a lot easier to violate coupling and modularity and not be detected. Anyone who is using a monorepo needs to be aware of this downside and deal with it. There are other downsides of multi-repo, and those dealing with them need to be aware of those and mitigate them. There is no perfect answer, just compromises.
That's why tools like Bazel are strict about visibility and put more friction and explicitness on those sorts of things. But this tends to not be the first thing at the top of people's minds when starting a new project... so in the monorepos I've worked on, it's never been noticed until it's too late to easily fix.
You don't even need to be that dogmatic to make this work either -- simply stipulating backwards compatibility between the two previous deploys should be sufficient.
The better version of this is simply versioning your backend and frontend but I've never been that fancy.
e.g: If your smallest deployable unit is a Kubernetes pod, and all your affected applications live in that pod, you can treat it as a private change.
It's not about migrating APIs or coordinating deployments. That's an impossible problem to solve with your repo. It's to update libraries and shared code uniformly and patch dependencies (eg. for vulns) all in one go.
Imagine updating Guava 1.0 -> 2.0. Either you require each team to do this independently over the course of several months with no coordination, or in a monorepo, one person can update every single project and service with relative ease.
Let's say there's an npm vuln in leftpad 5.0. You can update everything to leftpad 5.0.1 at once and know that everything has been updated. Then you just tell teams to deploy. (Caveat: this doesn't really work as cleanly for a dynamically typed language like javascript, but it's a world wonder in a language like Java.)
I can't fathom how hard it would be to coordinate all of these changes with polyrepos. You'd have to burden every team with a required change and force them to accommodate. Someone not familiar with the problem has to take time out of their day or week to learn the new context. Then search and apply changes. And there's no auditability or guarantee everyone did it. Some isolated or unknown repos somewhere won't ever see the upgrade. But in a monorepo, you're done in a day.
Now, here's a key win: you're really at an advantage when updating "big things". Like getting all apps on gRPC. Or changing the metrics system wholesale. These would be year long projects in a polyrepo world. With monorepos, they're almost an afterthought.
Monorepos are magical at scale. Until you experience one, it's really hard to see how easy this makes life for big scale problems.
Heh. This makes a couple assumptions that I only can wish were true: (a) that people won't go to monorepo until they hit some huge scale, and (b) that people will at that point have good test coverage.
I fix bugs that bother me in random projects (both internal and external) maybe once a month (most recently, in the large scale change tool!). For context, I've been at Google for ~3 years. I've only had a changelist rejected once, and that was because the maintainer disagreed with the technical direction of the change and not the change itself.
You want to make it easier to contribute so that people can send a patch and it’s more likely to be useful without too much back-and-forth in code review. Having common tools and coding standards makes that more likely.
I sometimes leave some repos at older versions of tools. Sometimes the upgrade is compelling for some parts of our code and of no value to others.
That's fine, but I think you should be careful whether you're pointing to the properties of using a single repsoitory in general; or the properties of tooling certain monorepo-using companies have built with no requirements other than supporting their own source control; or how uniform it can feel to jump into multiple projects when every project has been forced to a lot of the same base tooling beyond just source control; and/or a work culture that happened to have grown up around a certain monorepo - but for which a monorepo is neither necessary nor sufficient to reproduce.
I've worked jobs where the entire company is in a unified repository, and companies where a repository represents everything related to a product family, and places where each product was multiple gitlab groups with tons of projects.
The most I can say is that monorepos solve package management by avoiding package management. The rest comes down to tooling, workflow and culture.
I would be interested in hearing why it would be hypothetically worse if google had gone the other direction. Where they still spent the same amount of money and time from highly talented people on the problem of unifying their tooling and improving workflow, but done it to support a polyrepo environment instead. How would it have been fundamentally worse than what they got when they happened to do the same with a monorepo?
The only way this could ever work is if the change is really nonbreaking (do those exist in Javascript?), in what case you could script the update on as many repositories you want too. Otherwise, living with the library vulnerability is probably safer than blindly updating it on code you know nothing about.
Anyway, burdening all the teams with a required change is the way to go. It doesn't matter how you organize your code. Anything else is a recipe for disaster.
This is what tests are for.
> Are you proposing that a single developer/team clones the work of 100s or 1000s of different people, update it into to use the new leftpad, run the tests and push? ... Anyway, burdening all the teams with a required change is the way to go.
No, and speaking from personal experience, it's much more difficult to ask ~500 individuals to understand how and why they need to make a change than to have a few people just make the change and send out CLs. Writing a change, especially one that you have to read a document to understand, has a fixed amount of overhead.
(Also, you don't have to clone all the repositories if you're in a monorepo :) ).
If you want consistency so you can automate stuff, require consistency.
First, inter-repo dependencies are managed by pulling specific a commit or tag. So changing a library has zero effect on the program depending on that library.
Second, if you are introducing breaking change and you know not all client will want the change, you can have multiple branches. No, you do not want this as the default choice and not on the long term, but for the short-term transition, that is possible.
From that point on, clients of the library can upgrade to new versions with the changes on their own schedule. The client is never forced to upgrade until there a feature is absolutely needs. The library is not forced to support two versions of code.
That last point is not trivial. If every braking change need to go through this dual-support period in the same single code base it can become a support and testing nightmare. You need to duplicate tests and as the number of such dual-version of API increase, the compatibility matrix grows exponentially.
This is entirely avoided in the multi-repos scenario.
Maybe it's just where I've worked, but atomic commit carries with it some strong cultural norms that make it really tractable to have one version of everything. If code compiles and automated tests pass, the change is safe and may be committed unilaterally. Inevitably things go wrong, but the post mortems for incidents don't lead back to the lack of permission from the affected projects.
I wouldn't maintain software that can kill people in this way, but for everything else it strikes a nice balance.
I think the real intention of the "atomic commits" idea is a total ordering of commits between both the client and server code. Both a "big bang" strategy as well as the author's incremental change strategy can benefit from that arrangement.
The key is that at any given point in the repository's history, you can be sure that the client and server both work together. The point is not that each commit atomically completes an entire new feature, only that each commit atomically moves the repository to a state where the client and server still work together.
In that sense, the author's incremental commits actually do have that kind of atomicity.
> It's particularly easy to see that the "atomic changes
> across the whole repo" story is rubbish when you move
> away from libraries, and also consider code that has
> any kind of more complicated deployment lifecycle,
> for example the interactions between services and
> client binaries that communicate over an RPC interface.
This seems exactly wrong to me. Getting rid of complicated deployment lifecycles is exactly the job that people use monorepos to solve. Wasn't that one of the reasons they are used at Google, Facebook, etc? As well as being able to do large-scale refactors in one place, of course.You should be able to merge a PR to cause a whole system to deploy: clients, backend services, database schemas, libraries, etc. This doesn't preclude wanting to break up commits into meaningful chunks or to add non-breaking-change migration patterns into libraries -- but consider this: is it meaningful for a change to be broken into separate commits just because it is being done to independent services? What benefit does cutting commits this way give you?
What you want to avoid is needing to do separate PRs into many backend service and client repos, since: (1) when the review is split into 10+ places it's easier for reviewers to miss problems arising due to integration, (2) needing to avoid breaking changes can sometimes require developers to follow multi-stage upgrade processes that are so difficult that they cause mistakes, and (3) when there are separate PRs into different repositories these tend to start independent CI processes that will not test the system as-it-will-be (unless you test in production or have a very good E2E suite -- which would be a good idea in this situation).
I will say that, even in a monorepo, a big change might still happen gradually behind feature flags. But I think that generally it's nice to be able to deploy breaking changes in a more atomic fashion.
In our monorepo we have to treat database changes with care, like you mention, as well as HTTP client/server API changes, but a bunch of stuff can be refactored cleanly without concern for backwards compatibility.
And if only a single binary is produced, quite likely a single source code repo would be used as well - sounds like 'single developer mode', well 'small team' at most.
In this scenario, the single binary is the key encapsulation boundary, but your monorepo could be producing N binaries, each of which receives the change.
For example, if the change is to remove a single, unnecessary allocation in a low-level library function used across the repo, you can refactor it out and push the change as N binaries without worrying about compatibility.
This is handled by our deployment tool. It won't allow the new executable to run until the database has been updated to the new schema.
In fact, when I was last there in 2017, making a backwards-incompatible atomic change to multiple unrelated areas of the codebase was forbidden by policy and technical controls (the "components" system). You had to chop that thing up, and wait a day or two for the lower-level parts of it to work their way through to HEAD.
I would generalize this to say that the idea of deploying clients, schemas, backends, etc all at once is an inherently "small scale" approach.
For example different npm (cargo, etc) packages are controlled by different entities. Semver is used to (loosely) account for compatibility issues and allow for rolling updates.
A single company requiring multiple repositories for control reasons might be an antipattern and might indicate issues with alignment/etc.
> The benefit of changes straddling repositories
> is having separation of control.
Good point, although a monorepo with a `CODEOWNERS` file could be used to give control of different areas of a codebase to different people/teams.Also remember that non DVCS repos generally have find-grained access controls.
No, neither of those is why big companies use monorepos. Clearly the kinds of things you wrote are why the general public thinks big companies use monorepos, which is why this argument keeps popping up. But given making atomic changes to tens, hundreds, or thousands of projects does not actually match the normal workflows used by those companies, it cannot be the real reason.
Monorepos are nice due to trunk based development, and a shared view of the current code base. Not due to the capability of making cross-cutting changes in one go.
Otherwise, there is much more fear that a change to an important library will have downstream impact and when the impact does arise, you've moved on from the change that caused it.
Right?
Being able to rollback is indeed important and reducing the sizes of commits makes it less painful. If one project has a problem, unrelated projects won’t see the churn.
Or at least that’s how it was done when I was there. It’s been a while.
If you divide it up, most commits will land quickly and without issue, and then you can deal with the rest.
I personally don't have that much confidence.
Plus, when most people think LSCs, they think about the kind of stuff that the C++ or TS or <insert lang> team do, not someone refactoring some library used by a handful of teams, which themselves usually impact anywhere between ten to a few hundred files.
1. Add the new interface (& deprecate the old one)
2. Migrate callers over (this can take a long time)
3. Remove the old interface.
Even then, you risk breakages because in some cases the deprecation relies on creating a list of approved existing callers, and new callers might be checked in while you're generating that list. (In that case you would ask the new callers to fix-forward by adding themselves to the list.)
This three step process has to happen because step 2 takes a long time for widely-used interfaces, and automatic refactorings cannot handle all cases.
The only time you can consolidate all three into one commit is if the refactoring is so trivial that automatic tooling can handle every case (in which case, does the cost of code churn & code review really justify the change?) or the number of usages are small enough that a person can manually make all the changes the automatic fixer can't handle before much churn happens.
Whether or not that's the most common workflow is another story. Works great at the C++/Java API layer or trivial reorganizations, may not work as well when modifying runtime behavior since you have to do the 3 phase commit anyway.
Approving commits is part of any (non-global approval) LSC process; originally I meant that there's a lot more friction in a consumer generating code changes than approving them. If you received an email saying that you had to make a bunch of small changes to your codebase, you would probably ignore it and keep working on your P1/P2; on the other hand if you receive a small, reasonable CL from a library team updating their function usages you would likely LGTM without much thought.
On the other hand, if the change is: we need to completely overhaul the interface of this library/service, then you need more in-depth reasoning and the consumer team is responsible for making the change (similar to deprecating a service), since these large "deprecations" require a lot more knowledge of how consumers use the code.
There is also a middle ground where the changes are somewhat easy to reason about but still require some human intervention; then the task is usually assigned to volunteers.
Also, a monorepo forcing trunk-based development is great. Long-living feature branches are hell. I would even say to avoid feature branches entirely and use feature switches instead whenever one can get away with it. Every branch always runs the risk of making refactoring more difficult.
But I do like short-lived (max 1 sprint) branches for my own smaller scale projects because they group changes together. I name my branches after an issue number, rebase freely on that branch, and merge with a `--no-ff` flag when I'm done. My history is a neat branch / work / merge construct.
Not sure if this is just fear, but I believe trunk- and feature switch based history will end up an incoherent mess of multiple projects being worked on at the same time, commits passing through each other. I'm sure they can be filtered out, but still.
Why exactly? The best I can think of is that you may annoy people because they'd need to rebase/merge after your change lands and takes away a function they were using.
If you're literally just doing a rename, and you're using something like Java, why not just go ahead and do a global rename?
Then once some trivial refactoring inevitably causes some kind of breakage ("oh, somebody was using reflection on this class"), you'll need to revert the change. That'll be another set of full code reviews from every single owner. Let's hope that in the meanwhile, nobody pushed any commits that depend on your new code.
None of this is a problem if you have a library and two clients. But the story being told is not "we can safely make changes in three projects at once", it's "we can safely make changes in hundreds or thousands of projects". The former is kind of uninteresting. The latter is a fairy tale.
> At Google, we’ve long ago abandoned the idea of making sweeping changes across our codebase in these types of large atomic changes.
I suppose a single-developer code base in a fully checked and compiled language that doesn't support any kind of reflection has no particular added risk from renaming vs introducing a new name. Each time you remove one of those constraints, you add a little bit of risk.
If you have a giant company, it might be possible that someone is copying a jar file and calling it in a weird way that you don't expect.
If your language isn't fully checked, you might correctly rename all the Typescript uses and miss a Javascript use.
If the language supports dynamical calling, it might turn out that somewhere it says "if the value of the string is one of these string values, call the method whose name is equal to the value of the string". There's various IPC systems that work this way, and it will certainly be hard to atomically upgrade them. I hate that kind of code but someone else doesn't.
If your language supports you doing these things, you can create as many conventions as you like to eliminate it. But someone will have an emergency and they need to fix it right now.
Some people view the correct way of dealing with that problem is to insist on the development conventions, because we need to have some kind of conventions for a large team to feasibly work together.
But I guess the author leans towards the side that says "if it's valid according to the language/coding environment, it might be better or worse, but it's still valid and we need to expect and accommodate it". It isn't my preference but it's a viable position - technical debt is just value if you can accommodate it without some unreasonable burden.
It introduces changes to places which really doesn't need changes. We've done both at work, but I mostly prefer just making the old function(s) simply call the new one directly.
Then you won't "pollute" source control annotation (blame) and similar.
Not a 100% thing though.
We are using variations on this theme extensively at my company, in a large spectrum of projects, with great satisfaction.
Crucially, this method is orthogonal to using monorepos. It is simply a safety net for and good stewardship of your APIs.
Definitely a best practice in my tool belt.
I assume you are not creating some strawman every source code file is in a separate repository. That would be insane and nobody does that. If you are in a mutli-repo world, then how your break your repos up is an important decision. There are interfaces that are allowed to use within one repo that you cannot use in others, which allows things that should be coupled to be coupled, while forcing things that should be more isolated to be isolated. This is the power of the multi-repo: the separation of what is private vs public is enforced. (which isn't to say you should go to multi-repo - there are pros and cons of both approaches, in the area of where and interface can be used multi-repo gives you more control, but there are other ways to achieve the same goal)
Even if we take the one library example you have, the library has source file A.c and B.c (I'm using C as an example, but this should apply to any language). A is the main interface, with B helper functions. With a mono-repo it is easy for someone else to use B even though that is not supposed to be the interface, with a mutli-repo setup the interface to B isn't published, and so only A can get at it. Of course the point to this is if B.c is wanted as an interface you get notice of that need and can review the interface to make sure it is clean.
In a more complex variation of the above, related libraries in the same repo LibX and LibY can both use B.c, but libraries in other repos can't get at B.c.
Does this matter to you? That is your decision. In a simple situation it won't. In a complex project it needs to because complexity is the enemy.
Again, mutli-repo is only one possible solution to this problem. There are many ways to solve this problem, each with pros and cons. You need to figure out what is the best compromise for your situation, and then mitigate the cons to your choice. There is no perfect answer. I've touched on one a few of the pros and cons of them - and there are probably some that don't even apply to me but will to you.
However it doesn’t have to be one or the other.
Sometimes a single commit (or PR if you prefer multiple commits) can update the api and all clients at the same time.
Sometimes client callers are external to the repo/company. In which case backward compatibility strategy is needed.
There is no need to abandon a concept of a single repo just because you might not use one of the main benefits all the time.
Changing a model that is shared between different services (or a client / server) should be atomic. In practice, you need to think about backwards compatibility, and work in the three-step process outlined in the article (build and deprecate, switch over, remove deprecated code across a number of deployments / time).
If you don't have atomic deployments, that's one less argument in favor of monorepos.
We use a monorepo for our organization and have found that feature flags are the best way to manage the problem of exposing new functionality in ways that won't piss off our customers. We can let them tell us when they are ready and we flip the switch.
Once a flag goes true for every customer, make a note to drop it. This is important because these things will accumulate if you embrace this ideology.
The ultimate goal of Reviewpad (https://reviewpad.com) is to allow the benefits of the monorepo approach independently of how the codebase is fragmented at the git level. We are starting from the perspective of code review and debugging (e.g. code reviews with multiple PRs, or code reviews across multiple projects). For people doing microservices like us, the ability to have a single code review for both library and clients has been quite positive so far.
But even when problems come up that you didn't see in the tests, it is easier to revert the work if it indeed is a single commit.
The real reason why I still sometimes prefer many small commits is to reduce the chance of merge conflicts.
In my past job, I wrote a script to update all deployment manifests in tens of repositories. All commits arrived at the git server at the same time. Needless to say the whole team had to stop all work to wait for the ci/cd triggers @@@
because everything is in the same place, you can, in theory can find stuff.
However in practice its also a great way to hide things.
But the major issue is that people confuse monorepos for a release system. Monorepos do not replace the need for an artifact cache/store, or indeed versioned libraries. They also don't dictate that you can't have them, it just makes it easier _not to_.
You can do what facebook do, which is essentially have an unknowable npm like dependency graph and just yolo it. They sorta make it work by having lots of unit tests. However that only works on systems that have a continuous update path (ie not user hardware.)
It is possible to have "atomic changes" across many libraries. it makes it easier to regex stuff, but also impossible to test or predict what would happen. Its very rare that you'd want to alter >0.1% of the files in your monorepo at one time. But thats not the fault of a monorepo, thats a product of having millions of files.