Josh: Get the advantages of a monorepo with multirepo setups
github.com
github.com
Most of the time, when an org chooses to move to having a monorepo (rather than just being left with one by accident of history), the key advantage they're striving to attain, is the ability to make changes to cross-cutting concerns across many distinct applications/libraries, with single commits/PRs. To change an API, and all of its internal callers, atomically, without having to worry about symbolically binding the two together with dependency version constraint resolution.
Which is to say, the key advantage of a monorepo comes from having the whole monorepo checked out locally.
When each commit shares a common timeline it is really easy to rebuild build 1 with the exact same dependencies.
These are the "obvious" solutions to this problem, the first ones the average software architect would reach for. What would lead them to ignore these options and choose a monorepo instead, if not for what I mentioned above — the ability to make atomic changes to cross-cutting concerns?
First you commit/push to one repository
Then you update the submodule pointer in the parent repository
They would have been, except the experience working with them was somewhat problematic for those who have been down this path, and the message that many devs have received is to avoid them. I kind of agree with going to submodules, but confidence doesn't seem to have been rebuilt.
For example, you want to have a project in the monorepo, but also share it outside of your organisation, and even have external contributions brought back in.
Or, you are working in a project in the monorepo, but you don't want to (or can't) check out the entire repo. You can still checkout just that project and its deps.
Or, you are working in a polyrepo-using organisation, and you want to experiment with using a monorepo while continuing to allows devs to work in the polyrepos that it is composed from. No big-bang cutover necessary.
Avoiding "here's my new library version, go see if it breaks your shit" was the goal - you make a change, you run the tests, you see if the whole company's code can still build or not. Having fully-separate projects in directories in the monorepo using published dependencies was considered an antipattern (though it was very hard to keep some teams from doing that).
The disadvantages of the resulting monorepo weren't "this directories are so big to keep checked out when I'm just working on one specific project" it was "our old build times and build tools are dying under the strain and even trying to move to a 'monorepo friendly' build tool might be an intractable problem because our dependency graph has become such a mess of spaghetti."
A monorepo that was done well from the start so you don't have the slow-spaghetti-build problem from months or years of "oh it's easy to depend directly on this full other module, let's just do that" sounds very appealing. We just didn't pull it off in practice, and this project would ... maybe... help in the early stages by letting people have more restricted checkouts? But only if you already know what you're doing anyway.
Re the tooling issue, I’ve frankly never seen it solved. I presume Google’s done it, because they’re a bit messianic about the whole thing, but outside of them, I suspect all monorepo implementations fit into the “Communism” model of working great in theory.
Most Google APIs are written in protobuf, which is backwards compatible over the wire. This makes most changes painless, although caveat emptor as things can still get messy as they do in the real world. With something like a Java interface, it doesn't really matter if those APIs change because all the code is built and released together at HEAD.
That discipline to commit to an API and meaning for each field and not keep tinkering with it in messy ways is an under-valued and under-evangelized aspect.
Right, this is the thing that winds up following from the monorepo - the monorepo problem goes away if we just release everything, the whole fuckin’ system, into production every time we make any update to master. If that’s your system, hats off to you, I’m not nearly that good of an engineer.
Can you give a more concrete example of the problem? I don't think I'm understanding.
Short version, you’re right, coordinating releases with feature flags, branches, etc. solves this problem, but it solves the problem in a multi-repo world too, and a monorepo doesn’t obviate the need for those solutions as many of its proponents seem to suggest.
There are probably legitimate uses for monorepos, but an awful lot of people seem to position them as a silver bullet that lets developers stop having to worry about coordination in a multi-service world, and that’s just not how it works out.
As someone pointed out though, proto mitigates a lot of this.
Amazon's Brazil system doesn't sound like it does the "test the whole company's code" aspect in the same way.
> If you create a package at Amazon, you specify an interface version like "1.1". As long as changes are backwards-compatible, the interface version is not changed. When Brazil builds a package, it appends an additional number to turn it into a build version like "1.1.3847523". You can only specify dependencies on interface versions.
So if I'm working on a non-backwards-compatible version, my users aren't going to pull it down immediately in the build (if I properly call it version 1.2.x), so the build won't show which of them it broke.
Even if you depend on the latest version instead of pinning to a major version, how would I force a company-wide build of everything that depended on me using my not-yet-landed branch? I can run those tests in a monorepo on my branch directly, I don't see how I'd do it in a manyrepo setup without a very specific set of "every project in the company specifies its dependencies as depending on specific, easily-overrideable, commit SHA versions."
But it's important to note amazon doesn't monorepo like you are used to thinking, so it is actually on everyone to take the updates on the schedule they've set via automation or if it's an emergency via ticketing / policy. You are merely seeing if it's going to work, the automation does the rest, possibly over weeks.
And I am simplifying things here to make it easier to understand. Understanding brazil from not using it, is about as difficult as understanding git rebase coming from CVS, eg there are concepts you didn't even knew existed that you have to understand how to use to operate at large scale.
You are absolutely right about the main motivation of using a monorepo: Allowing upsrteam library maintainers to see downstream usage of their code and make the required downstream changes themselves at the same time they change their libraries.
Also like you say the easiest way to get those advantages is to just check out the monorepo locally, so if there are no other reasons preventing you from doing just that, go for it.
However there are a few reasons why this is not always sufficient:
Size: The repo might be so large that cloning it all will makes local tools (git cli, guis,...) slow to use, or in the most extreme case require to much disk space for your machine. To address this there are some git native tools like partial clone and sparse checkout, so size alone is not really the the main issue for us.
History "pollution": Having a lot of somewhat loosely related projects in one tree means a history that shows all the changes. Yes git can filter them, but once again that might be a performance concern, but once again not really the biggest motivation to create a new approach/tool.
Permissions: In some organisations (like the one I work for) it is not possible to give all developers access to all the code and thus the advantages of monorepo get lost just by trying to comply with data protection standards. The only solution with native git is to split the repo at legal (not necessary technical) boundaries and try to coordinate the changes across those. Loosing most of the benefits described. Josh does not have a full blown permissions system yet, but the concept certainly allows for it and implementation is work in progress.
Sharing with others (aka, distributed VCS): This is the biggest motivation for using something like Josh. The partial repos are repos in their own right and all the distributed features of git can be used with them. In a monorepo setup as you describe distributed workflow is sacrificed for monorepo advantages. Only developers in the same monorepo see the same sha1s and can easily exchange changes. In Josh the same library can be part of different monorepos at different organisations and while the monorepos have different history and therefore sha1s, the “projected” or “partial” library subrepos will have compatible history with identical sha1s. In this way Josh can serve as a bridge between organisations using different repo structures.
dependencies = :/modules:[
::tools/
::library1/
]
how are the canonical build artifacts for, say, ::library1/ determined, and how are they presented to the workspace?I understand that the partial repo layering is the key innovation that exists a layer below what I'm talking about, but I'm trying to understand how you can ergonomically layer never-build-twice logic on top of it.
What artifacts are to be build inside a given workspace is totally up tho the build system(s) and tools that work after the files have been checked out to a working copy at which point Josh is not involved at all.
PS: The title, in case it is changed, is currently “josh: Get the advantages of a monorepo with multirepo setups”
(No, git submodules are not it.)
I'd dream of somethong as simple as
git clone --partial foo/bar https://example.com/some.repo.git
And then everything would work normally. git clone https://example.com/some.repo.git:/foo/bar.git
And then everything works normally.This is the kind of talk about monorepos that makes me think they are a bad idea. Why would someone want to maintain a monorepo and then pretend it's not a monorepo? Not only just pretend it isn't, but invest not-insignificant time on the problem of pretending it's not a monorepo?
I am immediately thinking of the horribleness of how some of the (older) javascript frameworks re-invented the back button (and browser history in general) instead of.. ya know, using the browser.
With each project having its own repo, then you have to track the fact that Foobar 2.2 works with baizo 1.6-1.8 but not more recent versions.
Also conceptually it's easier when you are working with the client and the server at the same time, or the two mobile apps, and so on.
Of course people manage without this when the project has stuff that doesn't fit in a software repo (CAD designs, artwork, etc...there's a reason why that POS Perforce survives, for example). Solidworks has its own proprietary RCS that doesn't work with anything else.
IMHO if the project is relatively small (say <500K LoC) a monorepo is almost always the way to go. But with a big project it breaks down.
And it’s less about LoC, and more about the number of files and how much binary stuff you put in your repo (and how often it changes). Git is really bad when binary data is involved.
Git has the facilities to keep monorepos clicking along (shallow clones and sparse checkout) but they aren’t along the “happy path”
Keeps the binaries out of your repo, replacing them with pointers
A typo? You seem to mean that multi-repo is the way to go :)
The same story is true with things like APIs or types where two services need to stay in sync.
They did that because the browser didn't support adding to the history via JavaScript.
But even now that the browser does support adding to the history via Javascript ... is that really just "using the browser"? At some level in many modern web apps back button history is not just the browser. This isn't an ancient thing left behind with old frameworks.
You can update a library and all the downstream projects in a single commit. There's no race condition or caching problem of pulling an update without pulling/seeing the dependency update. You don't need to wait for dependency artifacts to build and propogate.
You can create a turn key build script that will build the world from source. You can skip any local artifact storage like Artifactory. You don't need to pull multiple repos in a serial fashion, no dependent pulls. You can structure your codebase such that if you pull one commit it can have no other dependencies.
The draw back is Git happens to not make it easy to pull just one folder. Other things like Perforce make it trivial.
As an example, Chromium [0] is a non-AOSP project that also uses repo.
[0] https://chromium.googlesource.com/chromiumos/docs/+/HEAD/dev...
I am not affiliated with the project.
JOSH claims to be reversible, so it could be used in either direction, which is where the multiple use cases come in. Treating subsets of a repo as their own repo, or treating multiple repos as one. I would say there is some application overlap between this and git submodule/subtree/subrepo and also tools like copybara.
I'd love if someone still working there were to write a nice post about that system, it was the first of such a kind I saw.
I think the best approach would be to have bidirectional links between the projects (if A needs B, then A has the stable version of B and vice versa). The point in that setup would be that "upstream" projects can notice when they are about to break tests in "downstream" repos and act accordingly.
It's a bit complicated to explain, but it works. That's why I hope they'll make a blog post :)
I don't know how far back you saw the Bloomberg system, but at this point it's basically the same as the Debian system (as in, debian/ subdirectories, .deb files, etc.). Versions of git projects are published as tarballs (source packages). Then sets of published projects are "promoted" and all projects that transitively depend on them are rebuilt and unit tested in a sandbox environment. If that process fails, the promotion fails.
Each source package can use any number of build systems, implementation languages, or project structures.
There's also a legacy subversion monorepo with a monolithic build system that builds on top of that, but it's slowly being phased out.
All that is an integration build including thousands of discrete projects. Those projects typically have additional CI/CD enrollments outside of the integration build system too.
Also, there are quite a few tools to manage the distributions, and it would be great if they were open sourced. Basically Bloomberg championing their approach, to gain the usual advantages of open source (developer familiarity, cooperating across companies for improvements, and so on)
> Those projects typically have additional CI/CD enrollments outside of the integration build system too.
Another thing to call out is that you can simulate the "promotion", so you can check in your PRs whether your change is going to break any dependency or dependant.
That's exactly one of the things josh can already do for you :) Josh's concept of workspaces is precisely this: define your dependencies (no matter where else they originate from in the monorepo) and then check out only those dependencies, along with any code that solely exists in the workspace. Your workspace checkout is effectively the "bloomberg"-style repo setup you described, as you only see your code and the code of your dependencies, but when you push, your changes get added back to the hidden, backing monorepo (the source of truth) where all related and pertinent tests are run, and your change can only be committed if those tests all pass.
Thus, your commit is your "release". Sure, you don't have exactly the same workflow, as there's no difference then between a commit and a release, as by your definition you don't release after every single commit, but the whole "release this change to everything else in the monorepo" is touted as one of the benefits: there's no massive integration headache if you have multiple breaking changes which you then need to work on resolving for everyone else.
source: I work directly with "chrschilling" - who wrote josh
There are others that aren't free (e.g. Perforce) and some that aren't quite dedicated to the monorepos & subtree workflow, but which handle it better by design (e.g. Darcs, http://darcs.net/).
But, mostly I'm thinking of Subversion.
The problem with using svn on the server is the lack of good tooling for things like code review. There's no SvnHub or SvnLab.
I mean, most of the people in industry today have known no character encoding except ASCII and its supersets --- and that's a good thing!
There's nothing about the git/mercurial object models that makes them intrinsically inefficient with monorepos.
What's inefficient is materializing the object database (when cloning) and the working copy (when checking out), when you're only going to need tiny portions of them.
Subversion doesn't have the first problem (but comes with extremely slow history operations), but sparse checkouts don't really solve the second because you have to statically know what to filter.
A better direction is instead to virtualize the filesystem, so you get the semantics of a real monorepo with git/mercurial, and you fetch only what's actually needed without any change to your tooling (it just needs to interact with the filesystem). It's also very easy to transparently implement caching and prefetching this way.
This is the approach that Facebook and Microsoft took, and I believe Google too.
Its certainly not against the design goals as official tooling supports shallow and sparse checkouts, they're just experimental features.
We should just strive forward and continue to make these features good.
It works against GitHub's design goals. You could work on multiple, logically distinct projects in the same git repository easily, if you so choose. git was designed for Linux's workflow.
People like to use distributed workflows even with monorepos. E.g. chains of commits, branches, rewriting local history, etc.
It's clear that people want a mixture of monorepos with distributed workflows. There are two ways to get there: add distributed flows to a monorepos, or build a monorepo layer over a distributed tool.
Both seem like valid approaches. The market will decide which approach it prefers.
I get the impression that things like LFS, submodule and subtree are hacks added onto git to try to make it behave like something it isn't.
As monorepos grow huge, this comes to be very costly or even prohibitive, and companies like Google simply don't use Git.
Here are some problems with alternative approaches that have been mentioned:
* VFS for Git: I believe abandonded by MSFT in favor of improved client-side tooling: https://github.com/microsoft/VFSForGit/blob/master/docs/faq.... .
* Sparse checkout: limits ability to use a build system to dynamically find any dependencies and rebuild them
* Submodules: can't atomically update both the parent and the child repo, have to manually update the referenced commit of the child repo in the parent repo, and each collaborator must manually update their child repo when the commit changes
Beyond that "mostly", it also configures git lfs, which likely will always be a git plugin and not directly in the client and the rest of it seems like stuff Microsoft is testing before upstreaming it directly into the git client.
this can also be solved by using a git mirror.
You still end up doing some kind of large checkout.
(We got our CI checkout time from 40+ minutes to well under half a minute this way.)
I’m currently using a monorepo with submodules and it works really well.
Dependency management is not a huge issue too, at least if you don’t have hundreds of submodules.
Imagine a company that develops strongly interdependent software in a monorepo, but needs to publish different subsets of this software to external entities which also expect a coherent version history.
Also being a server it does not require any installation or resources on the developers machine.
In addition to that over time more features where added that git-filter branch does not have, most notably "josh workspaces" which is a DSL for repo transformations.
With multi repository projects, this helps some thinking, as it is clear that the changes were not atomic between systems. They are literally separate at all layers, including the commit.
I sympathize with wanting a simpler view. I'm just worried on an inflated value proposition.
I'm not sure what, exactly, the problem is that you're talking about. When you deploy something, it's built from a specific commit.
Deployment is not atomic, it's a gradual process that takes some amount of time. Between when you (or your automation) chooses to deploy a system and when the deployment finishes is some window of time. The system will often spend much of that time in a partially updated state. You may also choose to canary changes, so you will have a mix of different versions in production at any given time. At companies where I've worked, the time from deployment start to finish for backend systems ranges anywhere from hours to weeks.
I don't understand how this relates to multi-repo or mono-repo concerns, however. The repo is a history of the source code (intentional changes by humans), it's not a history of the state of your production systems (which are the results of automation).
It is amazing how many times I've seen folks think that just because they can get build time tests happy with changes in two projects, that they can safely send out the two changes.
Does a multi repo "solve" this? Of course not. But it is easier to reason that two projects clearly need two deploys. Versus having to remember that one commit could be N project deployments.
Why is it a foot gun? I don't know what the negative consequences are here.
Atomic commits are just to make development easier. It means that you can refactor downstream dependencies in the same commit that you make a change to an upstream library. This way, you either build and deploy the new library + the refactored downstream dependency, or you build and deploy the old version, but never some mix. It reduces the number of possible configurations that can be built & deployed, since you can only pick from a point in one repo's history, and you can't mix and match various points in the history of different repos. With multi-repo, it is harder to discover down-stream dependencies.
> It is amazing how many times I've seen folks think that just because they can get build time tests happy with changes in two projects, that they can safely send out the two changes.
Build-time tests don't catch 100% of errors. Some errors will still make it into production. I don't see how this problem is related to the multi-repo vs mono-repo problem at all. At places where I've worked, if you change project X and project Y, and your commits pass the build test, you still have to submit the commits in some linear order. The pre-commit tests for X will include Y or vice versa. X or Y will be rebased or merged on top of the other one, and the result will have to go through pre-commit tests.
I am just trying to understand what problem multi-repo is supposed to be addressing here, and I don't have the slightest clue what you're getting at.
My assertion is that the "build and deploy new code with updated downstream users" only works in a minority of cases. Now, I grant that this could be due to the micro service nature of where I'm at. And I also grant that it is nice when this can work. However, the times it goes wrong are usually not at all worth the risk.
I understand what "foot gun" means. Explaining what "foot gun" means is not helpful. What I don't understand is the problem that "shooting yourself in the foot" is a metaphor for.
> My assertion is that the "build and deploy new code with updated downstream users" only works in a minority of cases.
There are two main cases here: libraries and services.
Libraries can be atomically updated no problem, most of the time. You change a function and fix the call sites.
For services, you have to add to the API, deprecate the old thing, refactor the clients, and then remove the old thing after all the clients have been rebuilt and redeployed. At least two steps.
What I don't understand is how this would be different for multirepo or monorepo setups. In either case, removing some old piece of functionality requires waiting until the clients have been redeployed. Using new functionality requires waiting until the server has been redeployed.
> Now, I grant that this could be due to the micro service nature of where I'm at.
The teams where I've used monorepos are also the teams with the most buy-in to microservice architecture. One team I was on ran a service that consisted of something like twenty different microservices. These interacted with services run by other teams. Everything was in the same monorepo. It worked fairly smoothly, as I recall--we spent most of our time solving domain problems and working on our team's core mission, and I don't remember any problems arising from the monorepo setup.
It sounds like your experience is different, and I was hoping that you would share some of that experience.
Yes, you can do similar with reviews that span multiple projects. However, I have seen it happen far more times in single repositories than I have in multiple ones.
Again, I am not in any way offering a panacea. I'm just saying that seeing things as atomic at one level leads people to think they are atomic at the next level. And this is a mistake I've seen many many times.
Yes, I have seen people manage it somewhat well. But every effort I have been involved with that tried to merge code history between projects has had more faults of this kind than the other projects I have been on.
I still don't see advantages to multi-repo here. Even within a single project, or single service, deployment is often not atomic. If you make a change to service X and deploy it, you end up with minimum two versions of service X in deployment until the deployment finishes rolling out.
So we have automated tests that cover that case... each service has to work correctly when combined with not only the repo HEAD, but also with the currently deployed versions (perhaps more than one! in the case I'm thinking of, it was only ever the two previous versions, but other teams had longer horizons) Testing against multiple versions is a bit more expensive, so it's only done as a pre-deployment check, rather than a pre-commit check. You cut a branch for deployment, and when one commit doesn't play nicely with previous versions of the service, you cherry-pick a commit to revert the faulty commit, and run the tests again.
Putting it in your face that you are changing two projects is about the best I can offer here. I would cede that this, too, is mainly a best intentions. But having as close to 1:1 between code changes, deployment changes, and code reviews at least puts things on mainly equal footing.
That is, when the explanation of why some code can't deploy together is that they are in two projects, and you can see that by them being separate repositories, that feels easier than knowing that two parts of a single repository have to deploy separately.
If you don't want people making changes to multiple projects at the same time, it would be trivial to add a pre-submit check to enforce that anyway. There's no need to switch your entire repository layout just to remind people that the code in different folders belongs to different projects.
That is, I would not push to move from one form to the other. I do like multiple projects being as independent as I can make them, though.
You would need to do it for JOSH and whatever your build tool is. You'd probably need some additional git commit hooks to ensure your build tool of choice config is in sync with josh.
You're right, but from my experience with using josh it really hasn't been a pain point. This, however, is coming from the context of having a build system which (appropriately) checks out all relevant josh workspaces as part of the monorepo verification build, and builds all of those workspaces along with building the relevant parts of the monorepo. Thus, if someone adds a new dependency in a workspace project build, the main monorepo verifier build will stop them from committing that change without also making sure the dependency is added to the workspace file, because the associated workspace build part of the verifier will fail (due to the missing dependency).
Additionally, if you're working in a workspace and you add a new dependency, your own local checkout build will fail due to that missing dependency until you add it to the workspace file (and then get that new dependency). As long as you have builds and tests covering your workspace, it's pretty easy to figure out if you forget to add the dependency to the workspace file.
Lastly, at least from my experience, it's overall been a pretty inconsequential price to pay, as new dependencies aren't added frequently after a project finishes its initial start-up phase.
There may have wrong description (FIXME) but a sort list is found here: https://github.com/icy/git_xy#why
Am I getting the correct impression?
git submodules let you tack a second git repo onto an existing one. For example repoA tracks `repoB@version1234` at path `/foo/bar/baz`.
This on the other hand lets you take monorepo and checkout `/go/mysubservice` as a "repo" and treat it as its own repo. Then when you do git pushes etc, it translates the changes into the larger monorepo.
> a blazingly-fast, incremental, and reversible implementation of git history filtering
For example, to run one or more services locally, we use a single script that sits at the repo base - ‘dev.sh service1,service2,...’. This avoids a lot of headaches for our developers, as we enforce compliance when adding a project to the repo. Lint config? One to rule them all. Test coverage thresholds? Single one. This consistency is the biggest win in my opinion.
Similarly, our integration tests are very easy to write without commit skew.
Finally, sharing libraries has been painless - since we have common/ and common/third_party/ directories at the monorepo root.
My previous job used a collection of about 6 repos for different services and such, and it was a constant struggle to ensure the correct versions were used in development - especially if you were working on a bigger feature that wasn't yet released but required "future" versions from other repositories.
Why would one want to do n pull/merge requests, n separate reviews (of related code), and n deliveries for a single evolution, is something I can't understand.
* Devops - one repo, one deploy, let the developers figure out which repo is which.
* Developer - one codebase, let the devops people sort it out and write lots of tooling to make my monorepo work.
If this stuff was easy, everyone would be a developer.
* Culture
* CI/CD tooling support
* Codebase sizes
You just don’t need that stuff, there’s like 20 of you on a team and at best your app probably sucks and barely has users, and if it does have users, it’s probably some trivial bullshit.
You’re all a bunch of ordinary folks, so stop fucking up the workplace with your identity crisis. No, you are not an elite engineer, you are Bob, the guy who goes home every day and watches Netflix/plays video games.
As someone who has been maintaining a React and Django app solo for the past three years, two repositories or more is too much cognitive overhead to work with.
Never doing that again. Monorepo are easier for small apps.
You want to put stuff in the same repo, that’s fine. What’s with all the other bullshit?