We can call it a multi-monrepo, that way our brainwashed managers will agree to it.
We can call it a multi-monrepo, that way our brainwashed managers will agree to it.
Also it's awful that a simple git pull doesn't actually pull updated submodules, you need to run git submodule update (or sync or whatever it is) as well.
I don't want to work with git submodules ever again. The idea is nice, but the user experience is really terrible.
I'm not at my computer to see if modern git prohibits that behavior, but it is indicative of the "watch out" that comes with advanced git usage: it is a very sharp knife
git subtree tries to emulate that, but it does not scale to huge repositories as it needs to change all commits in the subtree to use new nested paths.
Ideally “cached on the network” could be a sort of optional side effect, like with Nix, but you can still reproducibly build from source. That said, I can’t recommend Nix, not for philosophical reasons, but for lots of implementation details.
I've always had a bad experience using submodules, they're imo the poor developer's versioning tool. It's useful when you use a language without a good build/packaging tool, but otherwise, I'm better off leaving the language-specific tool fetch the depended code.
Also when I try to think about reasons to have atomic cross-project changes, my mind keeps drawing negative examples, such as another team changing the code on your project, is that a good practice? Not really. Well unless all projects are owned by the same team, it'll happen in a monorepo.
Atomic updates not scaling beyond certain technical level is often a good thing, because they also don't scale on human and organizational level.
2. You create a PR that fixes the code and the problematic call sites in a single commit. It gets merged and you’re done.
In the multi-repo world, you need to instead:
1. Add conditional branching in your library so that it supports both the old behavior and new behavior. This could be an experiment flag, a new method DoSomethingV2, a new constructor arg, etc. Depending on how you do this, you might dramatically increase the number of call sites that need to be modified.
2. Either wait for all the problematic clients to update to the new version of your library, or create PRs to manually bump their version. Whoops - turns out a couple of them were on a very old version, and the upgrade is non-trivial. Now that’s your problem to resolve before you proceed.
3. Create PRs to modify the calling code in every repo that includes problematic calls, and follow up with 10 different reviewers to get them merged.
4. If you still have the stamina, go through steps 1-3 again to clean up the conditional logic you added to your library in step 1.
Basically, if code calls libraries that exist in different repos, then making backwards-incompatible changes to those libraries becomes extremely expensive. This is bad, because sometimes backwards-incompatible changes would have very high value.
If the numbers from my example were higher (e.g. 1000 call sites across 100 teams), then the library maintainer in a monorepo would probably still want to use a feature flag or similar to avoid trying to merge a commit that affects 1000 files in one go. However, the library maintainer’s job is still dramatically easier, because they don’t have to deal with 100 individual repos, and they don’t need to do anything to ensure that everyone is using the latest version of their library.
1. A critical security/performance fix has no other recourse than breaking the interface compatibility of a library. Far more common scenario is this can be fixed in the implementation without BC breaks (otherwise systems like semver wouldn't make sense).
2. The person maintaining the library knows the codebases of 10 teams better than the those 10 teams, so that person can patch their projects better and faster than the actual teams.
As a library maintainer, you know the interface of your library. But that's merely the "how" on the other end of those 30 call sites. You don't know the "why". You can easily break their projects, despite your code compiles just fine. So that'd be reckless of an approach.
Also your multi-repo scenario is artificially contrived. No, you don't need conditional branching and all this nonsense.
In the common scenario, you just push a patch that maintains BC and tell the teams to update and that's it.
And if you do have BC breaks, then:
1. Push a major version with the BC breaks and the fix.
2. Push a patch version deprecating that release and telling developers to update.
That's it. You don't need all this nonsense you listed.
In multithreading this would be basically mutable shared state with no coordination. Every thread sees everything, and is free to mutate any of it at any point. Which as we all know is a best practice in multithreading /s
Semver provides just a few bits of information, not nearly enough to cover the whole gamut of shared and distributed responsibility.
The comparison with multithreading is not really valid, since monorepos typically linearize history.
I could have some comments on your "overlapping responsibilities" as well, but your description is too abstract and vague to address, so I'm pass on that. But you literally described the concept of library at one point. There's nothing overlapping about it.
Changing code under their nose risks breaking bunch of projects. We can also fix this by rather communicating, right? But if we CAN communicate... then we can go back to the previous option (telling them to update) as it becomes just as viable.
Communicating is always essential, and can't be avoided.
What happens when you roll this out and partway through the rollout an old version talks to a new version? I thought you still needed backwards compat? I'm a student and I've never worked on a project with no-downtime deploys, so I'm interested in how this can be possible.
If you’re changing the interface of an RPC service, then you can’t do that in a single commit, and need to fall back to something like the second approach, but with even more caution to make sure you properly account for releases and the possibility of rollbacks.
I leverage type systems and write tests to catch any mistakes they might make.
With git, you get a local stage/history which lets you rework/reorder your commits for clarity before pushing. It also allows for more options to resolve conflicts, although this increased ability has brought its own problems.
I assume that this is due to its design center around distributed repositories and patch files. But it’s still there even when using a central repo and a monorepo structure.
Good separation of concerns is like earning compound interest on your code.
Just keep the dependencies generic and tailor the higher level logic to the business domain. Then you rarely need to update the dependencies.
I've been doing this on commercial projects (to much success) for decades; before most of the down-voters on here even wrote their first hello world programs.
If you need to handle different versions talking to each other in production it doesn't seem any harder to also deal with different versions in source, and I'd worry atomic updates to source would give a false sense of security in deployment.
It's much more annoying to deal with multi-repo setups and it can be a real productivity killer. Additionally, if you have a shared dependency, now you have to juggle managing that shared dep. For example, repo A needs shared lib Foo@1.2.0 and repo B needs Foo@1.3.4, because developers on team A didn't update their dependencies often enough to keep up with version bumps from the Foo team. Now there's a really weird situation going on in your company where not all teams are on the same page. A naiive monorepo forces that shared dep change to be applied across the board at once.
Edit: In regards to your "old code talking to new version" problem, that's a culture problem IMO. At work we must always consider the fact that a deployment rollout takes time, so our changes in sensitive areas (controllers, jobs, etc) should be as backwards compatible as possible for that one deploy barring a rollback of some kind. We have linting rules and a very stupid bot that posts a message reminding us of that fact if we're trying to change something sensitive to version changes, but the main thing that keeps it all sane is we have it all collectively drilled in our heads from the first time that we deploy to production that we support N number of versions backwards. Since we're in a monorepo, the PR to rip out the backwards compat check is usually ripped out immediately after a deployment is verified as good. In a multi-repo setup, ripping that compat check out would require _another_ version bump and N number of PRs to make sure that everyone is on the same page. It really sucks.
Atomic deploys are not as important, because you can still decide to version your APIs or releases even if you're using a monorepo.
That being said, you can use multiple repos and still mostly avoid trouble by choosing how to cut your codebase (HR software is likely not going to depend heavily on presale, for instance). The metric to optimize is to minimize the required version bumps.