With a monorepo, you don‘t need to think about which commits work with which other commits. One commit ID is a full description of every subcomponent.
No clue if that’s fixed today but it soured the idea of monorepos for me, and that’s not incorporating how often submitqueue would go down.
I've just switched to a company with many tiny repos and IMO it's a huge hassle. There's no automated integration testing, and manual integration testing is a huge pain to set up.
I don't even know how I'd create integration tests that run as part of the CI. What version would I use? If you need to change both sides (very common), coordinating the commits and releases is very painful.
The "monstrously complex" build-and-test pipelines are a significant cost, but the alternative is higher release failure rate and moving slower overall IMO.
1 repo == 1 bounded context == 1 isolated unit of deployment.
Which, well, the Uber people clearly have since these big changes exist.
The most obvious counterpoint: many systems of record need to scale command operations (writes) and query operations (reads) separately.
So you absolutely would have separate programs, ASGs etc for those two roles
But that said, no I don't think you need a meta repo. You need something. Call it a Change set. MVP would be say merge commits in 4 repos. If any fail then revert. If they all succeed then deploy in a specified order.
But when your team owns a repo, then at least the damage is contained within your team.
Google solves that problem with heavy NIH syndrome (its hard to get promoted by utilizing an external third-party lib, better develop your own), and writing tests, yes. And for those few third-party libraries that google still depends on, updating them is a big PITA.
i’m finding it hard to think of a system at a large company that would have no internal libraries used by multiple projects, that wouldn’t require ever making a cross project change.
What helps maintain polyrepos sane
1. A healthy methodology for dependency management
Evolve APIs with deprecations. Decide on a healthy amount of time for a deprecations to live. Set up alerts when deprecations reach certain amount of time. Set up dashboards (e.g. Grafana) to track dependencies. Help other teams update, and think about how to make updates less painful.
2. Define good boundaries and API contracts to adhere to. The most important thing for a good API contract is stability.
3. Don't prematurely split into microservices. Better ideas for stable API contracts emerge the longer you can wait.
'A well-designed system should not have X therefore you don't need Y.'
When 'A well-designed system should not have X' is just a matter of opinion, or doesn't acknowledge trade-offs that might make X the better option than Z, then this isn't a useful argument against Y.
Sure, each project team could upgrade independently, but the company that’s chosen a mono-repo for their go code is likely to desire to have a single team tackle this upgrade.
In practice, you often end up with larger changes which are mostly janitorial. Bumping a version of a log library across the infrastructure, or a version of a serialization system. This simplifies dependency convergence and a best handled as if they are cross-project changes.
The other problem is that over time, good factoring tend to deteriorate. A large software project will invariably have people with different brains working on it, and they'll have different needs for the factoring. So the project naturally pushes itself toward situations where cross-project changes become a necessary thing.
Good factoring can probably be quantified too: dependency chains should be shallow and graph connectivity should be low. The more dependants a module has, the more stable its API contract should be.
at the risk of coming up with a contrived example: let's say you own a service that need to deserialize a datetime in a request in a format you don't currently support. assuming you own the stack, you need to a) update your date library b) update your webserver stack c) possibly update an intermediate webserver stack that includes primitives like logging, telemetry, tracing, auth, service discovery and d) your actual service.
If a->d are all independent, separate components, you have to orchestrate those changes through 4 separate repositories. And god forbid something you did at the lowest point in the stack is completely unworkable higher up.
There's all sorts of rocket science you could do to orchestrate these changes, but it ends up being contrived and edgecasey.
Most of the pain from monorepos can also be addressed with a dash of rocketscience (see:bazel), but the end model tends to have
a) have an easier mental model for the user
b) allow for consolidation of infrastructure work. Ie, your build/ci/language tooling teams can focus their efforts on one place
c) can coordinate changes across the entire stack within one field of view
d) can coral some of the worst, disparate instincts of a growing engineering org (ie, tons of teams optimizing for local maximas without internalizing knock-on effects).
e) fewer weird edgecases.
I like having examples, even contrived ones, but I'm not sure I understood this one. Can you elaborate on what you mean? Is it about adding support for a new serialization format for dates in requests to a service? Why would this affect the webserver stack and logging/telemetry/tracing/auth primitives?
I find that a lot of organizations have really strange thoughts on how to factor things into separate microservices and libraries. Usually I approach by asking the following question: if this was an open source-library or service (e.g. like elasticsearch), would you use it? If not, then its probably not a great candidate for a separate thing - lets try and come up with something else.
One way to handle CI/CD is using standardized pipelines e.g. you tag your repo with a tag `app:node` or `lib:js` and the github org pipeline scanner will find it and assign the standard `app:node` or `lib:js` pipeline to it.
A way that I like better but most tools unfortunatley don't support it yet is for the infra teams to publish libraries that are essentially functions taking some parameters and generating (standard) pipelines/configuration. Those can then be tracked together the same as other dependencies.
1. I don't see how this is a problem specific to polyrepos. Have an "open PRs" link in the onboarding handbook that gives you a view of pull requests from all repos in the organization. GitHub automatically shows you notifications from all repos. If engineers still chose to focus on one or two repos after that, I'm not sure why.
- Have a (Grafana) dashboard where you can see the latest / newest stuff. Use standard GH tools you use for OSS, such as follows etc to keep up.
2. Don't prematurely split into multiple repos. "No monorepo" doesn't mean not having poly-package repos. It means thinking what the sensible (library or service) API boundary is - treating your projects as you would treat library / service development. In this case a separate repo with lib3, lib2 and lib1 sounds like a good way to go - at most one repo per orthogonal internal framework (e.g. core-react-components). Repo dependency chains should be as shallow as possible, and differenting between public and internal packages is important.
3. Help other teams upgrade. If you are responsible for repo A, once you publish a new version tagged appropriately with semver, use the dashboard to look at your dependants and work with them (or rather, for them) to upgrade. Think of your dependants as internal customers, and make sure you add enough value for them to justify the upgrade effort. Cultivate a culture that values updates.
4. There are other alternatives to `npm link` e.g. see `yalc` https://github.com/wclr/yalc
Another pet peeve of mine is that the real issues get lost when you try to generalize. The article attempts to do this but that makes it hard to evaluate its claims. The best way to evaluate (alternative) solutions is to take a more concrete example repo.
For scalability of this model I'll just point to the OSS community; individual maintainers often several dozen active repositories, but also they have an API contract worthy of a documentation website, versioning scheme and planned deprecation, and they typically avoid cross-project dependencies
Besides the "N pull requests" problem you now lose history whenever you move a file across repo boundaries, and you'll eventually have straggler projects staying on old versions for years - so switching to a new way of doing something essentially means supporting both versions forever. Code for common needs gets duplicated, or worse, split out into yet another repository and /then/ duplicated, because no one wants to check if any of 100 repositories rely on the buggy behaviour they want to fix.
I find this to be this a level headed explanation of the advantages of monorepos: https://danluu.com/monorepo/. I'm surprised to see so many comments summarily dismissing them, as if sanely managing thousands of smaller interdependent repositories doesn't require at least as much investment in custom tooling.
The project should stay a full monolith until the factoring is more clear.
> I find this to be this a level headed explanation of the advantages of monorepos: https://danluu.com/monorepo/.
There are two issues I have with this article. One of them is when it describes drawbacks of multiple repositories, its not specific enough e.g.
> That sounds like it ought to be straightforward, but in practice, most solutions are cumbersome and involve a lot of overhead.
The other is that it assumes you have Google scale of resources to throw at the problem. If thats the case you can make anything work / monorepos or polyrepos. The issue is that small-to-medium sized organizations will not be prepared to invest the amount of resources needed to keep a monorepo working well, as the org will largly need to rely on existing available (OSS) tools which often have poor monorepo support. (Bazelifying everything has a significant cost, bazel rules are often not generic enough to work with e.g. the variety of JS ecosystem tools)
If one business change requires code changes to ~every module in your source tree, then of course microservices, and therefore separate repos per service or whatever, make no sense at all.
No, that way lies sorrow and despair. There should only be one version of any dependency being used in your company. Any deviations should require like VP-level approvals or something.
The more clients a library or service has, the more expensive this is, and it's an ongoing maintenance cost for every client. Since people changing the service don't feel the full pain of this maintenance cost, changes keep on happening, and eventually clients get culled because it's too expensive to keep on maintaining them all.
From the outside, this looks like the company abandoning venerable but still working product, and makes people scratch their heads, wondering why.
People often conflate these two things because often teams actually want to have centralized dependencies (so that you're forced to update or die as you said). If that doesn't work for you you can choose to have modules (or groups of modules) keep their independent set of dependencies, all while keeping the code in the monorepo.
* Breaking changes can be done at once. Very helpful for runtime deps.
* No chance that a repo is out of date.
* Upgrades are atomic (may be hard to test a system in a half state).
I've seen people going crazy with this bullshit of monorepos to the point every single directory was a "package" when it could be just a plain module import.
If you want to do microservices, then putting everything back into a single repository and enforcing everyone to use the same version and every change to require upgrades and deploys across the board is totally backwards.
Either do a monolith and have that consistency, or do microservices and allow teams to follow their own rules as long as they keep APIs stable.
Nonorepos + microservices is just a demonstration of everything that's wrong in the technical aspects of our industry. Just applying absolutely everything you read about without even considering if it might be better or not for your specific use case.
At the risk of sounding memetic, the question is not so much "Can you?" but "Should you?". Should you store a monolithic application across multiple repositories? Should you store a distributed application in a monolithic repository?
I agree with @likortera that you should not. And that's because...
> The way you store your source code doesn't have much to do with how your system is deployed.
I disagree with this statement.
Your architecture (monolithic vs distributed) imposes certain assumptions on other aspects of your distribution pipeline. Your workflow (in this discussion "how you store your source code") should support these assumptions, not hinder them.
For example, one benefit of microservices is that cross-functional teams can develop independently of each other. And yet one cited advantage of monorepos is that everyone is on the same version of dependencies all the time. In short, your teams are not independent after all.
Note that I'm keeping the example extremely generic to illustrate this inconsistency, a conflict of interest if you will, that I see people commit in this topic, because in my experience, these questions are not purely technical but involves product/business factors as well. Maybe for most of the people (operative emphasis on "MAYBE", because who am I to judge you), the discussion they need to have first is whether or not they are using the right architecture for their product in the first place.
If you choose to have a microservices architecture, you have to live with the fact that your teams/services will operate at different cadences. If you feel the need to impose a One True Library Version All the Time, then go for a monolithic architecture, and store your code in the same way.
One, teams that share a dependency version can still develop independently; they just share something in common. They already likely share other things in common: deployment target OS, cloud platform and services, shared authN/authZ frameworks, etc.
Two, a monorepo is just the SCM mechanism. As the parent was describing, it doesn't prescribe anything other than the code storage location and how branching, committing, etc. works. Yes, a lot of organizations prefer having a single version rule in their individual monorepos, but nothing about monorepos in general makes this a requirement. You can use multiple versions of the same dependency and still gain advantages from the single commit benefits, and even famous instances like Google's have exceptions where this is the case.
That's not really the case in practice. When people decide to choose mono vs multi based on their benefits, it's become a workflow philosophy in itself. If you choose a monorepo approach but use multiple versions of the same (in-house) dependency across components, you are just opening yourself up for a world of confusion. Sure, you can do it, but should you? Why choose a monorepo structure if you won't take advantage of its benefits?
> teams that share a dependency version can still develop independently; they just share something in common
My point about team independence doesn't mean they should not share anything at all. But rather, they now _update_ together at the same pace because the "atomic commit" that updated a dependency also updated my team's usage of said dependency, for better and for worse. My team might have a reason not to update just yet.
One of the companies I worked for, had the brilliant idea of putting EVERYTHING related to UI/frontend in a monorepo, where almost every single file or two were a different package. I used to joke there were more package.json files than actual js files (it was almost true).
I spent months saying this thing was a terrible idea. Nonetheless the "frontend infrastructure" folks wanted to do some CV padding and play with Lerna and their SV friend's cloud CI service startup, so they went ahead with it.
Months later the big problems started, among which one of the main ones was that they were pushing through every team's throats updates, breaking the product/features those teams were working on, disrupting their roadmaps, accusing each other's of low test coverage, doing hacks and workarounds, shit tons of crazy CI scripts for all the corner cases, much longer deploy times, most dev environments were a lot slower, deployment issues because now we had to deploy several different projects at the same time, etc, and of course not being able to upgrade to latest React because some team in the corner had an issue with it and they didn't have the time at the moment to fix it.
How did they solve all of this? In the span of 3 to 4 months they left the company. All four of them. Leaving behind an incredible amount of technical debt and nearly every frontend team totally fucked up.
What irks me is that some of these guys are pretty popular "youtubers" and spend their days giving talks of how great their work with monorepos and "frontend infra" is. They don't tell the messes they've caused of course.
Monorepos might be great if you're Google and have the resources and talent to do it right. For most companies out there, it is just creating a centralized problem that will eventually block everyone.
I'm a big proponent of monoliths, specially while you're not a > 300 person company. But if you're splitting your teams and services, then agree on APIs, don't break them, and let each team follow their own pace, with their own tools, and with their own schedules and preferences. Otherwise stick to the good ol' monolith and just separate things into modules/imports/whatever.
At my current employer, the main application has a "plugins" system with a very flexible and stable API. Every team around is just building "plugins" that can be installed into the main monolith, depending on each customer needs. This works fantastically well for a company with more than 1k engineers. No monorepos, no coupling, no interdependencies between teams, no parallel deploys, and each team manages their own destiny more or less.
The development view need not be connected to the deployment view at all, so unnecessarily coupling them can lead to worse outcomes. On the other hand, having the ability to couple them initially _and decouple them again in future_ works wonders for scaling.
I am consistently amazed that people with opinions on software architecture do not seem to recognise this seminal paper on the topic, or have the ability to re-synthesise it into tactics.
[1]: https://www.cs.ubc.ca/~gregor/teaching/papers/4+1view-archit...
When teams take the stability & versioning of their APIs seriously, the need to use monorepos to share that info is greatly reduced. A multi-repo approach is perfectly feasible when all components are working to established APIs, which also alleviates the issues mentioned in the article.
- easier to share and import common packages
- proto and thrift files are kept close to services and clients are updated globally automatically
- dependencies and go versions are managed globally and all services get the same security updates
- standardized build processes make it easier to manage large deployments
And honestly, after the initial repo download there were no visible downsides
Is the suggestion just "an easier way to avoid this problem is to not have so many tests?"
They also use mostly homegrown CI tools alongside phabricator for code review, or at least they did while I was there.
But also, why is it any different why are changes at a repo level all that much easier to track than, say, inspecting changed files and running tests according to what changed?
It supports 5 platforms, but uses 4 completely different build systems, including 2 custom ones (3 if you count depot_tools). There is very little overlap between the platform versions, meaning it's effectively 5 different projects smashed together into a single folder, and pretty much no way to use them in a cross platform project without some serious work. There isn't even a basic abstraction over the similar callback APIs between the platforms, although that's not a huge deal because the effort to write a basic abstraction layer is nothing compared to the effort of getting to a point where you can actually use it in a cross-platform project.
It's also funny that one of the build systems is GYP, which is basically a reinvention of CMake, except it's only used for the Windows build even though it can generate projects for the other platforms. Also, the VS project generator for GYP has been broken for a while (simple typo, trying to import OrderedDict from the wrong module. There's a PR to fix it, hasn't been merged for some reason), so it doesn't even work. Beyond that, it's also broken because GYP forces treating all warnings as errors, with a whitelist of warnings, yet the latest version (since yesterday at least) fails to build (tested on VS2019) because there's a warning that isn't in the whitelist.
You could try to fork it and fix these issues, but depot_tools doesn't provide a way to change the clone URL for repos, meaning you need to dig through the source code and wrap it in your own script that interacts with the internal APIs to do a simple clone (hint: fetch.py has a 'run' method that you can call with a custom constructed 'spec' object, which is a dictionary where you can inject your own url; just look at the hard-coded spec object for breakpad as a starting point). If you don't use depot_tools, then you need to manually clone all of the dependencies in the project since they're not even set up as git submodules.
There's also no versioning scheme whatsoever. Depot_tools seems to automatically checkout the latest version of everything (including itself).
I spent the past week wrestling with this monstrosity. Ended up successfully writing a Conan package for it that builds for Windows and Linux (there's one on Conan center, but it only supports Linux). I have 3 more platforms to go, but I think it'll be a better idea to just scrap everything and refactor into something more reasonable using CMake.
Instead of Breakpad, they also have a newer one called Crashpad, which is meant to improve reliability on Mac OS. Unfortunately, it depends on Chromium, so it won't work for my purposes.
...so all I'm saying is, maybe don't use Google as a role model for your project infrastructure.
/end rant
https://chromium.googlesource.com/chromium/mini_chromium/
https://chromium.googlesource.com/crashpad/crashpad/+/refs/h...
What's the issue you're having with Crashpad? Indeed the breakpad project is a mess by modern standards.
Who on earth can think that what they do, companies with the engineering power as these companies have, has to be also good for it's 30 employee startup just boggles my mind. Not talking about Uber, I have no idea about them and what they do. But I worked for smaller startups just blindly following what they read google or Facebook do and immediately thinking that's the best thing to do too.
It's ridiculous.
How do you orchestrate that? I guess this is the kind of situation where you really do need container orchestration.
"There is nothing more permanent than a temporary solution"
A monorepo means every service is on the same set of versions of every dependency at a given point in time. It's much easier to reason about what is fixed, and move everything to a known good version.
I worked at a similarly large and almost as popular SF based company and I can assure you every innovation there was 100% motivated by somebody's desie to put.it in their CV. Nothing of what was made there made any sense for the business, it was just people playing with toys. And worst part was people introducing this stuff (such as the one pushing for Lerna and an "UI" monorepo would leave the company after 1 year leaving a mountain of tech debt behind. But hey, they had Lerna in their CV now.
I haven't done any calculation on how long that kind of changes can be done with polyrepo setup.
If they can do that, their story can only be better for single isolated changes. It seems from outside, Uber's monorepo development experience vastly better than many small startups with variety of development methodology (no matter monorepo, polyrepo, or continuous deployment, or waterfall development, or GitFlow (a type of waterfall)).
Much easier to apply horizontal and vertical changes which extend outside individual services e.g. update the interface of a service, or fix a bad pattern.