We Halved Go Monorepo CI Build Time
eng.uber.com
eng.uber.com
I’m sure it’s fascinating for people who are into that, and I’m not saying there isn’t a substantive problem to solve here, but I’d personally never go anywhere near a problem like this. It’s too many steps removed from the value. I’d keep thinking “we are down a very deep rabbit hole. A side quest off a side quest. Should we be doing this?”
But I’ve always worked at smaller places, where things like toolchains and sales ops and marketing tech don’t take on huge lives of their own.
Only valuing easily-measurable work is, frankly, a modern organizational disease. It's like searching for the keys under the streetlight—and why? As a substitute for human judgement or resolving disagreements? As a way to make work more legible for executives? It leads to unforced errors and systemic problems, but we can't seem to do anything about it.
My demo is to pick a random crash and diagnose the root cause while talking. The mean time to resolution is low single digit minutes.
I showed this to an entire team, one person at a time, solving bugs as I went. I showed the junior devs, senior devs, and their manager.
No interest. None. Just… silence.
The tools are amazing, but the lack of motivation from the typical developers for learning to use them is even more amazing.
Ref: https://docs.microsoft.com/en-us/azure/azure-monitor/snapsho...
It works even better with DevOps source indexing added to the build pipeline:
https://docs.microsoft.com/en-us/azure/devops/pipelines/task...
A vaguely similar feature is Azure App Service memory leak diagnostic tools which take memory dumps on certain triggers or at intervals.
You can open these in Visual Studio and it’ll show the heap deltas over time.
Most of these couldn't be reproduced either. As in, you'd get a crash once a day in a page that would otherwise work successfully thousands of times.
How would you fix a problem where there's a stack trace only from a release? The scenario is: you can't reproduce the errors, you don't get line numbers, you don't even get function argument values.
I could solve these in minutes using this tool. Could you match that without such tooling?
As for your "side quest off a side quest" observation, I think that is a matter of perspective, about what personally motivates you, and about how work is organized in society. I find that working at a small company makes it easier to have a feeling of focus and purpose -- everybody is focused on one thing. But, I find that working at a large (profitable, well run) company makes it easier to feel like you can actually move the needle and have the resources do large things (and yes, this includes a connection between something like a 'build tools' team and being able to ship a quality product more quickly).
It is easy to miss that every race car needs a pit crew, every pit crew needs equipment and parts, getting parts requires suppliers, parts need to be built, tested, etc., all those people need to eat, all that food needs to be grown, and so on. Organizing all of that "work" into small single-purpose companies, that buy each others' stuff, is one way to go. Another way to go is to organize it into larger companies with sub-teams/groups/departments. So a company's larger mission might be "win races" but, if it is large enough, smaller support efforts in support of the larger goal can make a lot of sense.
I think tech companies should spend a lot more on speeding up their stack. I would love to do something like this (improving developer workflows in a giant organization) as my day job.
But saying that “it’s too many steps removed from the value” appears to be a misrepresentation of the value it adds.
> It’s too many steps removed from the value.
Don't be worried about being too many steps removed from the value. With that logic, CEOs, CTOs, etc are also too many steps removed from the value, yet they are valued very highly. Look at such infra work as multiplying the value of the entire company and in a way, being part of a "CEO team".
And this kind of reasoning is why tooling in most shops absolutely sucks
At scale, rabbit holes can be quite relevant. At scale, you can no longer trust that your personal view of the org is complete enough.
There's still plenty of waste but just because you don't need something doesn't mean it isn't in another team's hot loop. Best to come in with trust first and incredulity second.
There are tools out there, but they're not free and they involve some level of vendor lock in. I've worked on some monorepo build scripting myself.
My ideal VC solution would have the branch control of git with the checkout model and hook configuration of SVN.
I am not familiar enough with the version control ecosystem to know if there’s another product that does it, but I doubt it since I kind of suspect it would violate the CAP theorem.
And in fact, in this very article, the last paragraph actually states that they are aiming to increase the complexity of the change validation process?!?
> Additionally, it helped our team focus on increasing the complexity and features of our change validation process.
Personally, I haven't been super impressed with any of the engineering projects I've heard come out of Uber, but what they're doing obviously works to some degree so I can't be one to judge.
Uber is, partly, a software business. They need to ship software. They get direct value from shipping good software efficiently.
Look at the infrastructure required for that and what we spend on it. It’s massive.
The rest of the article is about their transition from their legacy infra like Jenkins to more modern and flexible things like buildkite (which is of course relevant for anyone in a similar transition).
If your company/project starts with bazel + remote cache + RBE you're pretty close to the end state they're trying to get to.
See also for OSS and paid solutions: https://bazel.build/community/remote-execution-services
With a monorepo, you don‘t need to think about which commits work with which other commits. One commit ID is a full description of every subcomponent.
I've just switched to a company with many tiny repos and IMO it's a huge hassle. There's no automated integration testing, and manual integration testing is a huge pain to set up.
I don't even know how I'd create integration tests that run as part of the CI. What version would I use? If you need to change both sides (very common), coordinating the commits and releases is very painful.
The "monstrously complex" build-and-test pipelines are a significant cost, but the alternative is higher release failure rate and moving slower overall IMO.
No clue if that’s fixed today but it soured the idea of monorepos for me, and that’s not incorporating how often submitqueue would go down.
1 repo == 1 bounded context == 1 isolated unit of deployment.
Which, well, the Uber people clearly have since these big changes exist.
The most obvious counterpoint: many systems of record need to scale command operations (writes) and query operations (reads) separately.
So you absolutely would have separate programs, ASGs etc for those two roles
I've seen people going crazy with this bullshit of monorepos to the point every single directory was a "package" when it could be just a plain module import.
If you want to do microservices, then putting everything back into a single repository and enforcing everyone to use the same version and every change to require upgrades and deploys across the board is totally backwards.
Either do a monolith and have that consistency, or do microservices and allow teams to follow their own rules as long as they keep APIs stable.
Nonorepos + microservices is just a demonstration of everything that's wrong in the technical aspects of our industry. Just applying absolutely everything you read about without even considering if it might be better or not for your specific use case.
At the risk of sounding memetic, the question is not so much "Can you?" but "Should you?". Should you store a monolithic application across multiple repositories? Should you store a distributed application in a monolithic repository?
I agree with @likortera that you should not. And that's because...
> The way you store your source code doesn't have much to do with how your system is deployed.
I disagree with this statement.
Your architecture (monolithic vs distributed) imposes certain assumptions on other aspects of your distribution pipeline. Your workflow (in this discussion "how you store your source code") should support these assumptions, not hinder them.
For example, one benefit of microservices is that cross-functional teams can develop independently of each other. And yet one cited advantage of monorepos is that everyone is on the same version of dependencies all the time. In short, your teams are not independent after all.
Note that I'm keeping the example extremely generic to illustrate this inconsistency, a conflict of interest if you will, that I see people commit in this topic, because in my experience, these questions are not purely technical but involves product/business factors as well. Maybe for most of the people (operative emphasis on "MAYBE", because who am I to judge you), the discussion they need to have first is whether or not they are using the right architecture for their product in the first place.
If you choose to have a microservices architecture, you have to live with the fact that your teams/services will operate at different cadences. If you feel the need to impose a One True Library Version All the Time, then go for a monolithic architecture, and store your code in the same way.
One, teams that share a dependency version can still develop independently; they just share something in common. They already likely share other things in common: deployment target OS, cloud platform and services, shared authN/authZ frameworks, etc.
Two, a monorepo is just the SCM mechanism. As the parent was describing, it doesn't prescribe anything other than the code storage location and how branching, committing, etc. works. Yes, a lot of organizations prefer having a single version rule in their individual monorepos, but nothing about monorepos in general makes this a requirement. You can use multiple versions of the same dependency and still gain advantages from the single commit benefits, and even famous instances like Google's have exceptions where this is the case.
That's not really the case in practice. When people decide to choose mono vs multi based on their benefits, it's become a workflow philosophy in itself. If you choose a monorepo approach but use multiple versions of the same (in-house) dependency across components, you are just opening yourself up for a world of confusion. Sure, you can do it, but should you? Why choose a monorepo structure if you won't take advantage of its benefits?
> teams that share a dependency version can still develop independently; they just share something in common
My point about team independence doesn't mean they should not share anything at all. But rather, they now _update_ together at the same pace because the "atomic commit" that updated a dependency also updated my team's usage of said dependency, for better and for worse. My team might have a reason not to update just yet.
One of the companies I worked for, had the brilliant idea of putting EVERYTHING related to UI/frontend in a monorepo, where almost every single file or two were a different package. I used to joke there were more package.json files than actual js files (it was almost true).
I spent months saying this thing was a terrible idea. Nonetheless the "frontend infrastructure" folks wanted to do some CV padding and play with Lerna and their SV friend's cloud CI service startup, so they went ahead with it.
Months later the big problems started, among which one of the main ones was that they were pushing through every team's throats updates, breaking the product/features those teams were working on, disrupting their roadmaps, accusing each other's of low test coverage, doing hacks and workarounds, shit tons of crazy CI scripts for all the corner cases, much longer deploy times, most dev environments were a lot slower, deployment issues because now we had to deploy several different projects at the same time, etc, and of course not being able to upgrade to latest React because some team in the corner had an issue with it and they didn't have the time at the moment to fix it.
How did they solve all of this? In the span of 3 to 4 months they left the company. All four of them. Leaving behind an incredible amount of technical debt and nearly every frontend team totally fucked up.
What irks me is that some of these guys are pretty popular "youtubers" and spend their days giving talks of how great their work with monorepos and "frontend infra" is. They don't tell the messes they've caused of course.
Monorepos might be great if you're Google and have the resources and talent to do it right. For most companies out there, it is just creating a centralized problem that will eventually block everyone.
I'm a big proponent of monoliths, specially while you're not a > 300 person company. But if you're splitting your teams and services, then agree on APIs, don't break them, and let each team follow their own pace, with their own tools, and with their own schedules and preferences. Otherwise stick to the good ol' monolith and just separate things into modules/imports/whatever.
At my current employer, the main application has a "plugins" system with a very flexible and stable API. Every team around is just building "plugins" that can be installed into the main monolith, depending on each customer needs. This works fantastically well for a company with more than 1k engineers. No monorepos, no coupling, no interdependencies between teams, no parallel deploys, and each team manages their own destiny more or less.
The development view need not be connected to the deployment view at all, so unnecessarily coupling them can lead to worse outcomes. On the other hand, having the ability to couple them initially _and decouple them again in future_ works wonders for scaling.
I am consistently amazed that people with opinions on software architecture do not seem to recognise this seminal paper on the topic, or have the ability to re-synthesise it into tactics.
[1]: https://www.cs.ubc.ca/~gregor/teaching/papers/4+1view-archit...
The more clients a library or service has, the more expensive this is, and it's an ongoing maintenance cost for every client. Since people changing the service don't feel the full pain of this maintenance cost, changes keep on happening, and eventually clients get culled because it's too expensive to keep on maintaining them all.
From the outside, this looks like the company abandoning venerable but still working product, and makes people scratch their heads, wondering why.
People often conflate these two things because often teams actually want to have centralized dependencies (so that you're forced to update or die as you said). If that doesn't work for you you can choose to have modules (or groups of modules) keep their independent set of dependencies, all while keeping the code in the monorepo.
But that said, no I don't think you need a meta repo. You need something. Call it a Change set. MVP would be say merge commits in 4 repos. If any fail then revert. If they all succeed then deploy in a specified order.
* Breaking changes can be done at once. Very helpful for runtime deps.
* No chance that a repo is out of date.
* Upgrades are atomic (may be hard to test a system in a half state).
'A well-designed system should not have X therefore you don't need Y.'
When 'A well-designed system should not have X' is just a matter of opinion, or doesn't acknowledge trade-offs that might make X the better option than Z, then this isn't a useful argument against Y.
at the risk of coming up with a contrived example: let's say you own a service that need to deserialize a datetime in a request in a format you don't currently support. assuming you own the stack, you need to a) update your date library b) update your webserver stack c) possibly update an intermediate webserver stack that includes primitives like logging, telemetry, tracing, auth, service discovery and d) your actual service.
If a->d are all independent, separate components, you have to orchestrate those changes through 4 separate repositories. And god forbid something you did at the lowest point in the stack is completely unworkable higher up.
There's all sorts of rocket science you could do to orchestrate these changes, but it ends up being contrived and edgecasey.
Most of the pain from monorepos can also be addressed with a dash of rocketscience (see:bazel), but the end model tends to have
a) have an easier mental model for the user
b) allow for consolidation of infrastructure work. Ie, your build/ci/language tooling teams can focus their efforts on one place
c) can coordinate changes across the entire stack within one field of view
d) can coral some of the worst, disparate instincts of a growing engineering org (ie, tons of teams optimizing for local maximas without internalizing knock-on effects).
e) fewer weird edgecases.
I like having examples, even contrived ones, but I'm not sure I understood this one. Can you elaborate on what you mean? Is it about adding support for a new serialization format for dates in requests to a service? Why would this affect the webserver stack and logging/telemetry/tracing/auth primitives?
I find that a lot of organizations have really strange thoughts on how to factor things into separate microservices and libraries. Usually I approach by asking the following question: if this was an open source-library or service (e.g. like elasticsearch), would you use it? If not, then its probably not a great candidate for a separate thing - lets try and come up with something else.
One way to handle CI/CD is using standardized pipelines e.g. you tag your repo with a tag `app:node` or `lib:js` and the github org pipeline scanner will find it and assign the standard `app:node` or `lib:js` pipeline to it.
A way that I like better but most tools unfortunatley don't support it yet is for the infra teams to publish libraries that are essentially functions taking some parameters and generating (standard) pipelines/configuration. Those can then be tracked together the same as other dependencies.
Besides the "N pull requests" problem you now lose history whenever you move a file across repo boundaries, and you'll eventually have straggler projects staying on old versions for years - so switching to a new way of doing something essentially means supporting both versions forever. Code for common needs gets duplicated, or worse, split out into yet another repository and /then/ duplicated, because no one wants to check if any of 100 repositories rely on the buggy behaviour they want to fix.
I find this to be this a level headed explanation of the advantages of monorepos: https://danluu.com/monorepo/. I'm surprised to see so many comments summarily dismissing them, as if sanely managing thousands of smaller interdependent repositories doesn't require at least as much investment in custom tooling.
The project should stay a full monolith until the factoring is more clear.
> I find this to be this a level headed explanation of the advantages of monorepos: https://danluu.com/monorepo/.
There are two issues I have with this article. One of them is when it describes drawbacks of multiple repositories, its not specific enough e.g.
> That sounds like it ought to be straightforward, but in practice, most solutions are cumbersome and involve a lot of overhead.
The other is that it assumes you have Google scale of resources to throw at the problem. If thats the case you can make anything work / monorepos or polyrepos. The issue is that small-to-medium sized organizations will not be prepared to invest the amount of resources needed to keep a monorepo working well, as the org will largly need to rely on existing available (OSS) tools which often have poor monorepo support. (Bazelifying everything has a significant cost, bazel rules are often not generic enough to work with e.g. the variety of JS ecosystem tools)
If one business change requires code changes to ~every module in your source tree, then of course microservices, and therefore separate repos per service or whatever, make no sense at all.
Sure, each project team could upgrade independently, but the company that’s chosen a mono-repo for their go code is likely to desire to have a single team tackle this upgrade.
In practice, you often end up with larger changes which are mostly janitorial. Bumping a version of a log library across the infrastructure, or a version of a serialization system. This simplifies dependency convergence and a best handled as if they are cross-project changes.
The other problem is that over time, good factoring tend to deteriorate. A large software project will invariably have people with different brains working on it, and they'll have different needs for the factoring. So the project naturally pushes itself toward situations where cross-project changes become a necessary thing.
Good factoring can probably be quantified too: dependency chains should be shallow and graph connectivity should be low. The more dependants a module has, the more stable its API contract should be.
i’m finding it hard to think of a system at a large company that would have no internal libraries used by multiple projects, that wouldn’t require ever making a cross project change.
What helps maintain polyrepos sane
1. A healthy methodology for dependency management
Evolve APIs with deprecations. Decide on a healthy amount of time for a deprecations to live. Set up alerts when deprecations reach certain amount of time. Set up dashboards (e.g. Grafana) to track dependencies. Help other teams update, and think about how to make updates less painful.
2. Define good boundaries and API contracts to adhere to. The most important thing for a good API contract is stability.
3. Don't prematurely split into microservices. Better ideas for stable API contracts emerge the longer you can wait.
1. I don't see how this is a problem specific to polyrepos. Have an "open PRs" link in the onboarding handbook that gives you a view of pull requests from all repos in the organization. GitHub automatically shows you notifications from all repos. If engineers still chose to focus on one or two repos after that, I'm not sure why.
- Have a (Grafana) dashboard where you can see the latest / newest stuff. Use standard GH tools you use for OSS, such as follows etc to keep up.
2. Don't prematurely split into multiple repos. "No monorepo" doesn't mean not having poly-package repos. It means thinking what the sensible (library or service) API boundary is - treating your projects as you would treat library / service development. In this case a separate repo with lib3, lib2 and lib1 sounds like a good way to go - at most one repo per orthogonal internal framework (e.g. core-react-components). Repo dependency chains should be as shallow as possible, and differenting between public and internal packages is important.
3. Help other teams upgrade. If you are responsible for repo A, once you publish a new version tagged appropriately with semver, use the dashboard to look at your dependants and work with them (or rather, for them) to upgrade. Think of your dependants as internal customers, and make sure you add enough value for them to justify the upgrade effort. Cultivate a culture that values updates.
4. There are other alternatives to `npm link` e.g. see `yalc` https://github.com/wclr/yalc
Another pet peeve of mine is that the real issues get lost when you try to generalize. The article attempts to do this but that makes it hard to evaluate its claims. The best way to evaluate (alternative) solutions is to take a more concrete example repo.
For scalability of this model I'll just point to the OSS community; individual maintainers often several dozen active repositories, but also they have an API contract worthy of a documentation website, versioning scheme and planned deprecation, and they typically avoid cross-project dependencies
No, that way lies sorrow and despair. There should only be one version of any dependency being used in your company. Any deviations should require like VP-level approvals or something.
But when your team owns a repo, then at least the damage is contained within your team.
Google solves that problem with heavy NIH syndrome (its hard to get promoted by utilizing an external third-party lib, better develop your own), and writing tests, yes. And for those few third-party libraries that google still depends on, updating them is a big PITA.
When teams take the stability & versioning of their APIs seriously, the need to use monorepos to share that info is greatly reduced. A multi-repo approach is perfectly feasible when all components are working to established APIs, which also alleviates the issues mentioned in the article.
- easier to share and import common packages
- proto and thrift files are kept close to services and clients are updated globally automatically
- dependencies and go versions are managed globally and all services get the same security updates
- standardized build processes make it easier to manage large deployments
And honestly, after the initial repo download there were no visible downsides
Is the suggestion just "an easier way to avoid this problem is to not have so many tests?"
They also use mostly homegrown CI tools alongside phabricator for code review, or at least they did while I was there.
But also, why is it any different why are changes at a repo level all that much easier to track than, say, inspecting changed files and running tests according to what changed?
How do you orchestrate that? I guess this is the kind of situation where you really do need container orchestration.
Who on earth can think that what they do, companies with the engineering power as these companies have, has to be also good for it's 30 employee startup just boggles my mind. Not talking about Uber, I have no idea about them and what they do. But I worked for smaller startups just blindly following what they read google or Facebook do and immediately thinking that's the best thing to do too.
It's ridiculous.
It supports 5 platforms, but uses 4 completely different build systems, including 2 custom ones (3 if you count depot_tools). There is very little overlap between the platform versions, meaning it's effectively 5 different projects smashed together into a single folder, and pretty much no way to use them in a cross platform project without some serious work. There isn't even a basic abstraction over the similar callback APIs between the platforms, although that's not a huge deal because the effort to write a basic abstraction layer is nothing compared to the effort of getting to a point where you can actually use it in a cross-platform project.
It's also funny that one of the build systems is GYP, which is basically a reinvention of CMake, except it's only used for the Windows build even though it can generate projects for the other platforms. Also, the VS project generator for GYP has been broken for a while (simple typo, trying to import OrderedDict from the wrong module. There's a PR to fix it, hasn't been merged for some reason), so it doesn't even work. Beyond that, it's also broken because GYP forces treating all warnings as errors, with a whitelist of warnings, yet the latest version (since yesterday at least) fails to build (tested on VS2019) because there's a warning that isn't in the whitelist.
You could try to fork it and fix these issues, but depot_tools doesn't provide a way to change the clone URL for repos, meaning you need to dig through the source code and wrap it in your own script that interacts with the internal APIs to do a simple clone (hint: fetch.py has a 'run' method that you can call with a custom constructed 'spec' object, which is a dictionary where you can inject your own url; just look at the hard-coded spec object for breakpad as a starting point). If you don't use depot_tools, then you need to manually clone all of the dependencies in the project since they're not even set up as git submodules.
There's also no versioning scheme whatsoever. Depot_tools seems to automatically checkout the latest version of everything (including itself).
I spent the past week wrestling with this monstrosity. Ended up successfully writing a Conan package for it that builds for Windows and Linux (there's one on Conan center, but it only supports Linux). I have 3 more platforms to go, but I think it'll be a better idea to just scrap everything and refactor into something more reasonable using CMake.
Instead of Breakpad, they also have a newer one called Crashpad, which is meant to improve reliability on Mac OS. Unfortunately, it depends on Chromium, so it won't work for my purposes.
...so all I'm saying is, maybe don't use Google as a role model for your project infrastructure.
/end rant
https://chromium.googlesource.com/chromium/mini_chromium/
https://chromium.googlesource.com/crashpad/crashpad/+/refs/h...
What's the issue you're having with Crashpad? Indeed the breakpad project is a mess by modern standards.
"There is nothing more permanent than a temporary solution"
A monorepo means every service is on the same set of versions of every dependency at a given point in time. It's much easier to reason about what is fixed, and move everything to a known good version.
I worked at a similarly large and almost as popular SF based company and I can assure you every innovation there was 100% motivated by somebody's desie to put.it in their CV. Nothing of what was made there made any sense for the business, it was just people playing with toys. And worst part was people introducing this stuff (such as the one pushing for Lerna and an "UI" monorepo would leave the company after 1 year leaving a mountain of tech debt behind. But hey, they had Lerna in their CV now.
I haven't done any calculation on how long that kind of changes can be done with polyrepo setup.
If they can do that, their story can only be better for single isolated changes. It seems from outside, Uber's monorepo development experience vastly better than many small startups with variety of development methodology (no matter monorepo, polyrepo, or continuous deployment, or waterfall development, or GitFlow (a type of waterfall)).
Much easier to apply horizontal and vertical changes which extend outside individual services e.g. update the interface of a service, or fix a bad pattern.
Zuul (https://zuul-ci.org/) was created for openstack to solve the issue of optimistic merges / PR queue testing.
When you use buildkite with own containers on AWS ecs you can use efs to do a git clone with reference. (https://git-scm.com/docs/git-clone#Documentation/git-clone.t...) Essentially what they do with a packed base repo, but you only end up sending what you need, not more.
The binary cache is available in other flavours too. If you don't use go, then sccache (https://github.com/mozilla/sccache) may be useful.
[0] https://buildjet.com/for-github-actions/blog/a-performance-r...
edited: wrong link
So even though you can more easily horizontally scale and handle infinite requests, the latency of each request will be much poorer than if you were just running on better hardware.
Naturally there is the caveat that using binary libraries isn't a thing in Go.
You either:
* Just don't test such changes at change time, break down and push the testing effort to downstream projects in disguise of "dependency management".
* Structure your engineering effort to avoid such changes, for example avoid having a base library shared between teams.
Neither is ideal for a corporation environment, because both harms velocity. We accept these in the "open source world" because we don't have better options, but the same does not hold for corps.
And now you get into the fun situation of having to handle ecosystemic asynchronous updates between your multiple repositories.
Our services aren't even particularly well-factored and we update median 2, mode 1, mean 3-4 when we change a core library. (Out of ~25.)
This isn't true in my experience at large companies. For this to be true, you would either need to have absolutely no shared core libraries, or core libraries that are so stable and unchanging that they never receive updates. Updates to those core libraries usually result in a cascade of changes across a ton of repos, tons of tests that need to be rerun, and a slew of now incompatible versions.
All of these problems go away with a monorepo. Monorepos are absolutely the right solution for 98% of companies out there because they will never reach the engineering scale where monorepos start to struggle. For the 2% of companies that do reach that engineering scale, these kinds of blog posts get written.
The mistake is reading blog posts like this and thinking is necessary or even relevant to any more than 2% of the companies out there. The vast majority of shops will never need anything like this.
I don't really see this as a mono vs. multi repo issue either. In a mono repo you can still choose which projects update and the more you choose the bigger the diff and more involved the review process. In a multi repo case you can build analogous tooling to generate downstream MRs automatically.
If you need to touch every downstream project regularly just to keep things functional, you don't have a mono-repo, you have a monolith. (Which can also be fine, really! But it's not the same thing at all.)
(Yes, I know with 25 projects we're nowhere near Uber scale. But based on my experience at larger companies, this seems to scale more or less linearly - in terms of downstream project count, library count, and developer time available to dedicate to such things. If anything, developer time available is what seems to scale super-linearly.)
Even a security fix is unlikely to affect every downstream user of a library except in egregious cases. But, if so, yes you update them all. (I'd say this is much less frequent than, I don't know, OpenSSL or log4j having a bug that makes us do this. The specific concern of an internal library having such a broad vulnerability is negligible in influencing our CI design.)
> What happens when they inevitably do want a new feature, and have to fast-forward through months or years of interface updates all at once?
You do it.
> What happens when a developer wants to work on a feature that cuts across systems, and now has to re-learn the "old" way of doing things?
You do it.
I didn't say "it has no associated costs." But you need to weigh those costs against the other costs of a monorepo, and other costs in your CI/CD design generally. "Take longer to update a really old project once a year" is a much lower cost for us than "have a dedicated CI team to wrangle the tooling we need for automatic downstream pushes / a monorepo."
It's unusual because a multi-repo environment makes it difficult to do, not because it's not necessary.
> We do that with all external dependencies.
Indeed. Why you’d want to also suffer that for internal dependencies when there is no reason to, I can’t fathom.
If you have a good process for external dependencies you can apply that to internal. If you don't have a good process for external dependencies a monorepo won't help (unless you make all your dependencies internal, which is even more expensive)
Yeah... but usually good process for external dependencies requires a lot of ceremony and takes maybe days to update one dependency. If the same happens to internal dependencies we will be just slower at making changes for no good reason.
This is more of a cultural thing. In fact, we are creating complicated solutions to mitigate this problem, only to go full circle after a while: people implemented Service Mesh via proxy sidecars, it added more layers (than just embedding this into the RPC library), but even doing so is easier than persuading downstream project owners update their dependency in a timely manner, so it takes off. And now they are going to chase for "proxyless" due to "performance". Oh well.
> the extra time to make multiple PRs is dwarfed by the actual update's complexity
Writing code is never the majority of the work. All the operational stuff is more complex without forced upgrades: coordinating the dependency updates, automating integration testing, pushing clients to migrate off old thick clients, determining what version of a specific repo is deployed, running a bisect to root cause issues, etc.
Also, having a rather global view of the impact of my code before actually make the change helps a lot, especially if you are working on low-ish level libraries. Technically this is not bound to monorepo, but people usually criticize the complexity introduced by having such ability for monorepo, so :)
At that scale, anything reducing complexity is a bonus. At that scale, monorepos are great for that.
You call it Merkle-style tree despite the fact it is obviously a DAG?
--
Edit:
> Large changes are also more prone to outages. If we land them outside the working hours, there would be limited resources to mitigate potential outages. To prevent this, we ended up asking engineers to get up early to deploy these changes
Huh, landing changes to HEAD and releasing to production are not decoupled?
All trees are DAGs so that's not really adding much.
And describing it as a Merkle-style tree conveys additional implementation about the implementation - presumably the tree stores the hash ("computed from all of its source files and inputs") which just saying "it's a DAG" doesn't tell you.
Or just avoid the fancy names and call it what it is: they walk the dependency graph to identify build targets that have a changed dependency.
Perhaps you’re complaining about it not being a tree because there can be multiple roots which seems fair enough.
[1] I think there are many ‘roots’ as roots are going to be leaf build targets like various executables or test results
https://eng.uber.com/research/keeping-master-green-at-scale/
Picking a language that has fast build speed, allows you to be more flexible, and able to test and deploy quicker
They had Go, wich builds very fast, but they failed to educate their developers to maintain a healthy build pipeline to avoid things getting too slow
This kind of developer will cost your company millions, make sure to educate them properly ;)
I'd never work for a company with slow build speed, part of the reason why i refuse to touch languages like Rust and C++, due to their insanely slow build speed, i refuse to live in a world like that
Use multiple repos. It doesn't hurt you know. And you can have proper granularity
The costs of monorepo seem to be much higher than just dealing with inter-dependencies. And you can always group projects if it makes sense
I'd be happy to see a counter-example, but I've not seen one yet.
Heck, even most individual developers use multiple repos.
Monorepos are usually ok (and even talking about team/project granularity) until you outgrow it.
Buildsystems like one-repo-per-project much better than monorepos. (Shipping a fix doesn't require rebuilding "the whole world" for example)
Seems to me that Git wasn't thought out for "monorepos" at first place as a VCS. Also, these companies should either use something else or not do monorepos.
Does this play nice with gopls (go language server) allowing you to jump around to definitions?
>is used most of the time with svn
It supports mercurial, git, and svn.
>I 'm curious why they opted for it in the first place, instead of a more modern platform.
It is a modern platform and works well.
„We solved an issue we created ourselves”.
Which sometimes is ofcourse what you have to do.
I was new to Phabricator, but I picked it up quickly. Even though it might not constantly pump out new features, Phabricator gets the job done.
Only very recently. Obviously after Uber adopted it.
> from what I remember and is used most of the time with svn
Nope.
> I'm curious why they opted for it in the first place, instead of a more modern platform.
Because it's quite good? It's not some archaic thing like SVN for which there's an obviously better newer option.
Having said that it is much harder to learn from others' mistakes than your own. Learning from your own success is hard too. Learning from others' success is the hardest by far. At least we have a glimpse (albeit indirectly) in this article of the mistakes they made.
I can't imagine deploying a functionality that needs to interact with 20 other services, if I can't have some zero-cost assurances like Go's typing and having all IDL available at developing time. This is of course is even better if at compile time I can check all contracts are still honoured. Monorepo is a really "cheap" way of getting all this for, close to, free.
If a product feature X can't be delivered without changes to 20 services, the architecture is too granular.
But the ~main whole point of defining distinct services which exist on different ends of a wire is so that they don't need to stay in sync at the code level!
As you say, a service implements a contract, written as an IDL or a JSON schema or informal convention or whatever, in order to express a promise to consumers. That promise needs to be kept as long as any consumer relies on it. If your search service publishes a protobuf that has a service definition called e.g. SearchV1, then your search service is absolutely obliged to keep supporting that SearchV1 service until it's no longer used by anyone.
If this isn't the case, and you can deploy changes to a service that violate the wire-protocol contract you establishes with your consumers, then this is a problem. And it's easy to do. I find that many programmers don't fully understand the weight and implications of the interface they define between client and server.
What you say is also true
- applications: front, backend, background-worker
- libs: database-orm
In a multi-repo layout, if you want to make a change in database-orm, you'll make your PR in its repo, test, and make a release with your changes.
Nice and easy right ? Well, you're not done. Now you have to make a PR to update the dependency on each repo using this library. If you're lucky, nothing breaks and it's quite quick.
But it's not always so easy : you notice that you actually broke something down the line in the backend. You have to fix it (in your library), and do it all over again. You can also have libraries depending on other libraries, multiplying the effort when you messed up something.
The monorepo handles that, you update one library and you can see what you broke down the line, and fix all of that quicker. Also, changes are easier to follow since one modification impacting several applications or libraries can be made into only one commit.
You can tell me that libraries should have a nice definition and the applications should be independent from the actual internals, but that's rarely the case. It's a tradeoff and lots of companies are going this route.
The processes you describe are great to make teams think twice how to go at a change, instead of "let's change everything and see how it goes" attitude.
I built a tool to bring up every microservice locally so you could test every change together so the separate repository problem went away. It coordinates vagrant LXC environments.
The version I built at my employer was integrated with chef and Ansible. The version I built at my employer handled cloning and pulling dependencies too. And could build in parallel and deploy to local load balancer haproxies.
But my open source version is barebones by comparison. In the future I shall try build developer tools open source rather than trapping work at my employer.
HTTPS://GitHub.com/samsquire/platform-up
Not entirely. Since the fleet isn't atomically updated to the next version you have to be careful about multiple versions being compatible with each other.
The text itself can be changed but it takes at least 15 minutes to mutate to a deployment. We still cannot generically mutate running code to other running code. Ksplice and other live kernel patching and Chrome's binary patching should be generalised and productised.
Patching methods at runtime is another thing that is possible (ruby, python and Erlang) but I am not aware of a general framework for deploying mutations to servers at runtime.
But yes - GitFlow or Environment Branch HELL is so common in the industry - its hard to talk about better approaches.
The solutions to monorepo taking 10x longer to build seem to be "Spend 10x longer on the manual developer parts doing 10x more procedure with more tools, more rules, more work"
One place change == X possible conflicts.
Ten places over longer period of time (a lot bigger PR) == XXX possible conflicts.
All depends on the size of the team and project.
Tho its hard to find a good place to switch from one approach to the other one.
I do not understand your point about conflicts as 3 commits across 3 projects with 3 different CI pathways is going to cause more conflicts than 1 commit across 1 repo. In my experience managing one code change or project across 3 repos is a 10X difficulty increaser in terms of repo management, conflicts, etc. It's not just 3X harder, it's 10X harder to me. The number of times I've seen a spelling mistake/naming difference/etc in 1 out of 3 repos because the PRs were done separately and no one noticed is too damn high.
The simplicity of having it all together strongly outweighs the benefits of multi repo in most situations IMO. The number of projects/companies/etc that would benefit from some highly engineered microservice-based multi-repo monster is probably less than 100 in my country, and 1000 worldwide.
This is the issue with monorepos -> That over time the more ppl work on them, the higher the chance of the CICD to fail due to conflicts. And till now pretty much NOONE in the world fully resolved the conflicts issue.
Companies introduce merge queues which cripple productivity, because every next merge "can" (but doesn't have to) break all previous PRs.
Instead of having a very easy Feature-Branch pipeline, you have to build some abomination that becomes a bottleneck as soon as the company starts growing.
--- Like I said in other comment - majority of IT still lives in GitFlow/EnvFlow hell. Some of us learnt from those mistakes and do better now.
Why build libraries and separate services if an update on some part of the codebase needs changes across all the platform?
This trend of microservices + monorepos sound to me like taking the worst part of everything.
If you do microservices, then let each team have their own stack, their own tools, their own libraries and be independent and just agree on the APIs.
If you want everything super consistent and share as much code as possible, etc then just do a traditional (well architected) monolith, maybe deploy it with different configurations for different scaling needs, etc.
Monorepos are to me a symptom of worse problems.
I worked for a.company where they had a rule that "every repository should be a monorepo". They didn't even understand everything that was wrong with that rule.
Yes, you get the benefits, but also all the drawbacks. Those drawbacks might make sense for extremely large teams. I want to cry every time I see a 20 engineers teams wasting company's money with this.
Since this is Go, it's trivial to point in-development branches of downstream projects to an in-development branch of a dependency before merging it for those projects' testing processes. (This still doesn't cover 100% of cases of course, but nothing does. The better answer is not to over-factor in the first place.)