Monorepos and Forced Migrations
buttondown.email
buttondown.email
I've worked at places that have done it both ways (forced migrations vs. ways to do migrations later when convenient), and when you can do the migration later, it often ends up being much later, sometimes infinitely later (i.e. never).
Do you want all the pain of migrating now? Or do you want to spread the pain over a longer period of time?
If the breakage is due to incompatibility (the API behaves differently, etc.), then breakage that the migration caused is probably still going to happen if you do the migration later. So you're not reducing the total amount of breakage/pain.
And if you spread it over a longer period of time, you are increasing the maintenance burden on the team that needs to provide compatibility with the old and new ways.
Everyone was forced to do their migrations (exemptions would generally be approved but only once). But they had fair warning and could schedule it to fit in with the rest of their work.
That being said, it's possible to do version pinning carefully at the executable level, just so long as each executable has exactly one version of each dependency in its dependency graph. But then you have a combinatorial explosion of dependency sets to evaluate and support. Theoretically possible, but not very practical. ABI stable languages and languages that ship as source code have a much smaller analysis space since compatibility can be assumed in many more situations.
Or avoid dynamic linking altogether.
Rust lets you have multiple versions of the same crate living in the same executable.
The problem starts when different dependencies have different opinions about ABI-important details. Like layout of data structures, layout of v-tables, use of concurrency primitives, or even stack utilization.
I don't have as much depth of knowledge in cargo (rustc probably can't avoid these problems). From what I can tell, it prefers resolving dependency sets per executable, which means it would be opting for the combinatorial explosion of compatibility sets problem. That's a valid tradeoff to select, though I am extremely skeptical that an organization with 30k Rust projects would find that dependency juggling is a non-problem.
Among other things it prevents diamond-dependency problem and other dependency issues. Google binaries are already regularly hitting the upper size limits for linker. If you needed to link 5 versions of each library it would greatly exacerbate the problem.
Having a monorepo also greatly benefits maintaining security. You only need to fix a security issue in the trunk, and the fix will automatically be applied in every binary that uses your library upon its next rollout.
> all to optimize consumer behavior to boost ad revenue.
I've been working for Google for more than 12 years and never ever have I thought about optimizing ad revenue. Not just that, but I also never heard about this consideration in any meeting that I attended over the years.
> Did I mention the report was commissioned, in part, by Google? (End of Aside.)
I don't see how this in any way invalidates the conclusion. It's not like this report was commissioned for marketing purposes.
> note git is not allowed at Google
This is very misleading. All Google open source projects, including Chromium, AOSP and all of github.com/google use git. What is not supported is a git frontend to the internal VCS. This is mostly due to the fact that support of any additional tool required staffing, so only a limited number of configurations/tools can be supported.
It's super weird to me that Google wouldn't staff this. Given how much of the external world uses git, I would have thought that keeping new hires productive would have made this staffing requirement a no-brainer.
But what do I know, I don't work for Google. I'd be super interested if anyone knows why this is the case, though.
You have two interfaces to the version control system: one native, that evolved from Perforce, and the second one based on Mercurial, which is pretty similar to what can be done with git.
At some point there was a decision to be made whether to support Mercurial or git, because they have very similar feature sets. I don't quite know why they made the choice they made.
To clarify, all of this only affects the client side of the VCS. It is not really possible to use either git or Mercurial for repository storage without some major trickery.
Thanks so much for the extra information!
[0] I have no real opinion on Perforce, but I always thought that the reason game studios used it was to deal with binary assets, and I'm surprised that Google would have lots of those.
Here's an article that I've found: https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...
It was mostly technical one. Piper itself is a distributed file system which has significant incompatibility with git's internal model and git didn't provide a good extension model to hack with. So to provide deep integration, forking was the only viable option at the moment.
Perhaps the equation might be different now since MS did some contribution for their own use and git maintainers have become less reluctant to accept features for monorepo use cases, IIRC. Still, it's harder to retrofit the existing distributed file system into the core git object model than to develop a new one designed with that in mind.
At any rate, it's extremely simple to move to Mercurial from the git world, particularly with trunk-based development.
What the author seems to be objecting to is that teams that manage upstream deps make breaking changes and those changes become interrupt-driven work.
From the perspective of the upstream devs that makes their lives tractable, because people get off of old versions and they don't have to support it, and security teams do not have to audit old code, and that prioritizes burning down tech debt and not putting it off endlessly.
The work that the author is complaining about is work that probably needs to happen one way or another. The monorepo as a forcing function to burn down tech debt seems to actually be working there. It would probably be a 10x bigger nightmare of years of tech debt if they didn't do that.
There might be tweaks that could be made to the process so that there were windows every year where breaking changes could be made so as to batch up the breaking changes coming down the pipe -- and require a variance if someone needed to push an emergency breaking change due to an external requirement.
A company shouldn’t make it impossible to make breaking changes but it shouldn’t make it easy either.
Semver is a good tool to build norms around. If your two year old project is either 0.x or 11.x you should probably be a bit embarrassed.
And I suspect you're going to find that inside of Google these are 10+ year long projects which were not designed well enough for that (because honestly nobody does that). And the breaking changes are most likely largely all well intentioned. Without at time machine its just work that needs to get done in one way or another.
SemVer also has its problems, and is not a panacea, for example:
https://hynek.me/articles/semver-will-not-save-you/
And I'm positive Google is aware of SemVer.
I still think that semantic version numbers are still helpful (for released versions; unreleased versions can just use the hash), even though it doesn't always help.
If you do depend on undocumented behaviours (or behaviours that are documented to be changed in future), you can specify that you require only one specific version, or a range of versions that they have been tested with.
They say "version numbers are unique, orderable identifiers of software releases", and I agree, but you can still do that and still use semantic version numbers too.
And while reducing deps makes sense to any developer experienced enough to have been bitten hard by it, Google faces the problem that due to its size there will be lots of highly specialized code dealing with issues that would too trivial for nearly any other businesses to bother extracting out into dependencies. You should always make things as simply as they can be and no simpler. At Google's scale simple things will be necessarily highly complicated and specialized.
You're also still talking about issues of concern to you which don't have any bearing at all on monorepo-vs-manyrepo. Fewer deps are clearly better in either model.
Still, I find this writeup to be a good glimpse into culture at Google that one engineer found to be, at least, more complicated than the rose-colored glasses through which we usually read about Google's engineering practices. This criticism also mirrors what people have said about Google's tendency to kill off their products that people depend on, which causes reticence on the part of users to adopt new products. (related: https://killedbygoogle.com/)
Something I've seen happen – management issues a directive, "Every team must use SemVer it is our standard". Now every team has version numbers 1.x.0. Almost nobody ever bumps the major version above 1 – people just make breaking changes on the minor version. There is no way to enforce the requirement to bump the major version, so no way to tell if people are doing it when they are supposed to (I know there is some tooling to detect backward incompatible API changes, but a lot of places don't use any such tooling, and even if they did, it will always be incomplete, there will always be backward incompatibilities it won't be able to detect.) So long as every team's version numbers look like SemVer, management can be told "Directive implemented, everyone is using SemVer now"
I wonder why OSS projects bother with 0.x.x. If you are never going to change the initial 0 to something else, why not just drop it? So you can do whatever you want while paying lip service to SemVer? Why not just do whatever you want, and ignore SemVer entirely?
Possibly due to being unsure, thinking too many changes will be made and they don't want to drop it.
> Why not just do whatever you want, and ignore SemVer entirely?
If your version number is a single number which always changes, then it will still be compatible with SemVer (since it will be the first number, and the other two numbers are then zero), although not very well in case sometimes it still is compatible you will not be able to upgrade without checking this.
However, if you do that, some people will start saying "this component is really unstable, they keep on making backward incompatible changes to it". Some people just look at the version number.
I don't like SemVer because I think it encourages focusing on the version number instead of the actual version contents, and because it puts all the focus on the backward compatibility of the public API, when that isn't always the thing consumers of a component should be focusing the most on. If I remove some obscure long-deprecated public API which nobody was using anyway, then following SemVer strictly, I must increment the major version–even if that is literally the only change in that release. Meanwhile, I can radically rewrite the implementation, even create a brand-new backward-incompatible next-generation API – but if I also include a compatibility shim which maps the old public API on to the new one, and if (to the best of my knowledge) that shim is complete (even if it actually contains subtle regressions my testing has failed to uncover), then following SemVer strictly that would be a new minor version only. Consumers should be much more concerned about the second release, but SemVer will mislead them into paying more attention to the first instead.
Allow marking old version numbers as "regressive"; any regressive version number is never considered compatible with any other version, whether regressive or not.
Allow a version to have multiple version numbers that alias each other; that might be appropriate in the case of radically rewriting the implementation in a way that is supposed to remain compatible.
However, that won't solve everything (nor does it even come close). It is necessary for package maintainers to notice if a version has security problems that older and/or newer versions don't, and to do whatever is appropriate in the given situation. Sometimes a program is compatible with multiple versions of a package even if they are not otherwise compatible (it is related to what you describe). Sometimes a program depends on internal details (or bugs) which are subjected to being changed. Sometimes there are other reasons why a user might not want to upgrade (including hardware compatibility). Sometimes conflicts are possible; depending on the situation there might or might not be a way to resolve this.
So, like some other messages mentions, using semantic version numbers won't solve everything, but they won't break everything either. Use or don't use semantic version numbers according to your choice, but whichever way you choose should be documented, and anyone using them should be aware of these considerations.
FOSS is helpful that you can fork software and modify it for your use if needed, in case you want some but not all of the changes that have been made, or if you can make your own improvements.
> The last two words are the crux of this issue: monorepos deliberately centralize power.
No it doesn't, it democratizes decisions. In a monorepo when someone breaks an API you have the power to roll that breakage back. In a typical environment you will have to deal with that breakage sooner or later anyway without any say on the matter. So a monorepo makes it much harder and costly for library maintainers to break downstream dependencies, while in a typical environment they can break their API every new commit and create a ton of work for everyone else without a second thought.
"But I don't have to migrate without a monorepo!"
Yes you do, but without a monorepo the requirement to migrate wont come from other engineers, but will come from some guy high up who demand that now everyone must switch to at least version X to reduce security issues and the huge maintenance burden of maintaining Y separate versions at once. So if you think that engineers in non monorepo environments doesn't spend a ton of time migrating dependencies then you are wrong. They migrate just as much, but they happen in bigger chunks and gets way more complicated.
I worked on a microservices project where downstream projects put "contract tests" into upstream dependencies. Ie, you can change your dependency if you want, but it has to match the contract/API we agreed with you or your change doesn't go in.
The whole git/mercurial thing is, as I understand it, related to the extensibility of the systems. Mercurial is hackable and extensible in ways that matter, git isn't (or at least wasn't)
The one version policy only really matters for third party dependencies. If you're making a change to a library that is entirely within the monorepo, you can do a three phase migration (add the new functionality, migrate users, deprecate the old functionality). Google is exceedingly efficient at this, and the vast majority of the time these migrations can be done without any effort by local owners.
For third party deps, you can't usually do that, but there's a workaround[2].
I also think that this post and the one here[1] are amusingly both on the front page at the same time. This one decrying the exact sort of mandate based approaches that "make SRE scale".
Also worth mentioning that
> The person making the new API is expected to change client code to match, but is not responsible for ensuring the change does not break the client.
Is just explicitly untrue. There are policies that are clear about this (this generally comes down to "we cannot break your tests, but if it isn't tested, we can't know we broke it")
[1]: https://news.ycombinator.com/item?id=28825352
[2]: https://opensource.google/docs/thirdparty/oneversion/#tempor...
As the ex-maintainer of numpy and scipy third_party at google, I really appreciated the one-version-policy. It was the only way to sanely handle upgrades of thousands of different codes (and google has thousands of different codes that depend on numpy). It certainly helped us managed the complexity associated with mixing various versions of numerical libraries.
I think we're talking past each other a bit here. Piper's implementation could have been git-compatible originally yes (I think this is what you're saying?). The thing this article complains about is the turndown of the unofficial, unstaffed (20% at best) git-wrapper client around piper. That was a hacky not-fully-git-like but still well loved (I used it for a while!) tool. Doing it right would have been a huge undertaking and, as I understand, would have required building a totally new thing, while mercurial was hackable and extensible and could be used as a base.
- monorepos are a good way to organise software, especially if there are many internal libraries or dependencies
- It is bad if your dependencies are always churning and breaking and disappearing underneath you
It sounds like the actual problem might be that there are too many dependencies or that they change more frequently. I think the OP was suggesting that the latter of these is the case.
(I should also mention that there are recognized issues with the cost of churn at Google, but this isn't really one of them)
Things have probably changed since I left, but I bet the author's experience is pretty rare these days too. Google's system works quite well for their scale.
Shouldn't the person/team who makes a change take ownership of fixing any breakage it causes, rather than passing the buck to their customers?
It sounds like the real issue isn't anything as technical as monorepos-vs-polyrepos or software versioning policies, but rather a management and culture problem?
That doesn’t feel scalable to me. I’d imagine some dependencies at Google run to hundreds of clients.
To take an OSS view, some packages run to hundreds of thousands. Imagine if the maintainers of react had to fix all their clients.
It’s still not scalable in the sense that the downstream can make any change they want. They are restricted to changes they can either manually or automatically apply to client code in a reasonable amount of time. But that’s the trade off for being a heavily depended on library (which a lot of the time is an explicit decision).
Is it really though? Sure, it's technically possible, but it feels like a massive burden to place one on dependency.
Not only do you have to maintain your own lib, but you have to understand how every other project uses it.
In my experience the alternative tends to lead to divergence of codebases and products. We've all seen companies that have some great products counterbalanced with a long tail of products that behave completely differently, or are orphaned, or take decades to move to new infrastructure or integrate with other products.
We've been GCP customers for five years and I can't recall offhand an instance of a forced migration due to a breaking change. They give a lot of lead time and plenty of notices, and then continue to support the old thing for a good period of time. It makes sense to me that they would treat outward facing APIs differently given that they span many organizations by definition.
Sure, but how many times have you deployed an update that you were sure was security? Most libs I work with are 90+% features or ease-of-use releases, the security is hand wavy to do with dependencies, or a very rare "oops".
In most cases, the library authors are responsible for driving the migration. This is different from third-party libraries adding causing pain to everyone else because they decided to change an interface without sending PRs to migrate all their users.
I’m sure there are multiple layers of redundancy, but it seems like you could get into a Facebook style situation where you need to go open a (virtual) cage somewhere to flip a switch.
When I worked at Google I worked in hardware platforms and I spent a lot of time working with hwops people, who have to go physically out to some server and press buttons (whilst chatting the results to me). If corp gmail, chat, and everything else was down, I might have some trouble getting in touch with those folks and verifying the right fixes (which is why the FB situation is so crazy).
Yes, but for I guess the author wanted to keep things simple for outsiders.
Piper is supposedly not much different than Perforce. Forgetting implementation details is just a centralized version control system that holds a monorepo and everyone* just test and build from trunk/master/head.
Especially around who the burden falls on when a backwards compatible change needs to be made.
Agree with the point in the post that their mentality leaks in the way they support products.
The question is if they are being overworked for this or if theyre mad that their cushy well paid perk filled 9-5 job is inconveniencing them twice a month to get out of their seat. Like what?
The only times I’ve seen that work is when absolutely no one actually runs the code. If you want to write code at places where it touches billions of people (for better or worse), then this seems like a better way to do things than what other companies seem to do.
Just because OP is payed well doesn't mean they can't point out frustrations borne out of what they view as inefficient process.
Minimally they also use monorepos, so there’s that.
And it isn't like those with dependencies has no recourse, if they think one of the dependencies breaks too often they just remove the dependency and live without it. The library maintainers of course wants to avoid that so they have strong incentives to not break API's. And they can't break API's without first notifying or fixing their users either since then their code will just get rolled back, at Google it is typically much more work to break an API than for the downstream dependencies to fix those breakages.
Not GCP, but Google runs some of its open source projects like this. I've been burnt by breaking Guava and Java Protobuf changes multiple times. Especially with Guava, I strongly recommend that library developers avoid it.
If you have something you might characterize as one large suite of applications, they do bring benefits.
If you have many independent products, as is common in the B2B-space, especially with several versions in production at different customers that have different customization demands, it can be an absolute nightmare. Team A and D urgently required feature X as some regulation went through in their market and their customers had gambled it wouldn't, now Team B, E and G can't produce anything meaningful for six weeks because they need to build a downstream workarounds to circumvent feature X because their customers aren't allowed to have it because it violates GDPR.
In those discussions I often bring up that it is important to consider the engineering organization as a whole and think of it as a system in it's own right.
Others have already highlighted the monorepo benefits associated with managing dependencies. When it comes to (forced) migrations - we are all familiar with the accumulation of technical debt at organizations over time. I hypothesise that absence of (forced) migrations plays a role in causing it.
In my perspective the majority of monorepo shortcomings are related to tooling. On the build side I have found Bazel to be a fantastic piece of software. The source control question is another story. This was actually one of the original reasons for me to co-found a devtools company - Sturdy (YC W21)[0] - taking a step back and thinking over the developer experience.
Eg, /v1/dosomething, I need to make a breaking change but I can still provide the old service too so I leave v1 alone and instead add /v2/dosomething
I think another poster [0] hit the nail on the head when they identified the problem as likely poor dependency behavior.