Why Google Stores Billions of Lines of Code in a Single Repository (2016)
cacm.acm.org
cacm.acm.org
We don't actually have a single repository, although Google3/Piper really is huge and dominant. Android, Chrome, Fuchsia, and myriads of other projects don't exist in the monorepo (well, they are mirrored there) because the tooling and libraries there are really geared towards server deployment, not devices.
karma farming?
and thought this might yield an interesting discussion
Some stories are great and worth seeing every few years.
Your anger gave away the fact that you do the same thing:
https://news.ycombinator.com/submitted?id=Tomte
The one downside is that it creates a lot of noise on HN. It forces new stories off the first page quicker.
Also, stories about Jerry Lee Lewis’ marriage to a 13 year old isn’t HN, for example.
About Jerry Lee Lewis: I disagree. The sheer fact that moral compasses were so wildly divergent between America and England, of all pairs of countries, so recently, is quite on-topic, in my opinion.
Also, I think it's great that you found only one story you disagree with in the thirty stories my submission page shows. Clearly you approve of my selection.
> I’m simply pointing out why we get lots of repeat submissions.
That's explicitly okay, as the moderators point out all the time.
I didn’t say it wasn’t ok. I’m sure I do it occasionally.
I’m waiting for the first person to write a karma farming AI. Take great stories, determine a decay time, then repost.
“HN Karma Farming with Deep Learning (in Rust)”
This probably doesn't even need any AI, just some scrapping and scripting to determine when it's worth resubmitting.
Somebody should implement this, seems like fun.
Of course, if it's done with deep learning... then it would be something else. Maybe even determine correlation with other themes in the other news of the day/front page, or month, etc.
Sure, write that karma farming bot. Seems like a good service to those of us who want to see good content we may have missed, generate new discussions on classic content.
Complaining about “karma farming” has always seemed like a weird jealousy thing to me. Where you presumably care so much about karma that you’re annoyed over your arbitrary concept of karma-fairness. Who cares? If someone can get karma by reposting classic content for HNers who missed it the other times, more power to them. They reposted the content to an audience that wanted to see it. You can’t always take it personally or lash out at others just because you’re not part of that audience.
I did point out a downside, but have at it.
Your total "internet virtue points" here are displayed in the top right corner (whereas on Reddit you have to go looking). The virtue points any given comments gains are not visible to other people (unlike on reddit) but losses are. That doesn't seem like a big net difference between the two platforms.
HN's karma system is just as good as Reddit's at faking a consensus and marginalizing whoever has even a slightly less popular viewpoint on a subject.
(I don't know many well-designed karma systems, but, that's at least the goal. The feedback cycle exists for a reason.)
It's true that repostiness per se is annoying, but it's more annoying for good submissions to miss out on attention because /newest is something of a lottery. Allowing a small number of reposts is a way to mitigate that randomness, until we can maybe work out a better solution.
And you're right this topic is on a circular timer.
(more people are aware of the challenges and stumble upon monorepo as a tool to deal with this among other things)
related:
* unison
* clojure (immutability, spec, …)
* blockchain
* microservice architectureWhen I was working on an open source project at Google (mod_pagespeed) we initially developed it in the main repo with exports to GitHub. Then we moved our source of truth to GitHub to make it easier to work with external contributors.
(Speaking only for myself, not for Google)
Prior to Fiber, yes, I worked in Google3. I've been @ Google for 8 years.
Also, one main advantage of a monorepo is you can make changes across many packages at once. The canonical textbook example is changing the API of my service: I can find all the instances in the monorepo of where that API gets used and change both the API and update all the calling instances in an atomic operation. You can't (or at least, shouldn't) do that to mirrored projects because the mirror in the monorepo is downstream from the source of truth for the project.
On the other hand, if you aren’t doing simultaneous/atomic deploys, then having a monorepo may actually be an obstruction.
In fairness, that's the canonical test for "should this be 1 repo"
Of course, the scale is much lower, since it’s a really a single OS.
This article doesn’t claim that. Its title doesn’t imply it, and its second sentence denies it:
”today the vast majority of Google's software assets continues to be stored in a single, shared repository”
When Microsoft bought WebTV they got this license, which enabled them to cheaply roll out P4/SourceDepot to massive teams for whom it was a big improvement over what they had been using. With access to the source, they could modify it to their needs and scale away.
Google tried to buy a source license early on, but Perforce had learned their lesson from what happened with WebTV and wouldn't sell them one. That's how Google ended up doing elaborate projects for years to prop up and eventually replace P4.
Source : I worked at WebTV, Microsoft and Google.
Not that I'm against the policy or how it works. It just isn't "avoided" precisely.
Once you're in the mono repo, you have to play by its rules (well...should anyway). Like in the case of npm packages I mentioned above.
Basically the more users a single dependency has, the harder it gets to update.
The flip side is that the monorepo (more easily than otherwise) allows the person doing the change to bear the burden of updating the call-sites (& users).
Thanks for pointing out the difficulty this introduces.
Programmers know the value of everything and the cost of nothing.
There is a reason why most globally linked runtimes (e.g. C++ or Java) do not encourage such mechanisms.
I generally agree but the problem is even worse :)
https://yarnpkg.com/lang/en/docs/selective-version-resolutio...
It isn't clear to me that it is worth the investment, even for Google -- though it clearly world's "well enough"
And besides, it's not like this gets any easier if you have multiple repos. Eventually you have to roll forward your dependencies, and doing the work that goes along with that.
Multi-repo makes integrations harder, because you find out about breakages later -- which means that it might be harder to identify the cause and more likely that someone depends on the behavior that broke you by the time you notice it.
But, multi-repo makes local development faster and cheaper -- you're insulated from the churn of everybody else's check-ins. You don't have to constantly refactor everything every time some dependency makes a minor tweak.
You get more done and have a smoother development lifecycle -- but you're going to keep falling behind your dependencies unless you invest in keeping up -- and that part of the process is more unfun the less often you do it.
I currently live in a monorepo. I don't love it. But I've also not loved the multi-repos I've lived in, so... shrug
A lot of work is going into making git scale better which is nice but there will always be small open source repos. You could mirror them but it seems like it would be nice to not have to do that.
Is there any research going into figuring out how to keep many repos in sync with a single commit? Maybe you could use submodules across your company...have a single master repo that essentially just keeps track of working version sets. That might work but it sounds torturous. Is anything on the horizon to solve this problem with git?
Internal research at google suggested that ‘git5’ users were objectively less productive than average, but there’s some friction around ‘git5’ so that doesn’t necessarily implicate git itself.
The general low quality of a lot of the Perforce tools has me scratching my head every day. How hard is it to make a log window that scrolls to the bottom? The p4proxy doesn't have a max size, you need to manually delete files on the machine to free space!?
As for a single command to see already merged/copied CLs between two streams, please enlighten me. I want the unsquashed CLs. It would be really useful for generating patch notes but because p4 does version tracking at the file level and you can merge some files from a commit and not others etc, it's not trivial at all.
In my personal opinion, using a git/hg like interface makes it lot easier to work on a complicated CL because you can maintain internal local branches and you can easily revert your incremental changes. That’s not at all possible in perforce. I just can’t see how git/hg interface can make anyone less productive.
They're not a good metric to judge a developer on. But when you can do a large controlled study (or even before/after with the same developer), without the developer knowing they're being watched, it's a good metric.
The one thing p4 does have is file level versioning. It's trivial to go back and pick single versions of files. Artists need that feature and most git clients hide it away or don't support it at all.
In the longer term, I'm hopeful that VFS for Git[3] (from Microsoft) and/or Mononoke[4] (from Facebook) will reach a level of maturity that makes it easy to host large monorepos using open-source distributed version control systems.
In the short-term most smaller organizations aren't going to hit the technical scaling limits of Git anyway. Most organizations could put all their projects in a single Git repo and not have any significant performance issues. Organizational issues like commit ACLs and code review enforcement are a separate issue, and often require additional tooling.
[1] https://gerrit.googlesource.com/git-repo/
See Facebook's discussion with git devs here: https://news.ycombinator.com/item?id=3548824
The long version is that many are trying. MS is making some headway last I checked.
Choosing a monorepo? Be prepared for:
- Investing in the build system
- Investing in CI/CD
- More painful upstream dependency management
- A serious investment in architecting modules and thoughtful dependencies
- Issues because of a bad deployment sequence
- Slow version control
Polyrepo your thing?
- Managing your own artifact repo (pypi, artifactory, etc)
- Painful cross-project changes
- Repeated effort around builds
Can't we call Github a monorepository, as it is a single area one can use to access code?
Implementation details likely differ from whatever Google and her engineers are cooking up, but at the end of the day I point to an archive and pull artifacts. Certainly not every Googler clones the entire monorepo...
In practice, if you're changing 100+ individual callsites, then you would probably make a backwards-compatible change, then use an automated system to send out and manage a bunch of different commits to clean up call sites, then clean up your old function signature once all the commits are submitted. If your code isn't a really widely used library then you probably have fewer than 20 callsites, though, so it's nice to be able to do it in one commit.
I meant to say "atomically commit", but my brain mixed it with mono-repo, to come up with monotonically, which is a math word w/ no relevance to branching (afaik!)
For branching, I just meant there's no easy way to branch everything at once. If you want to have "weirdExperiment" branch that affects code in 4 repos, you have to go branch 4 different times. For committing, same deal. You can't easily tag a single snapshot of the state of all projects. Etc.
Additionally, you generally get more flexibility of how to slice and dice your view of the repo, rather than being locked into code boundaries that were set once and are difficult to change.
At that scale I need tools to manage the monorepo. What about making tools to manage many repos together?
(I have worked for Ericsson previously for 7 years but that was before Eiffel)
[1] https://eiffel-community.github.io/
[2] https://www.amazon.com/Continuous-Practices-Strategic-Accele...
Why do they need to store so much data in a single DB, and could they use a different approach? Rewriting tools that already work well is a big of an extreme solution, but Google also has the resources to do something like that.
Well, we have examples of the opposite (Amazon uses a multi-repo approach) and my understanding is that it works fine, until it doesn't, when its awful. This often happens when there are security vulnerabilities, and Amazon is forced to, under time pressure, walk their entire dependency graph and pin all dependencies to newer versions, fixing forward as they go.
At the point where you have a lot of code, you only really have two options.
I see the choices as: a) rewrite one core tool infrastructure component or b) rewrite everything - currently and soon to be - built on that infrastructure.
It's also a question of level of abstraction. Why does my application care about database sharing, replication, leaders, followers, etc. It should only care about connecting & querying. The complications of running a system at scale should be hidden as well as possible from the code you've written.
From a monorepo perspective: why do I care where my code is stored? As long as I can control visibility, hide build complexity, and expose an API to other pieces of code I have all the functionality I need. Why should the nitty gritty of a package manager, VCSs, or other multi-repo concepts influence how I write my code?
Edit: or maybe Brobdingragian, as Swift said the publishers actually misspelled it.
This was a datapoint on a theory I have that projects where all the code is being actively stewarded (instead of abandoned code nobody understands) have less than about 15k (human) lines per developer. If Google’s ratio of generated code is near that, then with 40k developers they’re not far from that ratio.