Scaling Mercurial at Facebook
code.facebook.com
code.facebook.com
So instead of cleaning up the internal dependencies, they decided to rewrite Mercurial. That is the kind of thing Facebook likes to do: for example, when PHP got too slow, they wrote a PHP compiler....
The scale of these codebases is way outside what most people have experienced and it might behoove people to realize that a lot of conventional wisdom might not apply. E.g. last time I checked, Google's codebase was >2Billion lines of code.
Google has 30,000 engineers working on a "monolithic" codebase and remain relatively productive despite (or perhaps because of it). There are whole sets of different problems at this scale.
Thus far I've preferred the mono-repo mostly for dependency management reasons. Whether you have a lot of internal dependencies or external dependencies, you get similar benefits: - All the changes to internal dependencies are in your revision history, across the company. Shared internal library upgrades are picked up quickly and propagate throughout applications with low delay.
- It's easier to have shared versions of external dependencies "automatically" instead of establishing policies that need a human to enforce. This makes it easier to roll out new versions and bug fixes.
- It's easier to do system-wide improvements in code quality. Replacing common bad code with better implementations is something I've seen across code bases at both Google and Twitter and they have been beneficial.
I think nobody should be trying to implement a system like this if you've got fewer than ~600 SW engineers, though. Small groups don't have as much drift or system-wide refactorings that give you benefit.
Package management is simple, we push common packages into sonatype nexus with semantic versioning.
Personally I prefer strong versioning and allowing teams to upgrade common libraries at their own pace, but we've had a few times where we needed to quickly upgrade every repository (i.e., a security patch).
You have a few options:
1. You can unpublish the old dependency, breaking all builds until they upgrade. (Frowned upon)
2. You can write a script that identifies all repository owners that depend on you and send out a upgrade by X date email.
3. You can script pull request creation and submit hundreds of pulls to upgrade the dependency.
So far, things have worked out well, and I never want to go back to the monorepo.
As a "one-man-team" who uses at least 15 different repos, it's hard for me to imagine how a massive company would manage things within a single repo and no package management.
Also, aren't any of those monorepo companies concerned that a single rogue employee or stolen laptop could leak their entire source code?
I can only guess I'm misunderstanding what is meant by monorepo and how they're used...
Personally, I also favor a single repo. You manage it the way you manage separate packages: with organization and some discipline. The magic is that command line tools like grep, sed & awk--and static analysis tools for refactoring--work really well. You can change a method signature and it just works. I've been part of monolith-breaking before (most recently at Trulia) and it definitely adds friction to the development process to work across such a rich graph of package dependencies.
* Compilation is done centrally, you code against a mock or only the interface and the submit the code for test and final build.
* Or only libraries are supplied, possibly ofuscated.
There are dependency management tools that help enforce public/private code on a wider scale and that help the build tools make sense of it all. There are also ownership tools that say what people and teams are qualified to review code in certain directories. The config files for all these tools are checked into the repository.
There's no versioning though. If you want to change an internal API you just update it and all the callers at once: patches in source control are already atomic. For truly massive changes (more code than most companies have) this gets too unwieldy and there are special tools and strategies people use.
This means you can't have a project depending on an out of date (internal) library. Without that requirement, you don't have the situation where different libraries need to be synced to different versions. And without needing to sync different things differently, you can get away with just one repository.
I worked at Amazon previously, which has world-class tools for dealing with versioned libraries in bulk, and Google's approach is vastly better. You spend less time worrying about breaking other people's dependencies, and you don't have someone spending a day fixing libraries every couple of weeks.
FBShipIt has been primarily designed for branches with linear histories; in particular, it does not understand merge commits.
That's sad.Even with stolen FB code without dedicated infrastructure you still cannot do anything.
"Why didn't we think of this!!!?"
Maybe it's because these unknowns can't fathom working with such tightly wound systems that after all the code reviews are done your changes are irrelevant and need to be updated again and go through another review. In before straw man, but other than pointing out day to day issues, maybe you should try to understand the argument against trying to fit everything all in one place. Do you really want to be the maintainer of all those third party libraries you imported? Do you not allow third party and adopt "not invented here"?
Try again
Our code base has grown organically and its internal dependencies are very complex. We could have spent a lot of time making it more modular in a way that would be friendly to a source control tool, but there are a number of benefits to using a single repository. Even at our current scale, we often make large changes throughout our code base, and having a single repository is useful for continuous modernization. Splitting it up would make large, atomic refactorings more difficult. On top of that, the idea that the scaling constraints of our source control system should dictate our code structure just doesn't sit well with us.
I read it, but none of those things make me think better of their code.
Which works great until the day you wake up to realize your code base has become an incomprehensible mess and your engineering systems are outdated and your competitors are shipping better software at a faster pace.
I'm not really disagreeing, I'm just saying a balance must be struck. Really hate companies that view anything but feature work as a loss.
http://danluu.com/monorepo/ is a good explanation of why that's a pretty foolish take.
1. Holy god why would you let your code grow to such a massive, interdependent scale? Just release everything separately and versioned so that breaking changes don't affect everyone all at once. The idea of git being a bottleneck is absurd and you are using it wrong.
2. This is a very reasonable, practical approach to sharing code across a company. It reduces siloing and ensures that major refactors can happen in one pass without a ton of coordination. Better to fix the version control system than waste endless resources refactoring millions of lines of code.
Both reaction is valid. Having worked in both styles of codebase, I recognize that there are trade-offs in either case. The optimal solution depends on the project and the team.
Sometimes the path of least resistance--that is to say, the path to getting things shipped and, in turn, making money--is to let the codebase grow organically and worry about cleaning up any messy interdependencies later, once you have a better idea of what code you even need to keep around. In this scenario, it's important to recognize that developer efficiency is going to be an uphill battle in the long run, but if you are proactive about maintenance and tooling improvements then this approach can still be relatively painless.
Other times, especially when you're working on a tried-and-tested product with a clear API and a dedicated team, it can be productive to split it out and let the team manage their own versioning and releases. This becomes especially useful if the product is open source. (For instance, I wonder how Facebook manages its open source releases relative to its shared Mercurial codebase.) In this scenario, developer efficiency is usually less of a problem, as proper use of versioning can ensure faster, more agile updates to each product. But the downside is that your company as a whole can end up in a kind of versioning hell, where every project depends on a different version of every other project, and keeping everything up to date can require a huge amount of coordination.
So, in the end, pick your poison. My reaction, years ago, was more along the lines of #1, but I used to be much more of an idealist earlier in my career.
There is no atomic deployment of a large distributed system - so even if you can check in related changes in different areas of a codebase, how do you release them?
It sanitizes and syncs them to GitHub
https://code.facebook.com/posts/1715560542066337/automatical...
http://thread.gmane.org/gmane.comp.version-control.git/29510...
Their first attempt was in May 2014: http://www.spinics.net/lists/git/msg230487.html
This is years old. Am I missing why this is newly relevant?
About securing different parts of the repo, the mercurial server actually doesn't have any user authentication! You are meant to do that yourself with SSH or a web server, where you should be able to have more restrict access to some special folders.
About using a single repo, it does make sense to have all code that interact with each other at the same place. Imagine changing a variable name in some API and at the same time update all usage of that name in the whole codebase. And imagine the bureaucracy and people management for just making a variable change if there where separate repos witch you might not even have access to.
> For a repository as large as ours, a major bottleneck is simply finding out what files have changed. Git examines every file and naturally becomes slower and slower as the number of files increases, while Perforce "cheats" by forcing users to tell it which files they are going to edit. The Git approach doesn't scale, and the Perforce approach isn't friendly.
I don't know if Perforce's "unfriendliness" is enough reason not to choose it nor do I know if that was the only reason they rejected it. However, speaking for myself, I vastly prefer Subversion, Git, and Mercurial over Perforce precisely because Perforce requires you to ask for permission before you can edit a file.
There's also this ever-present feeling of dread whenever I invoke any p4 command -- if I absentmindedly submit something with unshelving the change that I had shelved after making a previous change but reverted when trying out a new change, I can wipe my old data and lose work. Or some issue comes up in a deployed environment, so I have to grab a new workspace (or dance through shelving/unshelving all my open CLs), have a 10 minute coffee break while it syncs, sync up to the release line and work on that one. Also my IDE doesn't like being switched rapidly between repositories all the time. Also it's difficult to just _commit_ some code after making a change and wanting to save progress. Every operation is just so stress inducing; how do you handle it :/