For a huge, non-open codebase there are some pretty large downsides to a fully distributed VCS in exchange for relatively few benefits.
For a huge, non-open codebase there are some pretty large downsides to a fully distributed VCS in exchange for relatively few benefits.
It's important to stress that Google uses Perforce and not git (at least for that monorepo, they use git/gerrit for Android).
A monorepo this size would simply not scale on git, at least not without huge amounts of hacks (and to be fair, Google built an entire infrastructure on top of Perforce to make their monorepo work).
And there are other non perforce like Piper interfaces.
You are exactly right that git doesn't scale though, go see the posts on git that Facebook's engineers made while trying, only to be met with replies to the extent of "you're holding it wrong, go away, no massive monorepo here", at which point they made it work with mercurial instead. Good read though, lot of good technical details. Can't find the link at the moment though :(, but it was from somewhere around 2012-13 ish.
Edit: here, looks like the original thread is deleted but here's the hn pointer: https://news.ycombinator.com/item?id=3548824
If all those things continue I think the only reason to use git over hg would be github. How long until they decide to support Mercurial too and people abandon git?
Yes. End of story. People will abandon things that don't support them for things that do and those that want to continue using something that fits their application will do so. Nothing to see here; we get it, you don't like git -- don't use it if it doesn't fit your needs. However, don't expect those who do like it to go out of their way in a way they don't want to please you. Just because there is a community developed around something and that something is open source does not mean they are required to accept whatever patches come their way -- often the best projects know what to keep out as much as what to let in. In this case, the git community has decided it doesn't want to do those things; more power to them.
I think you nailed the problem with Git here: it was created by one guy to support his pet project and as long as it works well for him all the other feature requests are low priority.
https://opensource.googleblog.com/2017/03/dispatches-from-la...
Edit: scratch that, that works but has no threading. Take two.
These work:
http://git.661346.n2.nabble.com/Git-performance-results-on-a...
https://blogs.msdn.microsoft.com/bharry/2017/05/24/the-large...
"The Google codebase includes approximately one billion files and has a history of approximately 35 million commits spanning Google's entire 18-year existence. The repository contains 86TBa of data, including approximately two billion lines of code in nine million unique source files."
Source: https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...
It prompted me to do a quick afternoon experiment with how git would handle a billion lines of code:
What are the other ones and the main differences, really curious
Was pretty much used exclusively back when I was in gamedev, not sure if that's still the case.
Now, imagine you're a huge corporation. Your code consists of millions of files that have been edited millions of times. It's never going to be released to the public. It's never going to be forked, much less by a stranger. You're going to have only one main branch and main build ever, except for maintenance branches. The complete history of everything that has ever happened on that repo is would take up many gigabytes, and developers are probably only ever going to need to look at and/or build locally 0.01% of that code themselves.
If you were going to design a version control system from scratch for the latter scenario and you had never heard of git or any other existing VCS, how would you design it? Would you come up with something like git? Probably not. People would just have local copies of the minimum of what they needed to get their work done, anything else would call some server on the VPN they were always on. And you would probably come up with some whole specialized server architecture with databases and such that wasn't that similar to a corresponding client architecture that it would also need.
A lot of companies don't use git.
I think one aspect of Git that is really important is forking, and having your own local commits. Merging commits and patches in svn were awful. You wouldn't ever allow someone random to join your svn repo, but if they can reasonably provide a patch, you could take it. Git makes that massively easier.
brrrr
For me the main feature was distributed nature. SVN is OK on a gigabit corporate LAN with dedicated people to manage & maintain the servers + network. Anything less than that, and it becomes slow and unreliable.