Git from the Bottom Up (2022)
jwiegley.github.io
jwiegley.github.io
Usually there are at least a few choices with different compromises. Browsers, OSes, languages, editors, all have multiple actually popular choices.
Does git win because of the decentralized, everyone-has-a-local-repo aspect? Or the staging area, which I understand other VCSes don't have? Or just speed or reliability? Or just being in the right place at the right time?
I suspect most GitHub users didn't realise that it supported SVN until earlier this year though.
https://github.blog/2023-01-20-sunsetting-subversion-support...
You'll still find different solutions in companies that really need it though, like Google.
Git is very much "good enough" for many people. The number of foot guns and lack of a well trodden path for many centralized corporate devs is quite high, but many people are content with it.
Personally, I think if the only advantage is git formalizes a merkle tree and sha based text comparisons with a poor cli ui, it's not very good. But many people think this is sufficient.
There's a powerful network effect with choice of version control tool, because others need to collaborate with you (even if just to pull or browse your repository) so the tool that starts to dominate will likely monopolize. This isn't the same with browsers.
How it got there is a different story. I’m old enough to have been around the Danish CS scene back when computer clusters were just an experimental thing at universities and when CVS arrived around the millennium. While it was mind blowing for its time, it was also terrible to work with. Similarly subversion wasn’t very nice (for many people) and it’s likely because the centralised model just isn’t very work friendly for a lot of organisations. Then came Git, which wasn’t the first decentralised version control, I believe Microsoft came first with their hellish team foundation which proved that distributed systems aren’t good by default. Anyway, at the time Git was such a breath of fresh air in how great it was to use. It was both easy and productive and it held such an increase in quality of life that it was a no brainier to transition to git for most organisations. At least the ones which didn’t have expensive team foundation or subversion contracts.
Which birthed the circle or dominance that might never get broken. I doubt any of the old players will in any case. If they were capable of building great version control they would’ve done so from the start, and since Gits only major problem is its multitude of options it’s hard to imagine that Subcersion would ever cut 90% of their features. Team Foundation mostly trundles along because some organisations still haven’t given up on their licenses (which is the reason a lot of Microsoft things exists), but with Microsoft owning GitHub it’s not likely that they want to do anything other than to keep it good enough to keep that revenue.
How fast is Git? Do you remember or have you tried other VCS-es older than Git? Try getting the first working copy, try looking at changes for a file through all its history, try making a new branch, try tagging. Git is fast.
Because it is fast it was selected by many and the rest is network effects. Now you can do something just as fast or maybe even faster, but everyone else is on git. You may do something with better UX/UI, but this rarely is enough. You have to do something much better and you will not be orders of magnitude faster than git, but that is now a basic requirement.
Your questions are actually covering 2 different categories:
(1) Why did (past tense) git win?
(2) Why does (current tense) git have a monopoly of usage now?
For (1), the various theories for git winning mindshare to create an insurmountable lead include technical and social differences:
- Git being faster than Mercurial. It had "cheaper" branches than Mercurial. And some blamed it on Mercurial being built with python instead of Git's C.
- index/staging area
- the Github free tier being more generous than Bitbucket(Mercurial)
- intangibles such as being created by Linus and high-profile usage by the Linux kernel contributors
Google Trends shows that Git+Github almost immediately outpaced Mercurial+BitBucket from 2008:
https://trends.google.com/trends/explore?date=all&geo=US&q=g...
https://trends.google.com/trends/explore?date=all&geo=US&q=g...
For (2) today, you mostly have inertia. Most developers are not going to bother to research and re-evaluate the VCS landscape and make a deliberate choice on any technical merits. Git is already too massively popular so just go with what everybody else is already using and just move on to something else. Even if another DVCS has some technical superior aspects (e.g. Fossil?), it doesn't matter because git is already too entrenched in the ecosystem.
An example of the network effects of (2) is Python's creator Guido van Rossum suggesting the move from Mercurial to git/Github:
2009 choose Mercurial/hg.python.org: https://mail.python.org/pipermail/python-dev/2009-March/0879...
2014 migrate to git/Github: https://mail.python.org/pipermail/python-dev/2014-November/1...
His 2014 suggestion to switch to git isn't based on technical features. It's about using what's the most popular to reduce friction.
I was sorry to see it go away, but mercurial took over for a while, and then later I switched to git because everybody else had done so.
Likewise for the 2009 mail.
But being a monopoly doesn’t mean others can’t use other clients. I was using a Subversion repository in university but really I was just using Git and git-svn (probably). One can do the same today.
So, besides the obvious drawback of having to know two VCS, why does it seem that not many people are doing this? Because I don’t hear that much about it.
> Does git win because of the decentralized, everyone-has-a-local-repo aspect?
of course not, hg and bzr (and arch and monotone) had that, both before git existed.
Indeed it feels like most competitors are competing by offering less. This is our way! This is easier! But you're not going to win the alpha geeks over with that path. Better needs to be bigger, needs to enable more, and we're just not seeing those offerings materialize; it's unclear how you would be better or bigger.
Having the core abstractions that work & are flexible enables bountiful innovation & improvements. Git not only started with sufficiently malleable core ideas, but it's had endless features, optimization, and tuning baked in over the years. It's almost too much to comprehend. Shout out to Taylor Blau writing on the Github Blog: the man has been doing an amazing job covering what's happening in with the Highlights From Git series, since September 2018. It's really given me a sense of how much work keeps being poured into git, has given me a view of how incredibly expansive the git ecosystem is. https://github.blog/2018-09-10-highlights-from-git-2-19/ https://github.blog/tag/git/
The way git stores commits---as blobs and hashes etc---is a really good abstraction. That means you just don't need to know this. You can think of a commit as being a complete copy of your working directory that is identifiable by its hash. That's it.
You can then think of branches and tags as labels for those commits, mutable and immutable, respectively. This gets you pretty far in effective use of git.
But the details are cool. Just don't go teaching any of this in a beginner developer course on git.
I would think that mentioning the big-picture concepts at the outset, and revealing the details as the need arises is an effective pedagogical model.
On the other hand, the minestrone of git terminology which the user is forced to learn and get comfortable with a priori (and moreover which has been used ambiguously in blog posts and whatnot) [1] make the learning curve much steeper than it should be.
Everyone learns differently, that's for sure. I remember my experience learning git from scratch was particularly scrappy, until I sat down for a few hours and read the "Pro Git" book.
[1] staging and index
Would be very helpful for Docker or Kubernetes, for example. The only one that comes to mind is From Nand to Tetris. https://www.nand2tetris.org/
Not listed there but I also found Ben Eater's 8 bit computer a fantastic learning tool.
Looking forward to sharing this with colleagues. On one hand I am excited because I think this bottom-up approach is the correct way of understanding git, on the other hand I am afraid that they will think it's difficult to understand and just not bother. (Again I can't fathom how programmers, who you know, learned programming, which is quite a steep learning curve in itself for first time, shy away from learning git properly, a tool that's just as important to their daily work as whatever programming language they are using.)