A look back: Bram Cohen vs Linus Torvalds (2007)
wincent.com
wincent.com
Wait, what? Git was the response from Torvalds to BitKeeper being proprietary and the author being upset with some kernel developers. It's basically taking what was then state of the art for DCVS and reimplementing it, not coming up with some new paradigm. It's a good program, and to my knowledge it was engineered well, but let's not rewrite history so Torvalds is the father of DVCS.
BitKeeper existed, a few projects used darcs, GNU arch and monotone existed but I don't think I ever encountered a project using either (the Canonical arch fork aside), and that's mostly a wrap for free or Free options.
He may not be the father of a novel concept which he could write a paper about, but Git and Mercurial existing certainly coincided with the enormous uptick in DVCS usage and visibility.
I suppose the difference in opinion about the term "father" just comes down to being about the theory behind the problem space or wildly successful implementation, and IMO "taking known solutions and putting them together in a more successful way" is novel enough to claim credit.
I very much doubt it would have succeed if it wasn't for those two factors.
As such I would put it on the same basket as C, PHP and JavaScript, regarding adoption vs quality.
As for Github, it is just the new Sourceforge, helped by the way they mismanaged it.
Agreed; when GitHub is bought by the wrong company (when; not if), we'll be in a world of shit.
Github aside, in what ways does git fall shorts of your expectation as a user? Maybe there are alternatives for you to use, but we need some clarity to decide whether they exist or not.
I used to criticize Mercurial for their half-assed branching, lack of history editing and the refusal to admit that git got this right. But Mercurial has improved (bookmarks, histedit, etc) and git has not, so my feelings have entirely reversed now.
The thing is, Git is pretty much the most horribly over complicated user tool in history. Learning how to use Git properly is more complicated than learning how to use Unix. Several times more complicated. I honestly would be more comfortable giving a modern team who had never seen a VCS before just plain old CVS, because even if it's horrible, at least it's simple and you can just work around the stupidity. Git will leave you trapped in a hellscape of reading manuals and howtos and bloating personal repositories, and not really do anything particularly great except feature branches.
DVCS is great when you need it, but when you don't need it, it's annoying as hell.
At the end of the day, a VCS is just a tool, one of many that I need to learn and use to do my job, and it's not worth the effort to learn both git and hg.
Somewhat related anecdote: years ago I decided to learn Dvorak, but eventually switched back, because by the time I'd become proficient, my ability to type in QWERTY's completely gone. From what I'd read online at the time this was unusual: most people who learned Dvorak could switch back (& forth) within a few seconds to a few minutes. It took me probably a week or two to be able to touch type in QWERTY again, and maybe a month to get back to my original speed. And by then it was as if I'd never learned Dvorak at all.
Anyway, the point of the story is: maybe I just have shit memory :-)
Using Git for everything is like riding a bicycle with four derailleurs to pick up milk from the corner store. Granted, this is what I do right now; just because some technology is complicated or annoying doesn't mean I don't use it. But I wouldn't recommend it to others.
Svn in particular was so awkward to port diffs between branches and carry diffs forward in time (git equivalent of rebase) that I built a system of shell scripts around patch files. Instead of creating a branch and committing changes (so painful to create or switch branches in svn), I saved the working state diff as a patch file and reset the working directory whenever I had to switch to a different task. I had a couple of shell scripts, one to save the current diff and reset, and another to apply a diff to the current working tree. And I had a third one to do a three-way merge to resolve conflicts when the tree had been updated since the patch file had been created.
Git is the first SCCS I used that made sense.
Without the right mental model, I expect git would be very difficult to use. I might not recommend it for a team I didn't think could grasp the concept of a DAG. Then again, such a team will basically be cargo culting every moderately complex technology they use, so what does it matter if their "committing_notes.txt" file contains git commands or something else?
and what about this line on the wikipedia page?
> This [community] version of BitKeeper also required that certain meta-information about changes be stored on computer servers operated by BitMover, an addition that made it impossible for community version users to run projects of which BitMover was unaware.
so not storing some changes locally, and being unable to access it if you don't have a paying version of the client seems to be really archaic compared to the simplicity of git. i could be misinterpreting what they mean by meta-information though.
Git is nothing like a reimplementation of BitKeeper - the closest inspiration and model is actaually Monotone, which had the same basic concepts but was slow and clunky. IIRC, Torvalds credited monotone/Hoare for using merkle trees (and other things)
Could you explain why, for a tool meant to be used by humans, the underlying model (which I agree is very elegant) counts more than the user interface?
It does not. But git exposes this all the time, so committed git users feel that's a good thing, because understanding it makes helps them to make sense of some of the worst parts of the UX.
The mistake is to believe there is a U<I|X> in the first place.
It does not, however, stop you from using one of the many graphical tools, that are fine for simple usage and will all eventually fall short.
You could use a GUI to make complex HTTP routes, through proxies and DNS records via drag&drop, and then one day you'll have some weird DNSSEC error that the maintainer does not care about, and you will have to explain that to whoever is losing money.
Acknowledging a problem is not as trivial as it was first thought is critical in our line of work in my opinion.
In reality this "reference implementation" is the sole interface 99% of the users use, so the lack of UX is an issue.
There are other tools that do it better and right, so insisting that there's any redeeming qualities to git here is being blind to the obvious. There's actual academic research describing git's problems here, for crying out loud.
And interestingly, all the academic research talks about is the high level UI. Not the underlying data model.
> In reality this "reference implementation" is the sole interface 99% of the users use
/That/ is the mystery to me. I don't see a reason why a better UX around the same underlying data model, talking to the same server (and thus interacting natively with e.g. github) hasn't picked up.
Heck, one could implement mercurial's UX on top of the git data model. In fact, that even has been prototyped. I can't find the repository anymore, but iirc it was somewhere on bitbucket.
Function matters more than anything else... and exactly what "function" means and encompasses can be discussed, but it's difficult to define function such as to exclude git's model from git's function.
The interface that humans should use would have been cogito, but after some time there was no interest anymore in this and it was discontinued.
It's just like the movies. You can make a bad movie from a good script, but a bad script never, ever made a good movie.
Tridge (another kernel dev) reversed BK protocol and repo structure, as trivial as they were. BK people got understandably upset about this, in part because they were providing BK to the Linux Kernel Project for free. So Tridge's little stunt basically killed any and all goodwill between BK and LKP and Linus needed a replacement. All alternatives had issues and everyone just kept going in circles, bitching and moaning. So in a true programmers way Linus sat down and wrote what he wanted. The end.
Bitkeeper's sole claim of ownership moral or legal was to the bitkeeper software which nobody violated. What Tridge did was come up with another way for the owners of the relevant data to access their own information something wholly reasonable and justified.
What you dismiss as a stunt proved in a stroke how much of a mistake the relationship with Bitkeeper was.
This is approximately like sony records telling you that you can't take your cd and put it in a Samsung stereo and you defending sony except no reasonable person supposes that creators of material objects aquire moral rights to control the use thereof.
This isn't true at all. Anyone could download the source code, tweak it, and submit patches. You can even use your own version control system to track your changes.
The project managers may have used bitkeeper, but to everyone else the only thing that bitkeeper provided was, at best, convenience.
It's also entirely irrelevant if people use different tools to do their job. No one was hindered by anyone else's personal choice of tools. Anyone in the world was free to download the source code, change it, and contribute patches if they seed fit. Bitkeeper did not hindered this, nor did the linux development process changed once bitkeeper was replaced.
BitKeeper was better than the alternatives. This was literally Linus' rationale for using it.
Nothing what you say in the second paragraph contradicts what I said, but it also doesn't address the point at all. Do you understand why a difference in tooling and ease of working reduces openness?
http://sourcepuller.cvs.sourceforge.net/viewvc/sourcepuller/...
Your personal opinion of the license terms is irrelevant. You respect other people's license, because you want them to follow yours. If developers can not grasp that, how do we expect users to?
Everyone's freedom is so vastly more important than a trivial few's wholly imaginary "right" to profit. Software is anymore the building block of civilization, culture, business. Valuing the right to restrict others over everyone's interests is a strange inversion of priorities.
Talking so high about freedom and rights, and then calling other people's rights trivial when they do not fit in your echo chamber of righteousness.
Copyright isn't a right in the sense that anyone uses that word.
He made a connection to the BitKeeper server and typed "help".
https://lwn.net/Articles/132938/
The BitKeeper developer was an ass, not need to be polite about it. He issued DMCA requests to websites that were analyzing his license, claiming that it was copyrighted.
HN discussion: https://news.ycombinator.com/item?id=8650483
Content-addressing (ie hashes) enables decentralization (hashes are the same whereever you are) and integrity (changing content will change the hash). (Perhaps inspired by Tridge, who did rsync which uses hashes to test for content changes?) It's incredibly fast because commits are almost as simple as copying, and because Linus knows how to C.
Linus left out some features, like renames (which bitkeeper has), so this simple idea would be enough.
Some people rave about the index/cache, though it's separable from the core idea.
https://en.wikipedia.org/wiki/Monotone_(software)#Monotone_a...
git made several additional breakthroughs in terms of working in terms of snapshots instead of diffsets, keeping much less data, and performing many computations late rather than early. In essence, these result in a system infinitely more flexible and expandable than prior version control systems. Adding new ways to change the source code in CVS required a whole new data format (SVN). In git, that's a minor change.
The full power of git hasn't been anywhere close to exploited yet. The data model is very general-purpose, fast, and robust. It can do much more than just source control.
So I don't think this reasoning holds any ground whatsoever.
looks like you should have had a merge conflict in that sentence /s
On topic however, missing one such conflict is easier, if you have to spot it after the damage is done, instead of having to think about it yourself beforehand, no?
I surveyed a fair number of DVCS systems in the run-up to git. I'm hardly an expert, but gracefully dealing with merge conflicts was indeed something system authors were investigating. Darcs went so far as to have a formal theory of patch management [1], which guaranteed never to have a conflict (unfortunately, running the proof engine could take unduly long and in some cases, may never complete).
Anyway, I think it's a tricky balance to strike. You're right that if the SCM resolves the conflict improperly and introduces a logic error, that's really tedious to track down. However, I've encountered far too many cases of bad rebases or merges resolved incorrectly by humans as well. Sometimes it's a small change the developer didn't pick up in line that was edited by both. Sometimes it was a lack of understanding of incoming changes. Sometimes it was just a battle lost to the merge tool (e.g., seeing "<<<<<<<" in source files). In all those situations, I long for a smarter merge tool.
[1] -- https://en.wikibooks.org/wiki/Understanding_Darcs/Patch_theo...
During the merge window, Linus takes in most of changes for the next kernel release. All of which have been approved by someone else. With each RC release, the lieutenants still approve most of the new changes. It's a matter of when a crucial fix is needed for Linus's mainline kernel, then he'll accept a commit directly.
Tree structure is overrated. The way git handles it works better in practice. Files are an implementation detail, rather than something fundamental about code structure.
Using it as a first pass at content merging was mostly a performance optimization, and also only works if you track individual files as objects.
It might be a useful building block for a system that tracks refactorings as fundamental operations and knows how to do merges on your AST instead of on the serialized text form of your choice. But as far as I know no such system exists yet, and merging complex data structures is far more complex than simply merging their individual building blocks.
Cohen, as remarkable as always, comes across in this instance as trying to "sell" his product (codeville) and taking it a little personally when Linus isn't swayed. Linus clearly defines the problem he was trying to solve, and believes that git solves it better than codeville. He would have backed codeville if he thought it met kernel dev needs better than git. After all, he had jumped on Bitkeeper before, against other people's wishes.
"It's sometimes better to know that you don't know the answer, than it is to think that you know the answer."
- Linus Torvalds
"I am wiser than this man, for neither of us appears to know anything great and good; but he fancies he knows something, although he knows nothing; whereas I, as I do not know anything, so I do not fancy I do. In this trifling particular, then, I appear to be wiser than he, because I do not fancy I know what I do not know."
Plato: Apology, ~ 400BC
https://en.wikipedia.org/wiki/Trial_of_Socrates
for "corrupting the youth of the city-state and asebeia (impiety) against the pantheon (gods) of Athens."
So my claim above that it's about the religion vs science still stays. "Science" is of course relatively modern term, the older one was "natural philosophy."
A man who says he knows something is also being less of a wanker than a man who says he knows nothing, especially if the latter then uses that comment as leverage to claim wisdom over the former.
It's about the claims based on faith versus the acknowledging the scope of what we actually know so that we can actually find out, which produces the absurdities in the religions for which they have to shame themselves today, as these simply don't match what we know today for sure.
Socrates was sentenced to death for impiety, that's the part of his own, obviously unsuccessful, defense.
Whenever I hear a person say that 'we really know nothing', it always reminds me of Insane Clown Posse being angry that science has an explanation for how magnets work... as in, finding comfort in expressing ignorance :)
https://en.wikipedia.org/wiki/I_know_that_I_know_nothing
"Evidence that Socrates does not actually claim to know nothing can be found at Apology 29b-c, where he claims twice to know something. See also Apology 29d, where Socrates indicates that he is so confident in his claim to knowledge at 29b-c that he is willing to die for it."
So he was what we'd today call "a scientist" being ready to admit the changing but the finite limits of the knowledge while acquiring the new knowledge, not "an ignorant."
But it's also not surprising that not everybody even understood what was that about.
Damn people love this trope.
The thread was on HN last year, I think the title was something about Linus being to smart for his own good or something similar.
edit: "The curse of the gifted programmer (2000)", https://news.ycombinator.com/item?id=11077799
Bonus: the two times this link showed up on HN before and got more than a couple of comments: