Distributed Version Control is here to stay, baby
joelonsoftware.com
joelonsoftware.com
This is great. A few years ago, DVCS was weird and strange and SVN was the stuff. When people discover how awesome Git and Mercurial are, they stab SVN/CVS in the back.
I'm not sure why you're borrowing an assassination metaphor to describe someone's shift in technology choices. Technology aren't people.
1) Stop manually backing up versions, that's error-prone. Use RCS!
2) Stop using RCS, it doesn't scale. Use CVS!
3) Stop using CVS, it's too complicated. Use SVN!
4) Stop using SVN, it makes it hard to merge branches. Use Hg/GIT!
... And someday, we'll be saying 5) Stop using Hg/GIT. They ________. Use _______!
Same thing for hardware, and languages, and frameworks... and just about everything. shrug
It's much easier for a new entrant (startups) to innovate because they don't have the same baggage.
But that doesn't change my statement.
If you're not willing to break your business to change with the times, someone else will--and by the time you react it's too late.
http://www.onlykent.com/20100318/blockbuster-bankruptcy-2010...
In VCSes more than in regular software, you really have to get the design right very early in the process.
Similarly, in the case of changing mental models, most users and even developers of X won't be ready for Y when you write it. Trying to force the change onto those people will harm X and handicap acceptance of Y by confusing the Case For Y with the Case Against Killing X.
So it's hardly surprising.
The full quote makes it clearer:
"Subversion = Leeches. Mercurial and Git = Antibiotics. We have better technology now." (emphasis his).
Hg and git are different. It's not like SVN has a bug "it isn't distributed". It is fundamentally not distributed. So the progression is:
1. Stop manually backing up versions, it sucks. Use RCS/CVS/SVN.
2. Stop using RCS/CVS/SVN, they suck. Use git/Hg.
I can easily imagine a new DVCS system that would fix bugs in git or Hg. The git UI sucks for example. But I can't easily imagine a fundamentally different successor to DVCS.
(Not that such a successor won't happen - of course it will. I just can't imagine what it will look like).
But to me, SVN isn't one of those technologies that I'll look back on and despise. Sometimes you look back on a technology and wish it had never existed (the original EJB seems to inspire this in some people), sometimes you look back on it and see deep flaws but still have respect for it as a first crack at a problem (some people look on the original Struts framework this way), and some you look back on fondly even if you no longer personally use it (a lot of ruby/python folks seem to look back on perl this way).
To torture the analogy, I'd say SVN is a moderately effective antibiotic that has been replaced with a more effective one that has fewer side effects. But a leech? Come on, it was a good technology.
Nobody is driving 55 Chevy's around anymore. Some people dedicated aspects of their lives to appreciating the time and space when that technology was dominant, but their glory days have passed.
So has subversion's.
Perhaps it comes down to how you divide work up. I just don't like branches. They're mess.
Eg, if you want to split the development branch (trunk?) into two competing experiments, and then discard the losing branch at the end, a linear history of version/revision numbers split between two branches seems limiting IMO.
Eg2, if you are working on a new feature or set of code that you haven't yet pushed to the server, having an open stack of changes gives you the freedom to modify and improve those changes without cluttering the commit history. When you're done, you simply push the finalized set of changes on top of the target branch.
I spent years using CVS, then years using Subversion, now I have almost 2 years of git under my belt. There is something appealing about numbered versions, but there is no significant benefit over unique hashes of versions, or rather, any small benefits of sequential numbers is vastly outweighed by the tremendous benefits of unique hashes.
I don't think any developer would disagree that the ideal workflow is to do one thing after another sequentially. It's just that in the real world, reasons come up where you might need to create topic branches. Even if you rarely need that functionality, there's no reason to stick with a crippled system like subversion unless you really need the one feature that it's entire design philosophy optimizes (partial checkouts) more than solid fundamental changeset tracking and manipulation.
I'm no zealot—you won't catch me advocating linux vs mac, emacs vs vi, or ruby vs java. In most technology choices there is a wide range of tradeoffs and considerations. However Subversion is one of those rare cases where the tech is fundamentally flawed and serious developers need to move on, whether it be git, mercurial, darcs, or whatever. Subversion has a few use cases, and if you are not a professional developer then maybe it's deficiencies aren't very relevant. However, if you're slinging code all day, version control is your bread and butter, it will stay with you across languages and platforms, so it's insane to stick with a crippled platform that will always be that way due to fundamental design flaws.
I don't think so. I've been using git for about a year, and I still practically never branch or merge. Switching from task to task has a cognitive overhead that I prefer not to pay.
If you want to split trunk into competing experiments (and eventually discard one) using a "traditional" RCS like Subversion, you can do that. Easily. Worrying about the linear history of rev numbers would be like worrying about the ordering of hash tags in git. Just ignore them -- you always can get the change history of the branch on which you're working.
The second example is a feature provided by distributed version-control, not the choice of "changes" over "revisions". It's maybe easier to implement distributed version control with the former paradigm, but it's not impossible in either case.
This supports his point. If you think that a version is a changeset, as opposed to thinking of a version as a monolithic collection of files, you have already achieved enlightenment.
I don't know Mercurial but I imagine it can handle it as well as Git.
It's this feature that you'd have to pry from my cold dead hands. SVN now feels like a straight-jacket to me even when I'm working solo.
Why not just develop that new idea, without breaking the main line of work :/
Can you give an example of an idea you can't implement without breaking the main line of work?
I guess you will always be able to refute any example.
The important thing is: While you can put yourself under that restriction, you will pay for it.
Actually I think one of the biggest reasons I don't like them is that they're a kind of "hidden state". If someone made a VCS (or just an interface to an existing VCS) where branches were just represented by different subdirectories in the filesystem, then I might use them more.
I'd love a RCS that would let me see every branch that currently exists without having to check them out. The branches would exist as directories on my local filesystem. Any file operations (doing a directory listing, or accessing a file) on these local directories would actually be doing RCS commands and pulling data over the network, but I wouldn't have to think about that anymore.
Besides, you can do svn ls to see the branches.
I must be missing what you're getting at.
Am I right?
I have a hard enough time remembering the state of a single trunk, let alone if I had several branches on the go at a time. I'd fail spectacularly at remembering which branch has what on it...
It's easier to just do:
* Never break the build / functionality of code (Good idea anyway)
* Architect your stuff well with minimal dependencies so you can
can swap in/out modules/components easily.
Personally, I feel branching+merging is as much use/fun as filling out TPS reports.Even as I use svn today, I turn almost all my work into patches. For every bug I fix and every feature iteration, it gets turned into a patch. That way it's usually pretty easy if I have to apply it to an older branch for a hotfix or whatnot.
I generally hew to the extreme programming view of branching and continuous integration. Push early and push often. Browsing github, I seem to find a preponderance of projects where the branches are never push-ed back to a master copy. No matter what tool you are using, branches can introduce semantic changes that are hard to merge.
To sum, DVCS not that big of a deal for the way many people already use CVCS, when you end up working with a centralized copy anyway.
Mercurial actually has a whole lot more information: it knows what each of us changed and can reapply those changes, rather than just looking at the final product and trying to guess how to put it together.
For example, if I change a function a little bit, and then move it somewhere else, Subversion doesn’t really remember those steps, so when it comes time to merge, it might think that a new function just showed up out of the blue. Whereas Mercurial will remember those things separately: function changed, function moved, which means that if you also changed that function a little bit, it is much more likely that Mercurial will successfully merge our changes.
Joel's absolute right about one thing, though. Being able to merge correctly and reliably make a huge difference. It's what makes the distributed part of DVCS possible. Without a central repository, branches happen a lot more often, and merging has to work right. The contortions that teams go through to prevent branches are a thing of the past.
http://blogs.open.collab.net/svn/2008/07/subversion-merg.htm...
Not git. Git's repository model is very strongly snapshot-oriented. Of course the whole machinery for supporting changesets exists in git, but it is built atop a system that actively avoids "thinking" in terms of changes.
Linus on this: http://marc.info/?l=linux-kernel&m=111314792424707
"Subversion, CVS, Perforce, Mercurial and the like all use Delta Storage systems - they store the differences between one commit and the next. Git does not do this - it stores a snapshot of what all the files in your project look like in this tree structure each time you commit. This is a very important concept to understand when using Git."
Source:http://book.git-scm.com/1_the_git_object_model.html
However the diffs are there below the snapshot-based model as means of effectively storing and transmitting the snapshots.
I haven't bothered to read the blog post (Joel strikes me as someone with a huge ego and without any exceptional insights), but doesn't the above fact mostly refute his thesis? I like git and hg but they are just tools. You could do the same development model with patches.
a nice upside is that you can merge changes to each feature up into the "combined" branch as they get stable, to test side by side. when you finally finish, you can leave feature a behind by merging b into your release, or vice versa. or maybe even merge both into mainline.
sorry, the point I was trying to make is that cheap branching and merging is still an easier way to manage this situation than just ignoring branches altogether.
A: Create separate branches, manage merging/updating unrelated changesets.
B: Create 2 implementations of something in code, and manage nothing.
I guess it depends on how easy it is to isolate the part you want to create 2 implementations of, so it may depend on what language you're using as well as how you architected things.
Branching+Merging just seems like something extra I have to do, manage, and remember. Like filing. And I still can't see what benefit it gives for many cases.
But that's fine I think... I just don't get it. Maybe I'm too old ;)
An advantage of keeping an old code in its own branch is that it's less likely to become incompatible through random changes. In particular, I'm thinking of bug compatibility: there may be bugs you wish to fix in the new code, but you don't want to fix in the older code because of third-party dependencies.
This may give you millions of combinations, but since the various parts that don't need to interact can't interact, this isn't really a problem. OOP is nice when used by people that know OOP.
But I stand by the tongue-in-cheek suggestion that FactoryFactories are high altitude if not low earth orbit. And obviously, you can use factories and still switch between production and development with a single flag. So please don't interpret my remarks as critical of code I've never actually seen.
I don't see interfaces and plugins as being orthogonal at all, you can't very well have plugins without interfaces (it wouldn't be much of a plugin if the client depended on the "plugin" itself). Besides, orthogonal vectors don't get you to the same place. ;-)
In my opinion, the defining feature of a plugin is that it can be loaded and used with no modification of the code that uses it. On architectures with dynamic loading, this means you can drop a DSO somewhere and use it without code modification or relinking. "Factories", as usually described, would require some modification of the factory to support this new implementation (perhaps just a single line). A "factory" with a runtime-extensible list of implementations that it knows about, is a plugin architecture, but a plugin architecture need not look anything like a factory.
That being said, you clearly can implement indirection without plugins. I think of a plugin architecture as being composition at a coarse level. In the case of a factory with a runtime list of implementations, I think of that as a factory and a plugin architecture, possibly that the plugin architecture is implemented with a factory.
Factory factories, or dependency injection, is a lot more like Futamura Projections than a factory factory factory factory because you want a hammer.
On Futamura Projections: http://blog.sigfpe.com/2009/05/three-projections-of-doctor-f...
(And that's fu-ta-mu-ra, a Japanese name; not fu-tu-ra-ma, the TV show. I read it wrong the first time, and so did everyone I've admitted that to :)
(This is just my mental model, no need to read into it more than that)
Edit: I hope he does, too. I probably disagree with him more than half the time, but there are few other writers on software who consistently hold my attention, and almost no one who consistently makes me laugh.
I wouldn't mind a picture of Taco on each post.
Can anyone say if Mercurial is significantly faster?
After going through Joel's tutorial it looks like Mercurial keeps history from branches in the root much better than either TFS or Subversion. I think I would opt for Mercurial the next time I have to set up source control.
In addition to those conceptual differences, the actual act of transferring files between remote and local on git/hg is also much faster than SVN. Entire checkouts of git repositories are usually smaller (disk size) than the equivalent SVN checkout... of one version. This means less time on the network all around.
When you send commits to a remote server, git compresses them and sends the packfile over the network. The server will refuse non-fast forward merges (the branch you're pushing is supposed to be merged by you first, locally) so it pretty much has nothing to do but uncompress the changes and apply them. Unless you have a really slow network or are sending an insane number of commits, it's not going to take longer than a few seconds.
Cloning big repositories (like emacs, with ~25 years of history) can take a while if you have a slow network, but sometimes it feels like git can actually clone a repository faster than SVN can check out a single revision.
But repos like emacs' are rare. Often, SVN repositories end up being huge because they contain the code of many unrelated projects. In Git, having such a huge repository is impractical; each project should have its own repository. There is support for submodules though, so you can make a "master" repo that contains pointers to the child repositories, if you really have to.
If we talk drawings, photos, whatever binary data is needed for a project - what happens ? Are the binary deltas as good as SVN's ?
Even if the "binaries" are XML text - e.g. drawings in SVG - wouldn't I be out of luck trying to merge changes if 'Beth' added a squiggle and 'Cath' a square to different parts of the drawing ? (many tools, upon writing, reorder data 'ad lib'). Are there "merge tool" plug-ins ?
Also, this is really the same problem even for systems without solid branching capability, like SVN. If two people modify a binary image at the same time and one commits, then the other will get a merge conflict when they update and will have to solve the conflict in order to commit.
Edit: Uh stop upvoting me please. That was a tongue-in-cheek comment and, as sant0sk1 pointed out below, also lazy on my part.
"The last Joel on Software article[...]"
Thanks for all your attention over the years. I'll still be around, just not doing these essays any more.
I wish you nothing but the best with your new venture and with Fog Creek.
Thank you for all your articles. I've learned a lot from them over the years and consider them really first class. Hell, my printout of the Joel Test is one of the three things taped to my desk (The other two being Merlin Mann's 5 Inbox-zero words and "Your Company's App" by Eric Burke)
I've learned more about programming (as in doing so) than at my university. And it helped me recognize the patterns you dissected when on my first job -- made it much easier to quit.
http://www.fogcreek.com/Kiln/learnmore.html#hist_BranchAndMe...
That's about standard, the free 'gitk' does that much, too, for git.
I'm not an engineer, but a designer who likes to program and your thoughts on software are useful even for me. Now that you are venturing into the realm of VCs and being a media company I would've love reading (or hearing in the SO podcast) your insights over that.
Anyway, good luck. And I guess I have ten years of essays to dive into.
"I announce my retirement from blogging effective March 18th"
We work on relatively small code bases with a very small team of people, not large, unwieldy open-source projects with hundreds of contributors. Primary interest is in a log of what has happened to the code chronologically, rather than applying/unapplying specific revisions/hunks frequently. No need whatsoever for people to run with their own branches. With those requirements, git, Hg or Darcs would all present far, far more headache than they're worth versus Subversion.
In some other scenarios, the formula may yield a different outcome...