The State of Version Control: an Infographic
fogcreek.com
fogcreek.com
With all due respect to the FogCreek team, I'm not sure these numbers could be considered a representation of the development community in general. FogCreek makes Kiln, which according to Joel's (FogCreek CEO) blog is "a web-based version control and code review system based on Mercurial and tightly integrated with FogBugz." I know Joel's writing has broad appeal (I know I'm a fan), but it would stand to reason that there would be a disproportionately high number of Hg users in these results, would it not?
I wonder what the results would've been had GitHub polled their userbase? Or if you asked a random sample of Googlers (hint: 100% would say "Perforce professionally", and then a good chunk would say "Git personally")? Or if Fortune magazine had had CIOs poll the individual employees of their firms (I bet you'd see a whole lot more Perforce, and a fair bit of "What's version control?")?
This one has 1300+ questions!
"Adopted by Google" here seems a bit disingenuous; they offer Mercurial as an option for developers that want to host their OSS projects with Google, it isn't as if Google's in-house source control has moved to Mercurial.
EDIT ah the danger of opening up many tabs and then reading them later. Two people already beat me to the punch by 30 minutes:)
> 70% of programmers today are Windows based, 16% use Ubuntu/CentOS or other Linux, and 14% use Mac OS
Is this a skewed survey towards Microsoft orientated people, or is this the normal for the developer industry? (hard to say when you are in a bubble/niche of web startups)
While in the Bay Area, we might think these are skewed because everywhere we go we see Mac laptops, in reality, the employees at big corporations are going to outnumber us.
I think the survey is representative of the geographically-distributed, Joel-on-Software reading, micro-ISV crowd, and not much else. Micro-ISVs often develop on Windows (because that's what their founders are used to) and use Subversion (because it's fairly easy to setup and generally adequate for 1-10 devs). I suspect you'd get very different numbers at YCombinator or Google or Lockheed, though.
I imagine that the FogCreek developer newsletter is going to skew reasonably heavily toward Windows developers.
I asked why at my first job interview. Stated reason was that as our end users use XP (standard at their company), we should use XP. Makes sense.
I guess the cost of mac vs pc explains why macs aren't more popular in regular companies.
As for Linux, well, I think it's just about it being easier to buy PCs with windows than with Linux preinstalled.
If there's one thing that's killing corporations it's this narrow focus on acquisition costs and a near total disregard for support costs. People moan about Apple being "too expensive" all the time, but the reality is most computers are too cheap.
I hope one of the major vendors develops some alternative to Windows. HP is shovelling billions into Microsoft's pockets with nothing to show for it.
I used a windows 2k3/2k8 machine every day for 3.5 years at the same company developing software, and never had a support issue with it. On the other hand, I solely use Macs at home (and, since leaving the aforementioned job to work on my own company), and have also never had a support issue.
However, my grandmother has recently gotten ahold of a vista laptop and a MacBook laptop. She has really shown a greater affinity to the Mac.
Also, we have an internally supported and externally available standard Linux environment (LinuxCOE) which engineers (like myself), who run multiple OSes in their cube, can use.
Speaking for myself only here, I would say that the alternatives are available but we can't really ignore what the market demands.
The cost of an upgrade is the daunting part IMO. We need to buy a bazillion CALs (licenses) if we want to switch to Windows 2008 server-side, for example.
2) If you work remotely, RDP is way better. I've even streamed video over RDP and forgot that I was streaming it from a remotely located computer.
3) If using SSDs Win7 will have better perf and longer lifetime.
4) Virtual folders.
But I would say that if your machines are completely locked down, for example all XP machines are not on the internet, just some closed internal network, then I think XP might be doable. Once you touch the internet though, all bets off.
I've used Windows at work for my last two gigs. At home I use OSX and Linux (both desktop and server). My wife and kids use Windows (what can I say? She's the boss).
I've been posting more on HN recently, but I find HN links and discussions tend to fall sharply into three categories: general technical/business discussions between well-informed people, material about YC start-ups or other businesses that have no particular relevance to me, and flamefests that are full of people who think they know everything and clearly don't know much about anything. The first group is enough to justify my coming here. I'm trying to be better at ignoring the second and not feeding the trolls in the third.
Reddit is a curious middle-ground, because of the subreddit system. I follow a few subreddits that generally have high quality content and one or two that are mostly lighthearted, and I ignore the rest, including most of the main ones.
TechCrunch sometimes carries interesting news, but as with HN, a lot of it isn't particularly relevant to me and a significant amount is just the personal opinion/ego of whoever got to write on TechCrunch this week.
Given this balance, it doesn't surprise me to see the relative popularity of the sites shown as it is in the infographic.
This is from an unscientific survey of my friends and coworkers.
Maybe I'm used to always reading Linux market share at 1% so I want to see that big circle even if it is from a non-representational group.
I use TFS in my professional job and am pretty surprised that so many people like it. Everyone I work with either hates it or lives with it. Doing stuff like moving changes from one branch to another is unreasonably hard.
Every single thing it does is a little bit worse then the tools its replacing in my case.
The out of the box diff sucks, its slow, buggy, get latest dosn't, commits sometimes just fail to commit, the CI builds fail for random reasons, the ticketing system only lets you remove time from tasks (not track them).
Its stupid you have to use Visual Studio, Shell Extensions AND command line just to make it work correctly.
Anyone considering moving to TFS better seriously consider their needs before doing it. If you are purely a MS shop, are prepared to spend days fighting the tool, or are happy to go with the defaults, and have big beefy servers sitting around doing nothing then consider TFS.
Otherwise go SVN Jira Fisheye and Crucible. Trust me it will be better and cheaper.
Is it just me or doesn't this not reconcile with the relevant portion of the graphic?
HgInit.com is the site that got me to Mercurial. I now use it for all my personnal projects. Unintended side-effect though: pains with working with SVN at my day job became more obvious.
On the other hand, I do find it hard to believe that hg users love their tools more. Generally my impression has been that git attracts a lot of people who are, shall we say, passionate about their particular version control system.
Git's UI has always been very poor. Moreover, on Windows, merely installing a basic Git client requires jumping through silly hoops: if I wanted to run Linux on those PCs, I would be running Linux, after all.
For most developers, even those who are quite happy using CLIs in general and running on UNIXy platforms, having a UI that sucks is a serious disadvantage. Outside of those people working on major OSS projects like Linux where Git is the standard and those who like to use GitHub, Git has few compelling advantages over the other serious DVCSes to make up for their much better usability.
Also, just as an aside, Git is the only DVCS that has ever screwed up a project I was controlling with it due to a data loss bug. That puts it in a class with... well, only SourceSafe, really... in terms of how much I trust it to keep my code safe.
Hmmmm. Momentum is a moving target. I'd be blown away if there was another source control tool in history who's userbase has grown as quickly as gits has in the last few years. As for important, I don't really think any of my tools are 'important'. They just let me get important shit done.
Also, I don't even know what "Git's UI" is, but I'm sure you're right about it sucking. I do take for granted my level of comfort using CLI for source control.
I'd be curious about the details of your data loss bug. Most git data loss is user error (though I'll admit the ridiculously obtuse commands and concepts make these very easy to accomplish when learning the tool)
FWIW, I meant the CLI. I have rarely seen an interface that manages to make a relatively simple idea seem so complicated.
> I'd be curious about the details of your data loss bug.
It must have been a couple of years ago, so I can't remember the exact details and I imagine it's been fixed by now anyway. Basically, there was a problem (duly documented in their bug tracker; it wasn't user error) where Git would refuse to update one copy of a repository from another properly. If memory serves, it was related to switching between branches in some way.
It's possible that I was being unfair in calling it a data loss bug in Git, because I have a vague memory that they determined the remote repository itself wasn't corrupted if you knew how to rescue it. However, the effect was that the local working copy on my development machine didn't have the data in it that it should have, and ultimately I don't really care why my data isn't there, only that it isn't.
The Mac numbers don't surprise me at all, as the large influx of unix people to the platform are more used to using CLI tools.
(yes, I'm saying that on the whole, Mac users are less afraid of the CLI than Windows users. Shocker!)
I disagree with the blanket statement about Mac users. Developers using Mac are not Mac users by a long shot. For most of us I know it is a case of the Mac being a better *nix dev box than the alternatives.
But when switching to Mercurial, the gui tools weren't as good, so I learned how to use it via the command line. And, to my surprise, I found that it's much easier to use the command-line rather than the gui tools. Everything goes much faster.
So if you're a Windows dev, like me, and have never tried solely using the command-line tools before, I suggest you give it a shot. You may be surprised like I was.
(I DO think this means there's a great deal of potential profit available to someone who releases a GREAT git gui for OS X)
Personally, I use vimdiff.
What is it you'd want from a "Great Git GUI"?
You can't use Git completely without a GUI [1] and not having everything in a GUI means I'm forever typing some command and then typing another command (or several) to make sure what I expect is actually what happened. To "visualize", if you will. In a GUI I could just see it and this would save a lot of time.
[1] Diffs are visual (vi is a GUI). I'm sure some clown will show how it's possible to use ed or something so you can really do everything without a GUI, but the effort that will take demonstrates my point nicely.
Aside: would've been neat to see a "version control by OS" pie chart for GNU/Linux.
This mostly seems to be because of its performance.
Having used git and bzr extensively (and a bit of mercurial), I've found it's performance to be perfectly fine for general use on reasonably sized projects, and the ease-of-use to be far superior. Figuring out how to get Bazaar to do something you haven't done before is much easier than trying to track down a Git feature. Although I will admit that Git is improving in this regard, while Bazaar continues to improve it's performance and focus on interoperability.
Anyway, in summary: I really don't get why more people aren't using Bazaar.
* Emacs uses Bazaar for "branding" reasons IIRC.
If the survey was done badly, with an 'other' option with no way to write in what you use, then potentially all of the 'other' block could be bzr. It would just be left out by the survey writer not having heard of it.
I'm annoyed at how the very most basic workflows in Git seem awkward, like I'm working against the tool instead of with it. The simplest example is: I have a hacked-up tree, but I know changes have been made upstream, and I want to pull those changes:
$ git pull Updating 73f91c3..0ee9fa8 error: Entry 'README' not uptodate. Cannot merge.
It's complaining because I've modified README locally, which was also modified remotely. Every other reasonable version control system I've ever used will happily merge the upstream changes with my not-yet-committed local changes. But Git refuses. This is annoying.
What I have to do now is commit my hacked-up, non-compiling, possibly-swear-word-containing in-progress changes. I really dislike this. To me, "commit" means "take a finished bit of work and add it to the global history." I really dislike having to commit something that is extremely unfinished just because I wanted to integrate some upstream changes.
So what I usually do in this situation is "git stash", "git pull", "git stash apply." This works ok for the "pull" case. But what if I have multiple sets of locally-hacked-up changes? Like suppose I was working on one change when I realized that there's something else I should really fix first. "git stash" quickly becomes limiting, since you can't name the individual changes, so you get this list of changes that you don't know what they are or what branch/commit they were based on. In other VCS's like Perforce, you can have multiple sets of independent changes in your working tree. Not possible with Git AFAIK.
Anyway, I'll probably keep using Git, but I'm not as enamored with it as I once was. I used to figure this was all just porcelain issues that would be refined over time, but it doesn't seem to be getting any better.
Commit-before-merge is not a limitation, it is a feature of DVCS's. Merge-before-commit is broken by design, as you are modifying your unsaved work, and there is no way to get back to your pre-merge state (say you decide the merge conflicts are too much to deal with at the moment).
So you are correct, git will not let you merge if youmhave uncomitted work that would be affected by the merge, but it's really just trying to keep you from losing work, not trying to annoy you.
That said, this is one of the use cases for 'git stash' which will set aside your uncommitted work. You can then do the 'git pull' and then unstash (stash pop) your work. The advantage of this is that you can always undo the merge if you decide it's not what you wanted afterall. With merge-before-commit, you'd have no such option unless you manually set aside your work.
But what if I have multiple sets of locally-hacked-up changes? Like suppose I was working on one change when I realized that there's something else I should really fix first. "git stash" quickly becomes limiting, since you can't name the individual changes, so you get this list of changes that you don't know what they are or what branch/commit they were based on. In other VCS's like Perforce, you can have multiple sets of independent changes in your working tree. Not possible with Git AFAIK.
This is what branches are for. Do not be afraid to commit work in progress... you can always polish up that work before you share those changes. For example:
git checkout -b feature origin/master
edit, ut oh, interuption,
git commit -a -m WIP
git checkout -b bugfix-1234 origin/master
fix bug, git commit -m "fixed bug"
git push origin HEAD:master
git checkout feature
git reset HEAD^ # removes the WIP commit, but leaves its changes in your working copy.
I hope that gives you a better idea of how you can use branches. You might also want to spend some time reading up on rebase -i. Basically, start thinking of your localc branches as independent patch queues that you can freely edit, reorder, etc until they are ready to be shared. At which point you can push them out to the world.HTH.
The same thing could be achieved by having Git automatically create an "undo" commit before performing the merge. Then you could revert to pre-merge state with "git pull --undo", just like you can abort a rebase with "git rebase --abort."
Optimize for the common case. Of probably hundreds of merge-before-commit operations I have performed with other VCS's, I can't think of a single time I have wanted to undo this operation (after all, if upstream has changed you're going to have to merge sooner or later -- it might as well be now). On the other hand, I am annoyed by commit-before-merge every single time I perform a pull.
> This is what branches are for. Do not be afraid to commit work in progress...
I'll have to try the "git reset" approach, I hadn't thought of that before. But even that has issues IMO:
1. "git reset" is a data-losing operation if you call it with certain parameters. For example, "git reset --hard HEAD^" would throw away the WIP! I'm wary of making such a command something I type all the time, because there's always the risk that I'll call it wrong.
2. I have to do this dance of "commit, checkout, pull, checkout, rebase" just to pull upstream changes. And when I want to actually push the change upstream, I have to "merge --squash" and then delete the old branch (otherwise lots of branches will build up and I won't know which ones have been committed and which haven't). It's a lot of annoying overhead. I'd rather just work on master where I can just "pull", and branch only if I really want to work on two big changes in parallel.
> HTH.
I really do appreciate that you were genuinely trying to be helpful (as opposed to other replies). But my frustration remains that Git doesn't let me work the way I want to, and makes me perform contortions to fit its way of working.
I'm a regular on the git mailing list, and I've never heard of anyone desiring this workflow. It is unusual to want to merge into your work-in-progress. I'd go so far as to say "you're not using git as it was intended". Perhaps this helps a little:
http://www.mail-archive.com/dri-devel@lists.sourceforge.net/...
That said, it would be straight-forward to script/alias what you describe:
git config --global alias.cleanpull '!git stash && git pull && git stash pop'
But really, that's not how git is intended to be used.just like you can abort a rebase with "git rebase --abort."
Actually, rebase is much stricter than merge -- it won't let you start unless your working tree is completely clean. At least merge only cares about whether the files it needs to touch are clean.
Optimize for the common case.
That's not the common case, you've just been brain-damaged by non-DVCS's into thinking it is. :-)
"git reset" is a data-losing operation if you call it with certain parameters. For example, "git reset --hard HEAD^" would throw away the WIP
git config --global alias.popcommit "reset HEAD^"
2. I have to do this dance of "commit, checkout, pull, checkout, rebase" just to pull upstream changes. And when I want to actually push the change upstream, I have to "merge --squash" and then delete the old branch (otherwise lots of branches will build up and I won't know which ones have been committed and which haven't). It's a lot of annoying overhead. I'd rather just work on master where I can just "pull", and branch only if I really want to work on two big changes in parallel."merge --squash"? It really sounds like you're trying to use git as if it's subversion or cvs, and it just isn't.
You can check for merged branches with "git branch --merged origin/master"
If you just want to examine upstream changes w/o integrating them into your current work, you can use "git fetch" and then "git log master..origin/master".
And if you find you often need to be working on multiple branches at the same time, you can always make an additional clone and/or use the new-workdir script in the git.git contrib directory.
Perhaps git just doesn't fit your notion of how a VCS should work, and that's fine. But your annoyances with git seem to stem from trying to use it not as it was intended. :-(
<tangent>Git is a powerful VCS with a rather-awful CLI, but built on top of simple and elegant concepts. Trying to derive a mental-model of how git works from its CLI is fraught-with-peril and will lead you astray. It is worth learning how git works conceptually, and then mapping CLI commands to those concepts. For small projects, it doesn't really matter, but for large projects, git is extremely flexible and you can do things with it that I cannot imagine doing with any other VCS.</tangent>
Best of luck.
$ hg pull
pulling from ...
requesting all changes
adding changesets
adding manifests
adding file changes
added 1 changesets with 1 changes to 1 files
(run 'hg update' to get a working copy)
$ hg update
abort: crosses branches (use 'hg merge' to merge or use 'hg update -C' to discard changes)
$ hg merge
abort: outstanding uncommitted changes (use 'hg status' to list changes)However, if it is a direct descendent, it will try to merge for you:
$ echo b >> a
$ hg status
M a
$ hg pull ../a
pulling from ../a
searching for changes
adding changesets
adding manifests
adding file changes
added 1 changesets with 1 changes to 1 files
(run 'hg update' to get a working copy)
$ hg update
merging a
warning: conflicts during merge.
merging a failed!
0 files updated, 0 files merged, 0 files removed, 1 files unresolved
use 'hg resolve' to retry unresolved file merges
$ hg resolve -l
U a
In this example, the upstream repository made a change to "a" that conflicts with my local, uncommitted change to the same file. It uses its normal merge machinery to try to resolve the conflict.As far as I know, Git doesn't allow merges in this situation.
So there's no way to back out of that right? i.e., get back "a's" state before you ran update?
(Ah, I see from update's help that you have to use --check to prevent it from touching uncommitted files.)
Hmm.
No it isn't. Git is trying to be predictable.
> Like suppose I was working on one change when I realized that there's something else I should really fix first.
Branches are designed to resolve this problem. When you want to do something that takes a lot of time, you branch from the stable version. When you need to do something else, you branch from the stable version again.
"Tree is totally broken and doesn't compile, but Git made me do this to pull upstream changes."
There is nothing logical/reasonable for me to write in that commit message. I definitely don't want that commit to make it upstream. Sure, I could squash later, but why is Git forcing me to do something that doesn't have any value (write a "commit" message for a tree that is totally broken)?
If you try to use a screwdriver like you use a hammer, you're always going to be disappointed.
"git fetch" and "git remote update" both let you do the fetch without the merge. Once you have the updates, you can then decide what to do with them, which may have lots more options than just a simple implicit merge that most tools provide: rebase, overwrite, ignore for now and handle later when your work is in a cleaner state, etc.
Git has tools to manage this in the case of long running lines of development. Providing them for uncommitted changes is a much harder task (precisely because you can't name your work and roll back to it), and mostly unnecessary because commits are so lightweight. This is your fundamental issue with git -- commits still feel heavyweight to you.
Git's model really is fundamentally not "get everybody's local changes working with the exact same upstream code" (indeed for the upstream author, they may trust their code far more than the code they are pulling). Instead it is closer to "let everybody pick what changes and history to use in order to create their own coherent source". DVCS means each developer chooses what history to treat as authoritative. As such, temporary commits, temporary branches, and rewriting unshared history are encouraged.
The Wikipedia page about darcs says "Although the issue was not completely corrected in Darcs 2, exponential merges have been minimized."
I wonder what "minimized" means here. On the face of it, this seems like a pretty compelling reason not to use darcs for anything important.
I do miss the ability to check in changes that I only ever want to have local though. I haven't found a satisfactory way to do that in Git yet.
[edit: removed question that was answered by sibling post]
EDIT: Removed incorrect statement about 63% stat.
I hear this complaint a lot, but I do this all the time with SVN. Just have a directory (tempbranches) where you create your branches. It's probably something I do once a week (and not because I feel limited to only doing it once per week, but most of my work occurs in my local branch -- and I branch that one only when I want to do something that is more experimental, but will take a few days). I'll grant that its not as clean as a DVCS for this, but it works just as easily. The big difference is that we now have centralized accounting of this action, rather than it being distributed.
Um, are you sure? I find branching and merging in git (after three months using it) much, much, much easier than I did in SVN (after 7 years of using it).
The actual activity of branching and merging seem equally easy to me.
Now what for? I create branches to test merges with other developers code, branches before refactoring my in-progress stuff (so I can go back and pick up the original branch with all its history in case it was a dead-end). I do new branches for simple fixes, tracking down bugs (versioning all test-patches so you can check the next day/week what theories you've already tested is sometimes extremely useful). Sometimes I just do a new branch to get the feel of starting with a clean slate.
Once a month or so a look at all the branches lying around and delete those no longer useful.
All this is only possible because git branching is _cheap_. Orders of magnitudes cheaper than on any central rcs. And the best part is that git never ever loses anything. git reflog shows you _all_ the states you've gone through to the current one. So all these branches are actually valuable.
http://stackoverflow.com/questions/5035531/how-is-dvcs-git-m...
Also, the fact that commits have an atomic life is useful. If I cherry pick a change into a branch and then later merge it, Git generally knows what's up. With SVN this sort of scenario seemed to confuse things. Another advantage to this model is that it's very easy to compare (what commits differ between these, instead of a big code diff).
The differences may be subtle, but all in all I find myself much happier and more likely to use branches than I used to.
I think I do lose history though. And while history is important, if its just history, why don't people say that? People make it sound like you just can't do branches and merging in SVN easily, while you can.
Git will happily merge the result together, in all of the cases I've seen it will do so without even needing user intervention and everything ends up where it should be. With subversion this is a nightmare. If code changes location in a file suddenly merge'ing in more changes becomes problematic.
Recently my boss and I were working on different parts of the same project. I proceeded to move a whole bunch of stuff around to clean up the entire source tree, in the mean time he had added two more files and made modifications to some of the ones I had moved. I then proceeded to merge in his changes, git happily merged in his changes to the now moved files, and added the new files into the original location I had just moved, I moved them, committed and everything was happy. We tried the same thing in Subversion not too long ago and it was a complete mess leaving huge merge issues that took a developer a while to figure out what was going on.
We have a team working on various parts of code every day. Like literally every day. Merge conflicts do occur, but they're not the common case. It occurs infrequently enough that its not a big deal. And probably a quarter of the time the conflict is something you didn't want merged.
I proceeded to move a whole bunch of stuff around to clean up the entire source tree, in the mean time he had added two more files and made modifications to some of the ones I had moved. I then proceeded to merge in his changes, git happily merged in his changes to the now moved files, and added the new files into the original location I had just moved, I moved them, committed and everything was happy. We tried the same thing in Subversion not too long ago and it was a complete mess leaving huge merge issues that took a developer a while to figure out what was going on.
That I do see happen. SVN doesn't like changes to directory structure... merges or not.
I have never had a clean merge with SVN, ever. Whereas with git/bzr and friends doing "merge <otherbranch>" just works most of the time.
In practice, how is branching/experimenting different between the two categories?
Example from today: I was working on my deploy branch, where I typically only make to-be-deployed-within-the-hour microchanges. Then I saw another bug, so I squashed it. Commit. Then I saw another bug. So I squashed it. Commit. Then I saw a major opportunity for a simplification with a refactoring. Then I realized the refactoring was under-tested and that failures would be catastrophic, so I started adding extra tests. Now I'm fifteen commits past the last deploy, I have code that I'm 85% positive works on my must-not-fail deploy branch, there is one commit in there which addresses a bug which I want dead, and it is quitting time.
I have been in this state in SVN before. Recovery is NOT fun.
git branch all-the-work-i-did-today #Creates a new branch whose history looks exactly like deploy's does.
git reset --hard production_deploy_92 #Moves head of deploy branch to the tag of the last deploy, essentially forgetting commits afterwards.
git cherry-pick carefully_copy_pasted_hash #Nabs the one bug fix that I really wanted to deploy today.
Git makes my development and deployment processes better. Transformatively better, in some cases.
Embedded systems have to be damned-near perfect before you ship. Web systems, not so much.
Spolsky did a decent job outlining the headaches that DVCS is designed to solve: http://hginit.com/
I'm beginning to think that Ellison was right when he compared computing to women's fashion.
* Source: http://www.mokasocial.com/2011/02/four-things-about-fogbugz-...
Well, if you use hg. /cheapshot
A quick google indicates http://www.surveypirate.com/ as a tool which would allow large numbers of responses for free, or perhaps google documents.
First step would be some questions though.
EDIT: The OP has been updated to solicit responses from anyone visiting the page, which is pretty much what I was hoping for with this comment.
Windows + SVN is so easy to use (and so entrenched), I'd recommend to new developers & students to learn SVN.
Perforce may be a VCS for managers by managers, but at least it's a technically competent implementation of a bad idea.
Perforce had just under 4 hearts out of 5, with Mercurial and Git being just over 4 out of 5.