Why I Switched to Git From Mercurial
blog.extracheese.org
blog.extracheese.org
I started using Mercurial two months ago when I joined the Mozilla corporation, and now use it every day on one of the biggest and best-supported installations in the world. I also started using BitBucket for some of my Mozilla-related personal projects.
It might be that I don't fully "get" Mercurial yet, but I still find myself frequently missing Git. My first impression is that Git has a simple flexible model that supports a complex front-end, while Mercurial has a simple extensible front-end that ends up creating a somewhat more complicated model.
Any time I need to step outside the standard commit-merge-push workflow, I find the higher level of abstraction between Mercurial's user commands and its database makes it harder to understand what I'm doing, and sometimes prevents me from doing what I want. Things like rewriting history (e.g. for rebasing), or combining changes from multiple remote repos with different sets of branches, are still possible in Mercurial - but they usually require plugins to do well, there are more different incompatible ways to do them, and there are more opportunities to mess up your repo.
More concretely, I find that MQ is the best way to do many tasks in Mercurial that I would do in git with plain old branches and commits, and for many of these uses MQ is both more complicated and less flexible than the equivalent git commands. (But there are other uses where MQ is really better than the alternatives.) I also find that it's much more annoying to manage short-lived throwaway or topic branches in Mercurial, so much that I simply avoid using them much of the time.
But perhaps I'm just brainwashed (or brain-damaged) from too many years with Git. :)
Also, I am be interested to learn more about your workflow and why it is so difficult to do in Mercurial. So far, the only feature I've seen in git that I've wanted in Mercurial is the ability to manage publicly named branches using the bookmarks extension, and by extension, an easy way to track named branches.
The DVCS crowd looks down on centralized VCSses and points out that whenever the answer to a VCS question begins with "well, you should have done X before you did this…", it marks a failure.
Fair enough.
But I would extend that further — when I ask "how do I do X with this VCS", I don't want to be told "well see, you don't really want to be doing X". Yes I do want to, and I want a VCS that let's me do pretty much anything I want. That's git's approach.
In part he is dead on.
hg really can't handle large repositories; but to be honest if you're working with GB's of data then you're using the wrong tool :) period.
Point two is a "complaint" I agree with. The repository structure isn't particularly good - it "works" for the most part but sometimes you just want to kick it. And forget hacking around with it..
Silly things like "hg copy [file]" doesn't copy file history but does a remove/add & cp/rm are another ball ache.
Ultimately hg is a good tool; for lightweight, hackerish version control it is great. Beyond that? Yeh, use Git.
If you need something more complex or specialist then git is the clear winner.
Edit: iPad mistype.
But that's a personal preference. Git is just as good a choice.
But a revision control system failing to not corrupt or crash depending on the data you put into it is, frankly, no revision control system.
Just like a database failing to return the same data back that you told it to store and without telling you a problem occurred at store time is no database either (cough mysql cough).
I would group mine into two main categories:
* Error messages that are confusing.
* Requiring that the user knows a fair number of options to commands. You can get by in svn without knowing many options to pass to commands, but with git, you're going to sink if you don't.
Which ones? I have found them helpful, suggesting the right possible commands to use to solve a problem.
Requiring that the user knows a fair number of options to commands.
Which ones? "git add", "git commit", "git push", "git log", "git revert", "git checkout", "git merge", and "git branch" do not require any options unless you want some shortcuts or useful things that SVN cannot do.
That said, there are still some error messages that are quite good, if you understand git's model, but if you don't are nearly opaque. (Easy to google, though.) Git really is unfriendly to you if you're just trying to hack along without understanding how it works, and I don't see that ever changing. It's a power tool, not a butter knife, and while it can, should, and does have guards in the right place to keep you from really hurting yourself (you basically can't lose data in the repository, thanks to reflog; be careful with checkout, though), it's not ever going to be quite as simple as a casual user would like. If it was, it wouldn't be git anymore.
Granted, after a few times I just wrote it down, and now mostly remember, but it's not as simple as it potentially could be.
I was also a little shocked to find that I couldn't simply start tracking a remote branch after a local branch had already been created -- requires git 1.7. The workaround was simple (create a new tracking branch and merge the changes over), but to me this is an example of the UI being a little unhelpful.
I think it's because git's data model is so minimal and elegant, and the UI is anything but.
For me, it's not so much that the UI is quite bad (which it is), but that it could have been brilliant.
I'm talking about fundamental interaction with the version control system, not a pretty gui. git has pretty gui too.
If you want a really nice, consistent UI, check out darcs.
The index is one of the key components in git, but it doesn't really have a name or label in the UI. Say that it was called "index"; then the current "git diff" would be equivalent to "git diff index", and the current "git diff --cached" to "git diff index..HEAD". Less special cases to remember. IMHO, of course.
The same way "git checkout index foo.c" would fetch foo.c from the index (I don't even remember the magic incantation for doing that now). Etc.
Also, I still think that using the index should be optional: "git diff" should default to "git diff HEAD", "git commit" should default to the current "git commit -a", etc.
"git checkout" tries to do too much. It creates and switches to branches, and copies blobs from the repository to the working tree. The first form is reasonably safe, and bails out without --force if you do something stupid, the second does not. At least for me, the only way to learn the difference is the hard way.
The branch-switching should probably be done by "git branch" instead, and the fetching of blobs by "git reset" (which sometimes already does this, but with very confusing options; I never remember the difference between --soft, --mixed, --hard, --mixed and whatever).
That's only a couple of changes that could have been made a long time ago. Now it is definitely too late. Even if the git UI could be made smaller and more logical (with a huge amount of work), it's not worth pissing off pretty much every git user out there...
Obviously, it trips beginners up.
And it encourages committing stuff that may never have coexisted in the working tree - and thus have never been tested together.
Dulwich exists, but it is not a drop-in replacement.
When will I get a true platform independent Git (none of the msysgit stuff) ?
The only place bash doesn't exist is on Windows and installing Cygwin is the least of your worries in that case.
I think the real problem is the documentation, both official and unofficial. The best git intro I could find was gittutorial. It's OK, but a bit unfriendly. Everything else rapidly goes from telling you how to do stuff to teaching you how git works.
If somebody came from SVN and told me that hg was unintuitive I'd tell them it's pretty easy, and they just have to get used to a slightly different model. Then I'd tell they to try hginit.com. I'd also point out how they can still work while the network is hosed, and that even if the server dies they still have all the history. Cool, isn't it?
git users tell people to learn graph theory.
- edit files A, B, C, D, E in the default changelist. - put A and B in changelist #1 - put C, D, E in changelist #2 - oops, C is the wrong changelist. Move it to #1. - edit file F (starts out in default changelist) - put F in changelist #1 - send both changelists out for review. - edit file B based on review. Run tests to make sure both CLs work together. - commit one of the CL's
So the basic idea is that I have multiple changelists open simultaneously and the contents of the changelists are fluid. I can keep changing my mind about them until commit.
Of course, since Perforce CL's are file-based, this all breaks down if I want overlapping uncommitted CL's. Git handles that better, which is why I'm using Git now. But I still miss the ease of use of Perforce when working on multiple CL's. Sure, the commits in Git are only local and you can always edit them until you do a push, but it seems like I have to plan ahead more and commit more often, or I end up having to split or merge CL's using commands I'm not very comfortable with.
1. Hack on files 'foo' and 'bar', but instead of saying `p4 change`, say `p4 shelve`. Write down the changelist number for later; this changelist is stored on the server now. 2. `p4 revert foo bar`. You don't lose any changes; they're all on the server. 3. Hack on files 'bar' and 'baz'. Submit or shelve those changes. 4. `p4 unshelve -c (changelist number from step 1)`. (You'll probably need to resolve here.) Your original changes are back!
Git certainly breaks down if you're working on one massive branch, but it's designed for lots of branching. Go ahead; we're encouraging you to branch for every new feature, every bug fix, everything. Here's the above workflow in git:
1. `git checkout -b first-change`. Hack on files 'foo' and 'bar' and commit. 2. `git checkout old-branch`, where 'old-branch' is the branch that you were on before step 1, probably 'master'. Now `git checkout -b second-change` and edit files 'bar' and 'baz'. Commit again. 3. Checkout the old branch (master, likely) again. `git merge first-change`. `git merge second-change`. Resolve if required. Test. Push changes to the remote server.
Best of all, you don't need to branch before making your changes. You could actually move the `git checkout -b first-change` to after the edits on that branch, and it would still do what you want.
1. Problems with storing big files in Mercurial. You can use the bigfiles extension if you really want to manage large files with Mercurial.
2. Inefficient renames. Making renames efficient is a currently recommended GSoC project. I'm sure this will be fixed eventually even if no one picks it up this summer.
3. Destructive commands actually, well, destroying stuff. This point makes me feel like the author knew git first, tried mercurial out, and switched back, making the blog post a little misleading. I don't know how you could expect a delete not to be a delete unless you were familiar with something like git already. If I were to run a delete command in Mercurial, I would read the documentation on the command first to make sure it created a backup for me. Only if it then failed to create this backup would I complain. Also, I think I prefer the bundle approach of Mercurial. You can always rename the bundles to keep track of them. You could even write your own extension in 5 lines that ran a destructive command and immediately unbundled the created bundle to restore your changesets, exactly duplicating git's functionality.
The first two issues he mentions I would consider to be actually valid complaints against Mercurial, the third just a personal preference of the author. I would consider neither of them to be major game-changing issues.
Or do you prefer to manually backup the files of your project before doing any operation in your editor that is potentially destroying your files? I'd much rather hit ctrl-z as many times as needed once I noticed my last refactoring has shredded all the files it touched.
Being able to easily undo whatever I do to in whatever application I can think of is a nice feature - even more so if the application is what I entrust all my work to.
A branch may have its own, separate reflog which would be deleted, but that is only a convenience feature as the man page you linked to documents WRT the -l option; the primary reflog still records all operations made on the branch.
You might find it useful to actually test stuff before posting links to man pages that you don't fully understand. I know I do. :P
git gc
oopsie data loss. It is about what he did in hg, except in hg it is one command.
git gc only prunes nodes unreachable from branches AND the reflog, which I think by default is 30 days long
git stash drop does print the sha1 of the stash, and as long as you know that sha1, and it's not been gc'd, you can get a dropped stash back. (git fsck will also show the sha1s of dangling commits left from dropped stashes.) But while a separate reflog records existing stashes, that information is not retained when they're dropped. It's perhaps best to think of git stash as a more convenient form of git diff > patch, and you wouldn't expect that to keep a log of the patch file either.
So his whole reason for switching is I shot myself in the foot. And instead of learning more about his weapon of choice to avoid doing that in the future, he traded in his gun for one that is slightly more complicated.
Part of the point of using a versioning system is to avoid ever destroying stuff. Therefore, in a versioning system, you'd expect to have to work pretty hard to make a change you couldn't roll back from, wouldn't you?
Because it's version control; isn't the whole point that I can change things and return to a previous state?
I don't know anything about hg other than it being distributed, but if it destroys things, well, it's not a VCS as far as I'm concerned. That's crazy!
And it makes sense that there's a permanent delete feature, but I'd expect it to be outside of the normal workflow.
So why do you speculate he's using the wrong delete?
The fact that so many people consider "hg strip" an advanced command is part of the problem. Modifying history should not be considered advanced. Being able to recover from ANY command, including destructive ones, should not be considered optional.
Yes, so that deletes a patch. Anything in patches is basically in flux, and I wouldn't call losing a patch "data loss". If you call hg qdel "data loss", you'd call any sort of modification to a patch "data loss", since patches aren't versioned. If you want versioning with patches, use pbranches.
> Modifying history should not be considered advanced.
Maybe, but the only way I modify history in practice is through rebasing. I've never ever felt the need to modify history any other way. What use case do you have for modifying history in potentially destructive ways?
With respect to history modification:
First, I rebase a couple dozen times per day. I'm on a team that doesn't use merges unless we have a reason to (this makes it easier to bisect and think about history). I also create, destroy, and rebase many of my own branches every day.
Second, I amend commits a lot. I'll often spike some little piece of code I don't understand, then start amending the commit as I rewrite it with TDD, until the commit no longer contains any traces of the original spiked version. For more complex spikes and TDD rewrites, I'll do it over many commits, rebasing the spike over the rewritten version until the spike commit is empty and gets skipped by the rebase. Doing that in Mercurial would be... arduous. I can easily do multiple history rewrites per minute while doing this.
Third, I amend commit messages a lot, usually with "git rebase -i". Maybe I forgot the ticket number, or maybe the meaning of the commit changed (see the next point).
Fourth, I sometimes do drastic commit rearranging. This is harder to explain, but it usually involves splitting commits (in the simple case) or moving sets of related changes from one commit to another (in the complex case). These are sometimes at the file level, sometimes at the hunk level, and sometimes within a hunk. This is rarer than the others; I probably do it once or twice per week.
Fifth, I "git reset <ref>" a lot. It took me longer to start doing this, but it's useful in a lot of situations. For example, "oops, I accidentally created a merge bubble."
i don't know about git reset, but everything you mentioned before can be done with mercurial (mq, histedit, ...)
how do you manage patch queues in git? we need them because we are constantly backporting to different versions of our app. branches will mean n merges for n versions. stacked git? or is there a git native way to manage the same?
what about something like tortoisehg? gitk is pretty crude in comparison.
i assume one can glue a diff/merge tool like meld. is the experience similar when resolving conflicts?
thanks for any feedback.
I think it's fair to surmise that he must have been messing with commands he didn't fully understand.
The bottom-line is that in either git or hg you have to step past the screaming sirens and big red warning signs to truly lose data.
Although, I imagine quite a few users --force a command and, unaware of the afore mentioned facilities, write up a nasty blog post about switching to hg from git because git lost their data :)
Once you commit something, it stays in the repository in the commited version forever.
In SVN if you really want to delete a file, you have to stop the service, dump the repository to a text file, edit the dumped file to delete the data (or alternative, do not export the latest revision), restore the repository from the file, and restart the service. SVN has no native way to do that, you have to use some external utility to totally delete something.
Now, it can be argued that this is good or bad design, or that centralized VCSs are bad/distributed are good. But you are making a broad statement about VCSs and ignoring the existence of SVN, and this is misleading to anyone who reads this thread.
Ability to control and be able to reverse any changes you make to the source code tree is an essential feature of version control system - precisely because you don't know in advance what will work and what won't. The fact that people are not used to that simple idea just shows how broken are most of the other tools out there.
There is no reason not to want an unlimited undo for almost everything, especially today when disk space is so cheap.
"git gc" will prune the unused old histories for that permanent effect.
Unless they were also familiar with functional programming and immutable data structures; it's the exact same mental model.
"I would like to argue that none of the user-interface and high-level functional details are nearly as important as the fundamental repository structure. When evaluating source code management systems, I primarily researched the repository structures and essentially ignored the user interface details. We can fix the user interface over time and even add features. We cannot, however, fix a broken repository structure without all of the pain inherent in changing systems."
Keith makes a good point. While Git's interface is lacking in some areas, those things can be fixed in the higher level porcelain. Picking a system with a nicer interface whose plumbing is uglier will lead to problems further down the road.
A DVCS that is incapable of rewriting history is, for me, a nonstarter, so I'm not interested in talking about "well, if you turn off the interesting parts of Mercurial, it's perfectly safe!"
You can either use Mercurial in a mode where it's not at all equivalent to Git, but safe (extensions off), or you can use it in a mode where it's somewhat equivalent to Git, but might lose your data.
With Git, I can mutate history without the slightest bit of fear, often just to see what will happen. Does this giant branch I discarded two months ago rebase cleanly over master? Try it. If the rebase completes but I don't like it, I "git reset mybranch@{1}" and everything is good. And that safety net is always there.
In Mercurial, I have to stop and think every time. Even if the command does dump a bundle, restoring from it is a special case, and I'm going to have to go dig up the bundle file, and in some cases it might not be there and I'll be screwed. In Git, there are no special cases, I always know exactly how to recover, and the recovery mechanism is always there.
Once more, to be clear: In Git, it is always there. Always there!
Let me state it to be clear: mq is your problem, not Mercurial. Mercurial is perfectly safe.
This summarized, for me, why I use hg. It is a tool I am going to use everyday and it's not ok for this tool to have a crappy interface. My hard drive space, that I don't care about.
""" I'm sorry for recommending software with a confusing interface. But you'll be spending a lot of time with it; it's worth getting over the initial hurdle of confusion. """
Note that I call it the "initial hurdle". I know both systems very well, and Mercurial gets in my way far, far more than Git.
Has this changed, or was I always mistaken, or this guy talking about the sum of file sizes being a few gig?
Some of this can be fixed, but a lot of it is probably not going to change. When you git-add something it has to checksum the old data + new data. That's going to be a pretty expensive operation when you have 50TB of data.
There's no DVCS that I'm aware of that handles large binary files as well as say Perforce or Subversion.
Check out this project for more info on tuning Git for big files: http://caca.zoy.org/wiki/git-bigfiles
This is the first article that gives real pitfalls of one vs the other. I enjoyed the read and would love to hear more as well as counter arguments.
Must there? What if most people are switching based on the same vague feeling that there must be some reason behind the shape of the swarm?
(I'm a fan of both systems here, fwiw.)
While somewhat unintuitive in some of the commands or the default arguments you may need, git is perfectly decipherable and memorable once you do get through the initial period.
Additionally, as git is often handled on the command line (even by people who'd usually use GUI programs), it LOOKS more intimidating, but it seems a few weeks in, people are getting by with Hg (As they don't have to get what's going on) however can do pretty complicated stuff in git (as you have to get what's happening or you will be lost).