Git Reset Demystified
progit.org
progit.org
Of course, the feeling is strengthened by the fact that I still don't completely grok git, but shouldn't a tool that's there to help me be easier to grok? I understood dropbox in 1 minute. I learnt TDD in 3 hours. Why should (distributed) version control be so much more difficult?
I don't know how, just wondering if I'm the only guy who feels like this.
At the end of the day, for me it boils down to a comment by Xiong Chiamiov, made on one of Mike Taylor's rant-y blogposts about Git, which was much discussed on HN a while ago,
"I use it because the benefits outweigh all of the things that you mentioned."
Yeah, it's complicated. But a lot of people think it's worth it.edited to add links to the Reinvigorated Programmer posts, which I recommend reading because there's a ton of good commentary and suggestions for newcomers, even if the discussion got a bit heated:
http://reprog.wordpress.com/2010/05/10/git-is-a-harrier-jump...
http://reprog.wordpress.com/2010/05/12/still-hatin-on-git-no...
I'm convinced that indeed it's worth spending more time to grok [git/hg/bzr] and really get the power out of it, but I just can't shake the feeling that somehow, it should be possible to make something with the same benefits that works much easier.
I was hoping that maybe Kiln is exactly that, but from their promo-page all I can see is a rant about how cool DVCSs are, and not much about how Kiln makes Hg easier.
I don't think the current DVCS are the best thing ever possible, but I do think Git one of the best things we have now.
Articles like "Git for Computer Scientists" are great at this stage. Once you know how git's incredibly simple model works you are well poised to learn more advanced things. Rebasing can be hard for some people to understand and it's even harder if you don't know what commits and branches actually are.
After that you need little instruction and should be able to help yourself. It's totally worth it and it's not as hard as you think!
edit: Plus once you know how it works you get to skip big parts of articles like this and just skim around to pick up what you don't already know. It's pretty light reading actually.
"git reset" in one sentence: git reset allows you to move the HEAD's branch tag to any other git revision. (It is easiest to think about git in terms of labels moving around a tree of nodes.)
Mind, I don't mind tackling this kind of complexity when I do algorithms for getting something done. I'm perfectly capable of understanding what you mean if I'd really invest the time in it that I should. But why do I need to do such (relatively) complex yet abstract thinking when I want store a README, undo a mistake and then share it all with a colleague?
I'd recommend "Git from the bottom up" or Scott Chacon's "Getting Git" screencast. Both tackle Git for newcomers by describing the underlying data structures. Once I had a cursory overview of what Git was actually doing under the hood, so much else clicked for me.
You have a legitimate point, and I think it is because your level of understanding, your need to know, depends on the actions you want to take. At first, it is sufficient to know "git add <new file>" and "git commit -a". You will be challenged as soon as you need to do something beyond the very basic linear, like "git push" and "git pull". Branching, merging, sharing... Version control is just not dirt simple. There is fundamental theory to it. If you work with others, lacking that theory will lead to pain in any VCS (starting with a fear of merges and conflict resolution).
These things that you consider simple actions begin to add up, and at some point, your mental model needs to adjust in order to grasp the relationships amongst these things. Your model necessarily becomes more abstract.
I am not convinced that that means it is more "complex", at least significantly enough to be a real barrier to moving forward with git.
One day I finally sat down and starting going through the Pro Git book. After reading about the fundamentals of git I finally realized that this is ultimately no more complicated than a first year data structures class. It's just a stupid directed graph. All this wasted time scratching my head, only to find out that it's just a directed graph with labels attached to the nodes called 'branches'. All the crazy git commands with all the crazy options are just hacks that let you look at and fiddle with this stupid graph. Ultimately when you want your repository to be in a certain state, you first figure out what you want the graph to look like and then you use whatever git commands you want to make it so.
It's hard to explain how fundamentally simple it is. The actual interface doesn't make it seem like it's simple, and all the terms that everything has doesn't make it seem simple, but it is all a bunch of tools and terminology built on top of a simple data structure. Unlike many tools, it is easier to understand the internals of git than the externals. But once you understand the internals, then you can practically speed read an article like this in the same way you could speed read an article explaining for/while loop syntax.
But why should that be? For example, the git docs use different terms for the same thing (e.g. cache, index, stage).
Also, the git tools overload the same command for different purposes (e.g. git add tracks new files or stages pending changes from existing files; git reset can unstage pending changes or revert committed changes). These are different functional procedures that happen to share implementation details.
In git, a "revert" means you are creating a new commit that is the reverse of the commit you want to undo. The two commits exist and effectively cancel each other out. The danger here is potentially thinking "git reset" will revert a committed and pushed change (sometimes people push early). In general, this is probably not what you intend.
"git reset" moves branch pointers around. So when you "git reset HEAD^", you are really saying you want to move your current branch and HEAD (the thing that indicates where you are in the tree) to the prior commit. The current commit still exists but can be ignored.
Speculation now, but I think it is right... This same understanding applies to the "unstaging". The staging area is another tree node, and when you say "git reset HEAD" on the staged file, you are moving it back into the HEAD state.
One note on unstaging and other resetting... You can reset --soft or --hard. --soft (default) means that you want to dirty up your directory with the differences between the file's current revision and the reset revision. This is useful if you are cleaning up an unpublished branch (using reset to undo).
"git reset" could be called "git move-my-branch-tag" and make more sense.
This is just the internal stuff, you DONT need to know this using git on a day to day basis. I use git for all my projects and most of the time, i mainly just type `git commit -a -m "lol message" && git push`.
Do you need to know everything about MS Word in order to become productive with it? No. Apply the same concept to Git or any software really.
I guess the difference is that if there's something that Word can do that I want to figure out, there's good Help and I can get pretty far by scanning the menus. With git, good help exists too (in the form of blog posts like this article), except that the concepts are significantly more complex than those involved in making automatically numbered chapter headings.
In short, to my experience figuring out how to do something non-trivial with Word is exploration. Figuring out how to do something non-trivial with Git is more like studying an advanced CS class.
The purpose of the article though is to improve understanding beyond that and at an intuitive level. It's the difference between saying Bayes Theorem is p(h|e,c) = p(e|h,c)*p(h|c)/p(e|c) and linking to http://commonsenseatheism.com/?p=13156
The thing I like about git is that yes, as a whole it's pretty complicated, but it's really not that hard to get up and running quickly especially if you use github. Then it's a gift that keeps on giving as you learn more about it while also becoming more proficient. Kind of like vim in a way, which is notorious for its wall of a learning curve, but if you just want to treat it as a normal text editor you can tell it to start up in insert mode and use gvim, then proceed to learn from there.
Note that this comment was written by the author of one of the most popular git books (Pro Git)! His dead tree book is #88 among Amazon's Software Development books. His Kindle book is #16 among Amazon's Software Development Kindle books.
Also, it seems to me that git's behaviour encourages broken revisions (i.e. compiler errors, test failures) because your working directory doesn't match what you commit. And broken revisions will interfere with bisecting.
If you always want to commit everything use `git add . && git commit -a`.
With git, how you code is a non-issue. The whole idea is that you're able to to worry about commits afterwards. The index is a wonderful tool to help you untangle the mess of code that has not yet been separated into logical pieces. Working with git is not just writing out code and then committing it. Making proper commits requires time, discipline and practice.
Part of the problem might be the mindset that a VCS is supposed to record how the development happens, but if you think about it, that does not make much sense. The actual development process of a feature or even a bugfix is often riddled with experiments, trivial mistakes, sidetracking, and other largely uninteresting issues.
Once you have thought about what is logical to record into the repository and create a commit, you can test it. Git stash allows you to put aside all other work while you run tests, and git commit --amend allows you to fix the commit until it works.
Test your commits, and you will not have broken revisions or unbisectable history.
However, in practice i have seen those broken revisions in svn just as well. People forget to add untracked files or they skip the final unit test run. Basically, I believe your criticism is theoretical and no problem in practice.
I never understood that. I can do this just fine even with TortoiseSVN. Just click "commit" and select the files you want to include in this commit. I don't see how I need to keep yet another data structure / piece of "state" in the back of my head for that. I definitely don't see why I have to look for an alternate career because of that.
Am I missing something?
Yes. In git you can also do 'git add -p', which lets you interactively add pieces of a diff to the index, not just entire files.
Now, you could imagine an interface where you interactively select the pieces of the diff when committing, and I believe some VCSs do this. So, even though you were missing something, you weren't necessarily wrong. :)
But having the index can be useful if you want to build up the things you're going to commit over separate 'git add -p' sessions.
There's no need to imagine -- git-gui and git-cola can do this.
Click on a modified file, select specific lines from the diff, right-click, and click on "stage selected lines".
The interface is pretty simple: it just shows you each small piece of diff to the files you are adding, and you say whether you want to include that bit of diff in the index or leave it in the working directory. Alternately, it can drop you into a view of the diff in the editor, and you can add or modify diff lines as you please.
That's a correct argument for why the index has to be separate from your working directory. But I don't see why the index has to be separate from HEAD ... in my model, every "git add" would automatically be followed by an implicit "git commit --amend". So instead of building up your commit in the index you build it directly onto HEAD.
Since you're remotely sane, of course HEAD is a private branch not a public one (because you're perpetually modifying "history" on it). (And to start a new commit, of course you also need a command which advances HEAD by creating an empty commit.)
To put it another way, the index should be just another branch. The git commands are way too complex because they don't treat it orthogonally.
Phase 1: Choose which parts of the working directory changes should be in the upcoming commit.
Phase 2: Actually commit those changes.
The separation is not entirely clean in Git. While this is very appreciated for convenience (e.g. git commit -a), it blurs the line and does cause confusion.
Everything about SCM then pops out. I found myself needing tools to manipulate the data model in some way, and a suitable tool always seems to exist.
It's kind of like the Grand Unified Theory of SCM. With such a simple core, everything about SCM just falls out.
I prefer to understand everything in this way so git is perfect for me. But I can understand that it's probably far trickier for people who prefer to learn top-down by starting with workflow and processes rather than from first principles from bottom-up.
Early on, developers saved old versions of code via manual copying. That was really easy to grok (Label floppy with "1983 Aug 12", copy, put in box). But it sucked.
So developers invented things like SCCS or RCS, which stored metadata. But they were harder to grok, as now you had to understand the idea of files having history (and branches), and needing to be "checked out" using tools entirely different from standard copying, and being "locked" against change. But they helped, even though they still kind of sucked.
So hackers invented CVS and Subversion, which worked on trees instead of files, used implicit locking and merge algorithms, and generally made life better. But they were yet harder to grok! Now doing regular development meant that occasionally you'd get collisions: you needed to understand the 3-way merge algorithm just to write code. And they still sucked.
So Linus gave us git. Which adds a bunch of new hard-to-grok abstractions like the index and commit ID. But it fixes a bunch of problems too, even if 10 years from now we'll agree that it sucked too.
So pick your point along that spectrum and choose your tool to match your comfort level. In short: your complaints aren't anything new, SCM systems have always been confusing to learn.
It's easy to say that something sucks, but it's harder to say why.
commit: a snapshot of your files
tag: a reference to a commit
branch: a moving tag
HEAD: the current branch
index: the next unfinished commit
git-add: copy file(s) to index
git-commit: create a commit from the index
git-checkout: copy file(s) from a commit and redirect HEAD
git-reset: redirect current branch to another commit
git-revert: create new commit, which is the inverse of another commit
An interesting thing is that "the next commit" is an object you can work with in git. This concept is new for svn users. With git you can puzzle together a snapshot before it becomes a "real" commit. So mistakes like "Oops, i should not have commited that file" or "Oops, i forgot a little fix in another file" can be handled with git.
Besides, it doesn't help that the first commands you learn about git (like git add, git commit -a, etc.) actually hide most of the complexity in a treacherous way (especially since they do not advertise the existence of the index).
It's probably possible to learn git efficiently by approaching it as something abstract, not as that stuff you want to be quickly productive with. Sadly, you probably can't realize that unless you already wasted your time on the misleading tutorials.
The article is intended for deep understanding, not as a tutorial.
I like Git, I really do. This article helped me like it more. There has to be a way to get the benefit of that object model without all the grief at the interface.
In particular git checkout <file> which is (a) not obvious (before stumbling upon it, I was trying to find it in git-reset documentation), and (b) looks pretty innocent (as the git checkout branch is quite safe to do if you've got staged changes.)
copy source destination
The pattern of git push command
push destination source:destination
2. What is confusing about "git push <remoterepository> <localbranch>:<remotebranch>"
- blob (like a file)
- tree (referencing trees and/or blobs, like a directory)
- commit (referencing one tree and other commits)
- tag (referencing a commit)