Distrusting git
benno.id.au
benno.id.au
... "and I was getting ready to commit a series of important changes" ... Before doing so, I want to merge in the recent changes from the remote master, so I do the familiar git pull. ... "maybe I’m going slightly crazy after 3 days straight hacking" ...
Do I interpret this correctly as that the author has not commited any changes for 3 days?
With SVN there may be an excuse for this, but with Git the right way is to commit as often as possible, and then squash your commits before pushing them. With such a workflow the problem would have been a non-problem - just use git reflog and checkout your previous version.
Of course you wouldn't use a git pull then, but just rebase your local commits on top of master.
Learn how to use your tools, instead of complaining about them!
It was a clear case of git violating a promise of safety made in the documentation, and a nice study of tracking down the bug.
That said, hats off to the author for tracking down the problem. Regardless of workflow flaws, the behavior he observed is a bug, and I'm glad it's fixed.
git fetch
git rebase origin/master
instead of: git pull
which you can also do with: git pull --rebase
But really, the preferred way is to use topic branches. So if you're on branch "foo", this is how you integrate into master. git checkout master # because it always matches upstream
git pull # this will always fast-forward
git merge foo
git push
That does create a merge commit but it does it in the right direction (master 1st parent, topic branch 2nd). You get to see the parallel development which is good to preserve in the history. Otherwise, rebasing is nice because it keeps the history linear.Didn't know about `git pull --rebase`. Thanks for that.
With `git pull`, it is all done in one operation.
Of course, if you've already reviewed the code or are pulling from your own remote repository then `git pull` is likely fine.
1. I might choose not to merge. The upstream code may be bad, or it may not be ready for a merge. I may need to do some more work to prepare my own code to merge.
2. I might wish to perform a fast-forward merge (if one is possible).
3. I might wish to force a non-fast-forward merge (if a fast-forward is possible). This is often a good idea, as it helps keep groups of commits related to a particular feature isolated.
4. I might wish to rebase.
The reason this happens is because git uses directed graphs and there is no link between the point at which you diverged from master and the other changes came in. The merge commit creates this. Unfortunately, it looks like you just committed everything that changed on both branches and the history can be difficult to track and/or confusing.
Fetch then merge will make the history more clear, but you still create a merge commit that will include all of the file changes (correct me if I am wrong here). git pull --rebase will take your changes out of the equation, pull in the new history and then replay them against the updated master branch. This is nice because your commit will only include (at this point) any merge conflicts resulting from your changes.
For a better explanation, with diagrams, see http://progit.org/book/ch3-6.html
I don't think this is really true. You'll just introduce a single merge commit on top of both your commits and the other person's. Yes, if you diff that merge against your own work, you'll see all the work that other people committed, but I think it's fairly well understood that a merge commit is just that - you merging your changes with other changes.
It can be a bit tricky to read the log when merge commits are involved, but try the --graph option or a graphical log tool.
After you do a fetch you can go on to do a merge just as though you did a git pull, but if you break down the steps then you now have the option to do a rebase or other operation as you see fit.
git pull --rebase
is a good practice. That way you can abort and roll back if there are any serious conflicts.But, in any case doing the fetch would have shown that the file was not modified in the fetched tree, and it would have gone ahead and overwritten my changes (without even listing in the merge log that the file was changed).
FWIW, I agree that fetch + (merge | rebase) is in general the best way to go, but I think there is a case when you a pulling in a simple fix from head where doing a pull is legitimate.
After all the documentation says this is a safe operation:
"If any of the remote changes overlap with local uncommitted changes, the merge will be automatically cancelled and the work tree untouched" (from git pull --help).
If this isn't meant to be considered a safe operation git pull should abort if there are any changes to the working directory.
You're losing history that way: If the merge doesn't actually work, then you've got a screwed up file and no way to roll it back.
It's one of the great strengths of git that you can commit files even if someone else has changed them. It's a bad idea to merge when you have anything significant checked out (I'll leave temporary debugging changes checked out, or very small changes, but that's it). Heck, it's a good idea to check in every few hours, to track changes your making.
He appears to be using git with the svn workflow (svn update; svn commit).
I run svn diff multiple times a day. I frequently have many different patches related to different functional areas and bug fixes. Effectively I work with local branches implemented on top of svn. I have a whole suit of scripts that automate this. Seemed like the only workable way to use svn in a disconnected fashion to me.
Why don't you just use git-svn or hgsubversion and stop doing that informally?
Or at the very list use something like `quilt`, so that your changes are represented as a stack of patches on your svn "upstream"
It's still not cool to clobber work like that.
git fetch
git stash
git pull --rebase branch-name
git stash apply
All this cp business struck me as a bit of not understanding how to properly work with git. I've never had my work clobbered by following the strategy above. But I also don't go very long between commits on my working branch.git pull --rebase branch-name
I personally do a "real" commit before rebase, but "git stash" is really the exact same thing. I guess "git stash pop" is less scary than "git rebase HEAD^". Either way, you can always undo that operation via the reflog.
What actually happened here is that the working tree got clobbered. That is a vastly less problematic situation. Working trees get clobbered all the time: rogue "make clean" changes, system crashes, someone-stole-my-laptop, errant rm -rf, forgetting which tree your changes are in... I've lost working data to every one of these, and I never felt the need to blame my tools in a blog post.
So yeah. It's a bug (and a pretty embarassing one). It was fixed. Is there anything more to say? Move along.
"All software has bugs, but bugs that destroy data are pretty devastating." It's hard not to agree with this...
This behavior should be stomped out in first year as a mandatory course, right along CPR.
We are cool like that, we blame the user when out favorite tool screws the pooch big time.
A good tool must promote good habits, and make it difficult to make mistakes.
From the article:
git has just gone and destroyed the last 5 or 6 hours worht of work. Not happy!Completely different thing. The 3 days was relevant because I was tired and my first assumption is that I had done something wrong. It turns out there was a nasty bug.
Also, I don't think my workflow is that broken:
"Linus often performs patch applications and merges in a dirty work tree with a clean index."
[1] http://gitster.livejournal.com/29060.html
If I am pulling down some unrelated changes, it is not unreasonable to do that. That is one of the features that git provides.
I know perfectly well how to use git, in a number of different work flows, and I use whichever is most appropriate for the given project at the given time.
Also, I don't think my post was complaining. It wasn't 'OMG git is the worst, I'm never using it again', it was an analysis of a particular nasty corner case in git.
Its one thing to fuck up and get panned on HN. Its another to go looking for justification that you are right. Its yet another to select a single quote from an article to justify your position when the entirety of the article refutes your claim.
You did a great public service by calling attention to this problem (the problem being not using git properly - nobody is going to hold that against you). Don't ruin it now by getting all defensive.
That is exactly the behaviour I expect from git, and exactly what broke down in this case.
If you don't put your work into git, it can't track it for you. You can always uncommit if you don't like what you committed. In fact, you should consider your local history to be a work in progress; just like you're editing your source code to keep it clean, you should be editing your history to keep it clean. Git never deletes history, so even if you edit it, you can always get back to where you were. What you call "master" may have changed, but what was master before your rebase still exists, in its entirety, inside of git. (It's simply called HEAD@{0} instead of master. See "git reflog".)
Long story below.
However, if you want real fun try out what one centralized repository did for me once. I was (against my will) using Visual Source Safe in the 1990s (ick ick ick). Visual Source Safe at the time represented its data on a server with two RCS style history files (called a and b). When you committed both of these were updated (no idea why there were two) and then as a matter of policy Visual Source Safe re-wrote your local content from the repository. That is on a check-in: it wrote back over your stuff. Fast forward to the day the disk filled up on the server and a single check-in attempt corrupted a and b (so even if redundancy was the reason for 2 file, it didn't work) and the the server stayed enough up to force overwrite my local content. Everything lost for the files in question (no history, no latest version, no version left on my system, forget about even worrying about recent changes). Off to tape backups and polling colleagues to see if we could even approximate the lost source code.
I had a similar issue with Perforce on Windows - Perforce was case sensitive, Windows wasn't, and thanks to CamelCase, there were two files that had the same letters but different casing. (I don't remember the names.)
* Technically, "case-insensitive but case-preserving", which in practice seems to mean, "case-sensitive, except when you need it to be".
I still think it's a terrible default, though. I'm coming to OSX from BSD and using it as a Unix. Retrofitting case-insensitivity onto a Unix is bound to lead to ugly corner cases.
It generally works well, but the other way around does not work: many OSX software (most prominently — as usual — adobe's) play fast and lose with casing, and will break down on case-sensitive HFS+.
On top of which, it would have been an added complication for developers porting from Classic, and an unexpected change for users, in exchange for no substantive benefit.
In truth, the only time HFS+'s case-insensitivity has bitten me was extracting an SDK from Broadcom, and even then, the SDK wasn't about to run on OS X anyway (it's entirely Linux-based), I was just pulling out some documentation.
Indeed, I've had to make password systems case-insensitive for similar reasons.
I definitely recommend against using case-sensitive HFS+. It's a pain to fix because you have to reformat, and you can't restore from a Time Machine backup because Time Machine requires that the source and target file systems are case-compatible. :-/
I prefer case-sensitive filesystems because the separation between works is semantically meaningful in English.
Yeha, I wouldn't blame a tool for not dealing with the brain-dead almost case-sensitive HFS+ filesystem.
If your team is avoiding a tool with a 1 in 16,000 chance of failure then they'd probably also want to avoid flying (1 in 20,000 chance of death by failure), large bodies of water (1 in 8,942) and run terrified from cars (1 in 100) (source: http://www.livescience.com/3780-odds-dying.html.
The car stat seems rather high, and git won't kill you, but the general point is that a 1 in 16,000+ chance of losing a few hours of work is "s--t happens, find a workaround and get over it" odds.
Right?
Right?
No, of course not, and your point is obviously clear. "Blame the user" is still in vogue in some circles, though.
"If any of the remote changes overlap with local uncommitted changes, the merge will be automatically cancelled and the work tree untouched"
Either way it gave me a chance to re-factor the code, and do it better which ultimately ended up being MUCH better anyway.
I don't quite understand what point you are trying to make though. Ultimately the code loss was my fault, however it ended up much better for the source code I was working on because I had time to think about what had to be done and how I should do it, ultimately leading to less code that ran faster and was much better organised.
But the reason it went unnoticed for (horror of horrors) 2 minor releases, was that it only results from an unorthodox (although supported) usage practice.
I really wish merge, by default, was disallowed on a dirty tree. Yeah, it'd be fine if there was a command line argument to override this behavior.
That's not how probability works. That assumes a uniform distribution, but it's not uniform at all: if you run the affected versions of git and run a particular sequence of commands, you will hit this problem every time. Assuming that the Linux community's workflow doesn't change that often, the Linux devs would likely never have hit this because they don't run that sequence commands. Because it's not actually random, we do have control over it, and your suggestion that we all just "get over" the fact that an important tool lost user data seems pretty cavalier.
Yikes. git stash / git stash apply. Pulling into a dirty working tree is asking for trouble.
git losing data is Very Very Bad (and massive kudos to the author for tracking down the bug rather than just bitching about it), but if you're following a proper git workflow (pull to clean working trees, save often), you shouldn't ever be in a position to trigger this bug. That's not an excuse for git to break like that, but the reason that it was likely never seen in the 16k Linux commits is that it's not the "right" way to do things.
I can't explain why the opposite happens. Most of the people I know intuitively commit or stash their local changes before merging. They have this intuition even if git is relatively young piece of software. But then there are always a handful of people that I imagine who could do something like that. And I'm not quite sure why.
One possibility is that it could come down to the level of trust in computers. I don't think I could issue git-merge without git-stash/git-commit first—probably because I don't instictively trust programs to handle complex operations too well in the first place. Operations such as handling unsaved data or letting random commits from different place three-way merge themselves into a single branch. Or both.
This mechanism of distrust might be similar to how drivers who think they're bad drivers are, in fact, the best drivers. They underestimate their capabilities enough to assume everything won't always go right, and then they're a few steps ahead when something goes wrong.
Download: http://code.google.com/p/git-core/downloads/list
Announcement: http://git.661346.n2.nabble.com/ANNOUNCE-Git-1-7-7-tc6849424...
Do not use "cp". Please.
Copying changes to save them and reapply later is nearly guaranteed to quietly lose changes, reintroduce removed code, or otherwise screw up your work.
If you want to move changes stash them or commit them and then apply them elsewhere. Using cp throws out all of git's ability to help you do what you mean and not what you say.
Also, the data corruption was caused by a bug, yes, but the cp based workflow being used will result in a nasty suprise sometime in the future.
If you're on OS X, Time Machine will get you back to where you were recently (except if your home directory is encrypted, then it backs up on logout). Or use Dropbox/SpiderOak/other to keep the last n versions of your changes.
But I was impressed that, unlike so many others (myself included), the author went beyond just complaining. He actually made a real effort to identify the conditions under which the issue occurs. But I was blown away when he actually examined the source code and identified when the bug was introduced. Great work!
My favorite git flow in 143 easy to remember steps
1) git checkout -b MyFeatureBranch # create a feature branch
2) Code/Hack/Fall Asleep on Keyboard
3) git commit -am "Wow finally done with this tiny feature"
4) Go back to to 2 if needed
5) git checkout master
6) git pull # get all the latest changes
7) git checkout MyFeatureBranch
8) git rebase -i # squash commit comments if neccessary
9) Fix merge conflicts, git add, then git rebase --continue
10) git checkout master
11) git merge MyFeatureBranch
12) git push
13) PROFIT!!!
This is borrowed from http://reinh.com/blog/2009/03/02/a-git-workflow-for-agile-te...
code/hack
git commit -am 'did stuff'
repeat
git fetch origin master
git rebase -i -p origin/master
git pushGit is a versioning system that doesn't free you of making backups of your central master. And by master I mean the global central reference repository or whatever you call it.
So you screw up your repository using an unconventional workflow and now you and your co-worker don't trust git anymore ?
Well maybe you shouldn't have blindly trusted it in the first place. It's a better tool than many but still is just a tool you should use with care. As any tool. It has bugs.
That said, I feel your pain. Finding bugs in other tool can be a very frustrating experience. Well, shit happens :-)
But this is really a small corner case and I can see how it went unnoticed for a year.
I almost never use 'git pull' (I do git fetch and then "git merge" or "git rebase" depending on the results), but more importantly I never ever use pull when I have changes in my current working repository. I commit, and then I pull or pull --rebase etc. This way I'm really sure that my data is safe as git has a lot of safety features for committed stuff. all files are stored as objects and there is a reflog to help if you loose track of rebased branch etc.
Another thing that I sometimes do is 'git stash' before pull/merge/rebase etc. git apply later is also a very safe op.
<?xml version="1.0" encoding="utf-8" ?>
Perhaps your browser and mine disagree on when that declaration is acceptable.For reference, http://www.w3.org/International/questions/qa-html-encoding-d...
Edit: I think people took this the wrong way. I mean it's not pedantic to find the real problem. We're all better off because the poster I'm responding to showed us the real problem.
1. don't do what that guy was doing. asking for trouble
2. upgrade to git 1.7.7+. just to be sure.
Look, I have to deal directly with co-workers that misuse git. When I see them doing something bad, I tell them about it. The biggest resistance has been from those that refuse to change.
Here's a real-world example: "git messed up my merge again". What? I go to look into it -- well it turns out they don't understand git and work around it by copying files out of their sandbox, run "git pull", and then copy stuff back in. This is such a recipe for disaster (it easily can and will lose others' changes) that not telling him to "improve his workflow" is disastrous for the project as a whole.
Yes, git messed up his merge. But as others have noted, so could have a bad Makefile rule. All I'm suggesting is that by improving his workflow (committing early and often) then there is no chance of work getting lost, ever.
Sorry if my comments seemed rude or snarky. Actually, I tend to dislike "tl;dr" comments altogether. I don't think my comment added much to this discussion. If I could I'd delete it. cheers
If it was the comment saying '3 days of straight hacking' then I apologise for my prose being unclear. I was only bringing that up because my first assumption was that after 3 days of hacking I was tired and assumed it was indeed my fault.
It was a couple of hours worth of code that needed committing.
And doing a pull, which is meant to safely pull in unrelated changes, to test some things before committing, especially when it was a very simple change on head, is completely reasonable.
I'm not ignorant about the many ways in which git can be used, the post wasn't even particularly blame placing. Tools have bugs, shit happens. Most of the post was tracking down the bug, which is much more useful. Thankfully it was already fixed, so I didn't need to fix it myself.
I don't think its reasonable. 100% of posters here don't think its reasonable. You just demonstrated why its not reasonable. Yet you say its reasonable.
Here's my rule for version control: if an operation can go wrong, it will go wrong.