Git is an editor, and rebase is the backspace key
blog.izs.me
blog.izs.me
People can (and probably should) rebase their _private_ trees (their own
work). That's a _cleanup_. But never other peoples code.
That's a "destroy history"
http://www.mail-archive.com/dri-devel@lists.sourceforge.net/...He then gives specific advice.
It provides (safe) mutable history for local changes (the VCS tracks what is safe and what isn't), and a way to still be able to collaborate on changes without "surprising" the people who pull from you.
E.g.: http://draketo.de/light/english/mercurial/hg-evolve-2013-01-...
The extension is experimental, and I'm not sure how many guarantees are given wrt stability (probably best effort).
Now the extension is in the process of getting moved into core (where the behaviour, protocol, file formats will be frozen).
The text looked to me like one could extract all the commands, run them and check their output against the embedded reference output automatically. I don't know, it had such a rigid structure.
On a second look, the output contains timestamps, so what looks like shell output is probably only shell output the author has copied once into their text.
I'm sure there's an answer to that question, but I'm equally certain it's not the only question. It's not that git solves all this stuff perfectly, or that it does it in the best way possible, it's that lots of problems in SCM are human communication issues that simply cannot be solved by software. One of git's virtues is that it knows its own limitations.
That's what I like about Mercurial (also those days I use git more than I use mercurial due to $WORK).
We had some discussions as far back as mid-2010, at the time the feature was called LiquidHg. The first step was phase (differentiate between changeset states), and now it's evolve (make it flow).
I always wished that darcs was usable "in the real world", as the concept of composability is very appealing. Perhaps there is a compromise to be reached here (local patchs with key sequence points?)
- leaf nodes can and should rebase to ensure whatever they submit is "bisect clean"
- everyone else should never rebase because you'll destroy history
As a leaf node, you should always be committing and submitting the cleanest patches possible. The prime directive is "do not break bisect". There's much more wisdom in:
http://www.kernel.org/doc/Documentation/SubmittingPatches
stgit is a nice tool to manage your patches as a leaf node. It lets you go back and clean up individual patches in your history so that everything you submit is clean.
Maintainers don't rebase, period.
And recently (past 18 months or so), Linus has been yelling at them to not merge from mainline back into their proposed branch before pull request. The rationale is that when you do so, your pull request is now based on untested code.
Most dev communities will never scale as large as lkml, so not all the lkml conventions make perfect sense to adopt. However, like I said, there is still a lot to be learned from the community that's been using git the longest and pushing it the hardest.
When working on a very long and complicated feature on an unknown terrain, I commit "wip"s very frequently and rebase from time to time. I also frequently fork branches (from "foo-wip-1" to "foo-wip-2") before rebasing just in case I ever wanted to look at my messy history. When everything's done and thoroughly tested I delete a batch of branches.
This might sound like an overkill but it costs you literally nothing and might be useful sometimes.
I frequently find myself fixing or extending some open source project that my own systems depend on. Usually upstream will eventually take the patch, sometimes they might not. But I can't wait around to find out -- my own "master" branch gets deployed and lives a life of its own.
When upstream does get around the merging the changes, and I pull from upstream, the merge is 100% safe and clean because git sees my own commits coming back again. But if instead the changes were rebased, the chain of history is broken and git needs to go into conflict resolution mode. If I've changed other things since, I get a mess.
All of this gets even worse if you have things like continuous integration servers and git-based deploys. Those systems will happily deal with merges and break when you go and rebase a branch they've already seen.
Git is indeed an editor, but it's like edlin.
It's not Vim or Textmate or Excel. I want a Git that works like an editor, not requiring me to keep in my head (or in some shell bells and whistles) 'where i am', and what I can or should do now. Not a UI frontend with buttons for single-shot editing commands, but a real editor, that makes use of 2-dimensional space, to somehow really edit everything that matters to me about versions and sharing and changes and all that.
It's a really powerful metaphor. Microsoft did something very right when they released Windows Explorer with Win95, which treats your harddrive like it's a document that you're editing (i bet it's a ripoff, but that's not the point here). Copy, paste, undo, it's all there.
Can't we design something similar for distributed version control? How well would document-editing metaphors map to git? My underbelly feeling says that it might be a surprisingly good match, given the right visualization and editing primitives.
If Git is an editor, the obvious question is... what is it editing?
I'm guessing that the answer is "history", but that's fuzzy to me and I'd have to think some more for a better answer.
1. Something's broken. What part of the code is broken? (perhaps 5 minutes tops?)
2. git blame <file> (10 seconds)
3. which lines in the trouble area were changed recently (10 seconds)
4. git show <commit> to reveal what got changed in that commit and caused the problem
In 99% of these cases the git blame is just to see what the other programmer (or often myself) was trying to do at the time they broke the code -- in these cases it's obvious what's broken, just not why it was changed.
When it's nontrivial to figure out where the code is broken and I have no clue at which point in the history the code broke, it's still easier just to diff the broken code to a known good branch or commit and look for significant differences.
I guess where git bisect slows way down for me is that you have to devise code that will indicate definitively that the bug exists. It's really never faster for me than just eyeballing the troublesome code at that revision.
They do not need to be familiar with the code, the language, that module, the library, the developer or have anything more than the knowledge of how to check out/bisect the code base, and how to reproduce the bug.
For this it is an invaluable tool, and I can't count the number of times on the git mailing list that a simple bug report has allowed a developer to quickly locate the exact commit an issue was introduced and investigate further. A recent example is at [1].
[0] Here a user refers to any of 'technical user', 'developer', 'maintainer', 'original author', or really anyone who has the ability to check out the repository.
[1] http://thread.gmane.org/gmane.comp.version-control.git/21504...
One example of this is mozregressionfinder [1] that was created by Heather Arthur [2]. The tool automatically downloads Firefox nightly builds (in the same binary search pattern of bisect) and lets the user check for the bug they observed in each version. Once the nightly build where the bug was introduced is found, a Mercurial pushlog URL is displayed which can then be pasted into a bug report to aid developers in chasing down the bug.
[1] https://github.com/mozilla/mozregression
[2] http://harthur.wordpress.com/2010/09/13/mozregression-update...
But if you are working on a code base with hundreds of commits per month, I wish you a lot of luck. Your 10s of seconds might easily become days.
On a big codebase you might even want to automate bisecting as much as possible. See http://lwn.net/Articles/317154/ where Ingo Molnar says:
>>> for example git-bisect was godsent. I remember that years ago bisection of a bug was a very [laborious] task so that it was only used as a final, last-ditch approach for really nasty bugs. Today we can [autonomously] bisect build bugs via a simple shell command around "git-bisect run", without any human interaction!
Add to that "where bugs are guaranteed to appear in the same code path that causes them", and it becomes immediately apparent the value of bisect.
> I guess where git bisect slows way down for me is that you have to devise code that will indicate definitively that the bug exists.
Normally whoever found the problem should provide you with a test case. Once you fix the problem you'd add that to your unit tests to prevent future regressions.
Generalizing, whenever you are making incremental changes to a complex system where you don't understand all of the effects of your changes without lots of testing, it makes sense to write small commits so that you can backtrack to see when an unintended behavior first surfaced.
Maintenance programming is one field in which this approach is crucial.
I expect this is where the git bisect advocacy comes from. Bisecting down to the commit that caused the test to fail takes seconds. Then, you can start using the remaining process not unlike you mentioned to actually fix the bug.
It helps remove that initial hunt for the problem. When you narrow the problem to a specific commit it becomes quite clear what is wrong immediately, not having to filter through the myriad of changes that may have come after. That is, at least, where I have found it to be most useful.
Ah, if I could only be so lucky.
This is only one example, and I rarely have a need for git bisect in the code that I generally work on (since I can mostly keep it in my head, and regressions are rare). But in large projects it's extremely useful.
I just updated program X that I'm not a developer for, and it crashes a bunch in really weird ways. I ask on the mailing list and nobody's seen the problem and it doesn't reproduce for anybody else. I can get it to crash on my machine in a few seconds of using it by doing certain things. The backtraces from the segfault are in different places each time.
I clone the git repo, and bisect from the previous version that worked. It eventually finds the commit that broke things, with which it becomes obvious what was wrong.
I might work on a project for a week, write some code, do an intermediate commit, realize it's not needed, delete it. Then when I finally merge into master (rebase), that code I wrote is forever gone. Which I think is fine, but two months later I realize I could actually use it. But there's no way to get at it, is there? Or am I missing something?
That you can create arbitrary branches and keep them forever remotely or locally. If you think you'll need it branch off that point: git checkout -b branch_i_may_need then checkout the previous branch. a git log branch_i_may_need will always show the commit you needed at the top and git cherry-pick `git rev-parse branch_i_may_need` will cherry pick it into whatever branch you're on. That way you keep the commit and can still keep the remote repo clean.
What does anything buy except for a sense of satisfaction?