My Git Habits
blog.plover.com
blog.plover.com
I can understand the need for this kind of cleanup when pushing a fix to an open-source repository that needs pull requests to be self-contained, but for an internal company repo, how important is it to keep the commit history this clean?
If I have changes I'm not ready to commit and need to switch branches to work on another issue, I find that doing a STASH is an easier way to go. I can just stash my working copy changes, switch branches, then come back and apply the stash and keep going and then make one final commit with just the final changes I want to commit.
Other DVCS actually believe that being able to modify commit history is a bad thing and lean toward immutable commit history (e.g., Veracity). Git makes it pretty easy to modify commit history which is ok for local branches but can be easily misunderstood to break your branch if you're trying to modify commits that have already been pushed to a remote repo.
I like the idea of keeping the commit history clean but I'm not sure that it's worth the effort that it takes to manage the process. In the end, only your final good code is going to be merged into an integration or master branch anyway.
If you ever want to git-bisect your code-base to track down when you introduced a bug then you'll be very thankful that every commit is functional and passes your test-suite (apart from the tests relevant to new features on a topic branch of course). Having to hop forwards and backwards from each git-bisect point looking for a functional commit is a huge waste of time.
Breaking your changes into clean commits with proper explanatory commit messages also makes it much easier for people working on the code in the future to work out the intent of various parts of the code.
Think of a messy commit history as a form of technical debt: Sure it's quicker to just move forward with development, but if you have a history that breaks things down into clean commits with well written explanatory commit messages then you're making life much, much easier for whoever has to debug that code in the future. Odds are that that person is going to be you & by the time you come back to the code you've have forgotten all about it and be forced to spelunk through the commit history in order to work out what on earth you were doing. Think of it as a service to your future self :)
I agree that this property is very useful, but I disagree that it it is necessarily implied by the workflow as described in the article. By using "git add -p", he is constructing a tree that probably never actually existed during development - hence there is no guarantee that it works and passes the tests.
I strongly agree with you that a clean logical progression of commits is a good thing (especially for code review). However, making sure that each stage works and passes the tests takes extra discipline.
% git stash
% make test
...
OK
% git stash pop
Also keep in mind that the initial tree contains WIP commits that are unlikely to work or pass tests, so reordering commits can hardly make things worse.When you do need your commit history to be as neat as possible, git gives you plenty of tools to make it happen. Imagine the nightmare a task like this would be with Subversion. (I haven't used Hg a lot so can't say how it compares in this regard--I'd be curious to know.)
The workflow in the article seems like it's trying to be clever for the sake of being clever. I guess it works for him, but I think most people would find it too confusing to be practical.
Though I don't follow what the author practices, I don't think it is a practice that is done for sake of cleverness. There is a benefit in keeping the history clean in the main branch. My only wish is probably if git allowed one to achieve the same without so many hoops.
* I am hacking on some bit of code, so I set up a Perforce branch. While I'm at it I notice some ugliness and want to refactor it. I could make a new Perforce branch and put it there and then do some horrible merging apparatus but it's complicated and they live forever and it takes me out of the flow state, so I just make a commit. It is now very difficult to extract that commit from its context so that I could apply it independently.
* I make a mistake on a commit that is on a private Perforce branch, or maybe I did something out of order. The Perforce way is "you shouldn't have done that, then." and so my commit logs are full of "fix stupid typo" type commits that are basically just noise.
* I forget to add a file to Perforce and nothing tells me about it, ever, until we get to testing (or sometimes to production!) and Puppet won't run because it can't find some file <foo>. This is fixed with the 2012 betas and p4 status, but those require a server upgrade (because the p4 client is a thin wrapper around what the p4d server understands) and that isn't in the cards yet. This alone is enough to make me want to use git-p4 for everything.
* You can really tell that Perforce is built around a file system of RCS files sometimes. While it has atomic commits, many, many operations are based on individual file revisions (labelling, cherry picking, merging). This leads to the following:
* Cherry picking in Perforce is hard, because it's hard to get handle on the content of a changeset, as opposed to individual upticks in each of the files referenced. It can be done but it is a lot of effort, way more so than in git.
This is not the entirety of my complaints about Perforce, by any means, but they're the most visible UI limitations compared to what I'm used to from git.
And branching is Mercurial isn't nearly as neatly implemented as it is in Git.
But it would seem to address the issues you listed here. For example, you can rollback a commit if you did the "stupid typo" mistake. And you have a lot of power to "change history" and modify the DAG anyway you wish.
A couple of versions ago a new feature was added into core mercurial, called "bookmarks". Mercurial bookmarks are the same as Git branches.
So now you can use the "git style branches" (i.e. bookmarks) in mercurial, or you can use the regular mercurial-style branches (which I personally prefer to git's anyway). In addition you can use "anonymous" branches, which AFAIK do not exist in git at all.
How about trying to rename a file and change the class name at the same time? That always gives me trouble.
My typical workflow with Git is one of incrementally appending commits to a branch, and then using git rebase --interactive to bubble sort commits by impact.
For example, let's say I am working on a feature and I stumble across an unrelated bug that's easy to fix. I fix it. Then I commit just that fix. And then I go about my day. Sometimes, my fix may depend on a small refactor or ancillary change that's not finished. So I make the change anyway, and commit it, completely broken. After finishing the ancillary dependency, I use rebase to reorder the commits so that I can test the change in absence of the bug fix, and then again with it applied.
I've had feature branches grow up to 20 or 30 commits, where 10+ of them are totally borked or otherwise need to be re-ordered. By the time I've made sense of it all, I send it out to my team as 2 or 3 pull requests with 1 to 5 commits each. In the process, I look at my code diff over, and over, and over again. I find lots of bugs by inspection this way. I enjoy writing code this way.
This way my code can be reviewed as a series of logical transformations. It's much easier to review this way. I really appreciate it when my co-workers send me clean patches to review, so I respond in kind.
In summary: It's not about a clean history (although that's a nice side effect). It's about using Git as a tool to help you and your team reason about changes.
I do it pretty much the same way the author does (only I use Emacs and Magit mostly). The reason I do it is because my work has to make sense to other people, and also to me six months later. So I think of my commit histories as stories that must tell people how my software got from point A to B, and tell them clearly.
The chronological history of how I wrote some batch of code is just the first draft of the story. It's messy. It contains WIP commits, false starts, backtracking, and all sorts of cruft that tells more about the noise of software development (and when I stopped for dinner) than about how the software logically moved from one sensible point to another.
Once I've got a chuck of work all figured out, all that noise has to go. It must be edited away before I publish the work upstream. Otherwise, I'm just weighing down everybody who has to understand the work later (including me).
I think it becomes some what of a game to use all the more obscure corners of git when a good 20-30% of it goes a long way.
Using less commands is more.
if you are running as fast as you can for the first releases, what i do is use feature prefixes on the commits. And of course, a commit can have several prefixes.
of course, if i have to cherry pick something, i still have to look at each commit, but with prefixes it's easy. And i have the time i didn't spend thinking ahead for something i don't even know i will need.
I only use git-add -p when I've screwed up and didn't commit when I should have, so I have to split the current commit into two. It seems to me that rebase -i and merge --squash are better suited to re-writing history in the way that's being done here. I'm especially distrustful of any workflow that includes the line "I eyeball the diff".
But I'm no git guru. Is this a common way to work? Are there advantages over the alternatives?
Rewriting your existing history with git rebase -i is fine until it goes horribly wrong & you have to go groveling through the reflog to work out which commits you need to rescue in order to retrieve your lost work.
Rewriting your existing history with git reset can also go horribly wrong, which is why it's done on a separate branch here.
If you want all of my commits to be functionally and semantically separate and individually tested, I can give that to you. git-add -p is just an interface between that well-disciplined software-engineering expectation and what my brain actually does when it gets into flow.
Git allows documentation across files; and because commits are naturally associated with specific revisions, it can't get out of date.
{ Still, it does seem a lot of work, and it would be nice to document within the source itself; and in a way that helps and is needed by the code (so it can't get out of date). A bad example: specifically requiring/importing another class before being able to use it. This documents dependencies, and remains current or your code stops working (it would need to be an error to require a class without using it). It is a "bad example" because it doesn't help you, just raises a barrier then "helps" you cross it, like a stand-over man in an extortion racket.
What's needed is some immediate benefit (e.g. reduce code) to associating files/classes/methods in a crosscutting "module". }
I think he mentions that he does test these new commits that he creates as he goes about re-arranging history, but I think that should be emphasized more. I'd rather have a messy looking commit that passes tests than some nice looking commit that doesn't. bisect is powerful command that shouldn't be broken.
tl;dr , I didn't read that page. :(
Curious, what is the advantage of doing this? Why commit something locally just to reset it out the next morning?
> Such commits rarely survive beyond the following morning, but if I didn't make them, I wouldn't be able to continue work from home if the mood took me to do that.
Especially the part with the printouts and highlighters.
Could this actually be effective?
http://porkrind.org/missives/commit-patch-managing-your-mess...
(Maybe "need" is a bad choice of words: maybe it doesn't need those articles, but it's too complicated if it gets to have them. You don't get such avalanche of advice for a simple, no BS, tool).
Now, the complicated part means it's flexible --in the rare cases you need it to be. But it could probably use a facade that makes the common use cases more intuitive (there are some half-baked attempts that I'm aware of).
Personally, I use mercurial.
The existence of these articles doesn't tell us anything negative about Mercurial, or Git, it merely tells us that these tools are powerful enough to be used for Real Work in Real Workflows, and that there are few things people love more than talking about their workflows.
[1]: Random examples: http://stackoverflow.com/questions/448567/best-practices-in-... http://mercurial.selenic.com/wiki/WorkingPractices http://stevelosh.com/blog/2010/08/a-git-users-guide-to-mercu... https://blogs.oracle.com/kto/entry/mercurial_best_practices http://mercurial.808500.n3.nabble.com/Best-practices-for-mer... http://www.stevestreeting.com/2010/04/22/mercurial-queues-ju...
I long for the time when we won't have to do so much ancillary work to just sync up files. My bet: In 5 years, all these contrived workflows will be replaced by:
sync file1 file2 "fix bug 5782"In addition to syncing, I use Git to keep track of my history, to maintain multiple versions in parallel, to work on different features at the same time without interfering with each other and probably a bunch of other things I'm forgetting. All this just for projects with one programmer--for work and group projects, I use even more complicated features.
So yes, if all you want is syncing, perhaps Git is contrived. But if you actually want version control, it's actually pretty simple.
</snark>
The Photoshop Tips and Tutorials are about doing something new in a realm with inherently infinite possibilities, that is bitmap editing/drawing.
Managing source code shouldn't have "infinite possibilities" -- there are a few common use cases, and several more uncommon.
If you have "infinite flexibility" in your source code management system, you are doing it wrong, or at least inefficiently structured or streamlined, that is, in the wrong end of the scale of 1 => you do everything by hand, 100 => the computer does all for you as you want it to.
Not all workflows can fit an one-size-fits-all team, sure.
But that doesn't mean that you should have to micromanage the workflow even in the most common use cases.
I.e Git is more of a "source code management DIY kit" than a "source code management tool".
In my team we've had long discussions about how we want to code, what to branch, when to branch, when to merge, etc. all about how to keep things organized.
I think people like Git because of the flexibility. Different people and teams can use it in different ways. They don't feel boxed in to a predefined workflow.
I've used both git and mercurial. I would tend to agree that git has more of a feel of non-conformity when it comes to best practices and default ways that it behaves.
That said, mercurial has an equally bewildering array of choices when it comes to all the different add-ons that you can install to do similar things.
I think they're both great tools and I love the way that they've both innovated ahead of each other and copied the best of each other.
git add . && git status
git diff --cached | mate
git commit -am "this is what i did"
git push
and when it hits the fan... git reset --hard