Plan Your Commits
dev.to
dev.to
Interactive rebase is an amazingly powerful tool - a little hard to use at first, but the investment of effort to learn its capabilities pays off manyfold. If you're looking for an easy way to level up, and you haven't investigated it before, I can't think of a better choice.
Learning about interactive rebase greatly aided by local development flow, because it freed me from too much agonizing about commit structure before much code is written. It's a lot easier to decide how to structure commits to tell a story with the code in front of you.
What's the purpose of git commit tree at all if it doesn't truthfully represent the development process?
My feature development branch will have a commit for each logical change, typically 3-5/day done by staging individual chunks with "git add -p" (exclusively) as a "I'm done with this piece" checkpoint, still leaving all debug and still in-progress parts in my working copy only. But I wouldn't hesitate to do a "I'm moving from my laptop to my desktop" work-in-progress commit either. Later commits will frequently undo changes done in earlier feature commits, or just revert a commit outright if it proved to go down the wrong development path. You really can't commit too often in my opinion, and it's incredibly freeing knowing you can make sweeping changes and then back out of them at a moments notice.
For merging, I will rebase on master, squash everything into 1 or a few logical commits, and _review_ the resulting commit diff manually as if I were to send it to a maintainer. I will write an extensive commit message based on the code review as a kindness to my future self. The development branch still exists and can be referred to at any time, but I don't want to see that cruft when looking at the project log.
Pre-push commit history is purely for the developer. A push is a publish, which means you are now presenting work for others to view and understand, code review, maintain, and use to debug. There is no more value to people to see every commit for a typo-fix, rollback, and redirection than there would be if our text files kept a complete and honest history of your every backspace. Nobody else wants to watch the rambling movie. BUT that movie can be important to you for a time, and can remain on your local git repo.
If the commit history is getting long in the tooth, it does help to clean it up fairly often, and I have found that cleanup for long-running branches to be much more valuable to my memory than trying to walk or bisect through every minute of my development process.
A quick note about merges: A merge should have something more interesting than "developer merged branch foo onto master"; a merge introduces something important, and old branch names may not exist in the future at all or on the same commits. If one follows this rule of thumb -- make commit messages useful -- then the fact that you were updating your branch with the latest master becomes unimportant. It is actually noise and confusion for anyone having to dig through commit histories, due to the massive number of forks. In other words, always rebase onto master first before merging.
I look at commit history as a document that I can polish and present to others. I look at my code the same way, trying to make it readable for other humans. This is writing prose. I group like things together. Each commit is a solid microfeature. There are no typo commits.
This leads to a commit style where each commit has a well-defined purpose rather than just being a snapshot of the code at a point in time. You make frequent commits while writing the code, but once it's working, you go back and create distinct commits for each distinct change. While you're splitting your changes up, you sometimes realize you forgot something, made a mistake, or left some junk around you didn't mean to commit, which is a nice secondary benefit.
One squashed commit per feature is also terrible. A commit should contain a set of changes that accomplishes one thing, whatever that means to you and your team.
By example: this morning I was working on some huge refactoring and already had 5 short meaningful commits (not all with proper commit message yet) when I noticed some small bug creepd in so I added a line to fix it. That line then goes in it's own fixup commit for one of the earlier commits. Likewise, I noticed one of the new interfaces would be better if it had 2 more methods. So: new fixup commit again. A couple of commits later one logical piece of the work was done, so I did a rebase to get all the fixup commits amended to where they belong. Then I reviewed all commits carefully, rebasing again and squashing some commits because it made more sense to me like that, and adding/editing proper commit messages on the go. I work best like this (mainly for pretty large changes or new features which add a lot of code - for small fixes here and there it's obviously not as involved) and it does produce a perfecly viable story about the development process, useful to other readers, but only because of the rebasing. Otherwise it would be a mess. This to illustrate it's not always as simple and general as 'please don't do that rebasing madness'.
(i.e. what if git log just showed the highest level aggregate commit message by default)
commenting your changes properly doesn't have to involve destroying history - thats just an accidental failure of git that has caused a lot of people a lot of trouble
(Especially in the way that a lot of old SVN/CVS users came into git wanting a simple straight line view and finding tools like rebase instead of `git log --first-parent`, we've ended up in a bit of a mess where people expect to do the heavy lifting in "sandblaster" tools like rebase instead of better UI/UX tools.)
There are two kinds of commits: permanent history and work-in-progress current stack of commits. The permanent history should not be edited except in extreme circumstance. The WIP stack should be edited mercilessly until it looks as good as possible.
In git, the separation between history and WIP is usually enforced by social conventions (e.g. a decree to the effect of "never rebase master, only rebase features branches"). Mercurial codifies this social convention into the repository by marking commits as public or as draft depending if they've been shared on a publishing repo or not, and forbidding editing public commits without a heavy use of `--force`.
The argument for rebase is that some of your local commits (especially ones like "going home for the day" or "whoops typo") are just as useless a level of detail as your other random keystrokes.
Then there's the awesome `hg absorb` that will let me amend a whole stack of commits at once, automatically. It figures out which commit on earlier in my stack should get whatever change is in my working directory and automatically cherry-picks the right fix from my working directory to the right commit.
http://files.lihdd.net/hgabsorb-note.pdf
If I need to change gears, I make a temporary commit, labelled by some bookmark, and move to a different branch of my repo. If I want to throw out part of my working directory, I use `hg revert --interactive` or `hg shelve --interactive` depending on how permanent I want the throwing out to be. This uses the exact same curses interface as `hg commit --interactive`. I also rebase (i.e. change the base only, not rewrite the commit), fold/squash, and edit commit messages, authors or dates as needed. The best part is this all propagates metahistorically via changeset evolution (available now as an open beta on Bitbucket). It's all really nice, and I wish all of the Mercurial workflows for polishing commits were more well-known.
This way, at the end of the day you can just rebase the whole branch and squash all of the micro-commits in a whole commit implementing the whole new features.
It's nothing fancy, just rebasing (which is very very easy in git).
To use a trivial example, Python allows trailing commas for list items, eg: [1,2,]. With diffing in mind,
list_o_things = [
thing1,
thing2,
]
then, the diff to add "thing3," is a single line , rather than two in order to add a comma to the preceeding line.The solution, then, is the code should be "properly" formatted at all times, to avoid running into a situation that goes formatting -> logical change -> formatting leading to the problem you describe.
There are similar hacks, where git's diff implementation supports ignoring white space, and doing word-based (vs line-based) diffing, but what we're really after is AST-based diffing tool(s).
This is why I love gofmt(1).
The crucial idea is to make the changelog a first-class file that I can see and interact with in my text editor alongside my code. As I bounce around my text editor, the changelog keeps popping up in front of me, making sure it's never too far from my thoughts, and encouraging me to write in it. Everytime I save the changelog, editor automation records my latest update to the changelog as a new commit message. The changelog now becomes conversational, a record of my thought process as I work through a problem. Commit messages go from cacophonous to harmonizing.
Over time I stopped needing these training wheels, but they were very helpful for a couple of years. More details: http://akkartik.name/codelog.html
So with this style one can follow the style :-
Jot down the requirement/problem ==> Think of all solutions for the problem ==> Evaluate solutions ==> Choose best one ==> Jot down commit messages before hand ==> Implement solution ==> Use "git add -p" to separate logical units of your commit as per your commit messages already prepared ==> Tada !!
I have a rough history of what changed in case I need to revert but before actually merging the branch I want to make sure all is cleaned up for future reference.
I found it surprising how little people do that and just push with poor ungrounded messages
Similarly I am a very big fan of the phabricator flow with their arc cli which kind of works very similar
I've known people who don't care about how good the code is, they just keep committing, stating that's okay, because they will make sure at some point, most likely the head will contain fully working and tested code. When they merge they don't squash commits, they just merge because they believe full history instead of the glorified end state. I only do this to my personal project which I really don't care if the project ever takes off or if I am starting out the project for the first few weeks. I just need to get my project working.
I usually take a trade off. I use the --patch option often to split my changes into logical commits. I also try to make sure when I have a working code I commit and if I need to fix a small typo here and there I just squash and rebase. For slightly larger change I can add that changeset as a new commit.
I recommend folks to get used to using gir add -p option. It isn't hard to use the 'e' (manual edit) function, since the 's' split option cannot split lines further for something like:
+line1=2
+foo='this is food'
This line still here
You use 'e' mode to split.I'm not sure if such a UI exists for other coding environments, but I definitely wouldn't bother doing it if I was using command line to commit.
The '-p' flag is also supported on reset and checkout commands.
If it is, what do you assign to each commit?
This isn't as bad as it sounds because git has tools built in to handle merging commits together and deleting unwanted commits through `git rebase -i`
Think of it as cheap time travel.
As for what the commit message is, who cares. I use uuidgen -r but whatever works. If the proverbial manure hits the circular cooling device and you need to back to a working version, you will work your way back via git log -p and similar tools anyways. It's more of a controlled undo than anything else.
I find that testing really helps me define commit borders. For instance, I'll often commit once each new endpoint is tested, new class (if small) is tested, etc. Testing is great because it defines a natural place to commit (once the tests pass and the tests describe the functionality). However, some things aren't tested (shame on me) and I find myself saying to commit once something is working.
Keeping a clean working tree is critical to me. As is using a visual commit tool. I often see better commit messages and chunking done with visual committing
Sometimes we'll keep a running list of things we want to loop back to. For example, a rename to cuts across multiple files, some unrelated. I like those to be in separate commits for a cleaner history.
While it is, as others have pointed out, sometimes easier to commit "messily" and clean up the history.
This is very helpful in some cases, but I see it as tactical and not strategic. In particular it is easy to go far enough down a road that it becomes difficult to disentangle what ought to be separate commits, especially because you need to go back and run the tests on each independently before committing to master.
So if I see a decent commit point (the tests pass!) I will prefer to take it then and there, before turning to refactoring or to the next step in the work.
On those occasions where the ultimate solution was so unclear that we wandered around on a WIP branch for longer than a day or two, I sometimes forced into what I call Athena Commits (they spring fully-formed from the brow of Zeus).
I tend to hold to the short headline, but typically I find we write longer narrative blocks and then longer dotpoint lists of subsumed changes. I dislike this position but, if I have to do it, then my successors to deserve to know why it happened and what, in an ideal world, the shorter steps in smaller commits might have looked like.
One last tip: always always use `git add -p`. I've caught so many commented lines and half-baked thoughts and things intended for later commits this way.
If you're worried about losing work with the "git commit --amend" approach, some things to keep in mind:
* Git keeps your history even when you use "git commit --amend"; it's in the reflog. So you can do "git reset HEAD@{1}" to undo an amend operation, etc.
* All JetBrains IDEs (IntelliJ, PyCharm, WebStorm, etc) have a "local history" feature that automatically keeps track of every change you've made to your code, independent of what git commits you made. I've found it useful several times, especially when I mess some git command up. If you're not using a JetBrains IDE, I'm sure there are other similar tools.
Let's say you're writing code and you need to use component X, and as you dig into it, you realize it's not factored the way you want. Rather than lumping a refactor of component X into one code review with all the rest, it's often better to save your work (using stash or commit), move your git HEAD before your latest work (using interactive rebase), do the refactor that you want, and then apply your other work on top of that refactor. If you're doing more work with component X and realize the refactor needs a tweak, or if you notice a bug from the refactor, you can go back and fix that commit using interactive rebase.
When you're happy with the refactor, you send it out as its own code review with a commit message mentioning "This is preparatory refactoring for feature Y". Then both the refactor commit and the feature commit are self-contained and easy for you to look over and easier for code reviewers to understand.
Un autre monde est possible!
I also make an empty commit describing what the session's supposed to do, then amend it when done.
- When I work solo on a project, I do my "daily commit of work".
- On an open source project, I squash my commits at the end into one per "feature" or "issue".
- In a larger team, I usually work in a branch, then I submit the whole branch for review, this let people checkout the branch, test it and review it, without breaking their own workflow. After this, the branch is merged in the format that make the most sense for the project (usually, but not alway, 1 commit per branch is kept).
For the planning thing, I prefer to create issues (even for my own personal project).
In other words: you commit a lot, right? Do you have an issue for each commit? If not, then you're not using "planning" in the same way OP is doing, and your comments about "planning" don't apply.
When I didn't commit for a long time (mostly during long refactorings) I break up the changes into logical groups and commit them seperately using the interactive commiting feature. Using a good GUI to help me there improved that workflow a lot for me.
Actually, the brief description was the commit message of the last commit that modified each file, you know what I'm talking about, but somehow in that particular repository they did some trick to make these commit messages appear as descriptions of each file.
Or maybe it is just an erroneous impression I had at that time.
Sometimes you have a change partially complete and realize you need to make another distinct change in order to finish this one. Have the discipline to stash this change, do that one in its own commit, then continue working on this one.
I commit when a task is finished, or when I take a break or go home for the day.
Before I push I make sure that I have not broken anything. If it's still work in progress I hide it behind a feature flag.
I push after most commits.
Only using one branch, always ready to deploy.
There are many work flows and this is just one, that I use for the current project. I however find it very simple, it's just commit and push. No fancy stuff.
What you're saying about planning out/writing out what to do before the code, however, is still great advice. I end up doing that in the ticket description so I know what I'm doing before I touch any code.
e.g. refactor code to allow thing X to be pluggable, implement new alternative for thing X, load thing X if user has the right permissions all need to be squashed down to just "implement thing X"
1. Is master always safe and deployable? 2. Does the history tell the logical story? 3. Does the history tell the physical story?
That is, I prefer safety/deployability above all other considerations. Then I prefer a history that makes sense and can be comprehended. Then and only then do I care about the actual order and content of commits.
Rebasing down to a single commit satisfies 1, but ignores 2 and 3. In my experience 1 and 2 go together very well.
It's an overreaction to how messy our commits were before.
Some team members hate seeing merge commits, and our Stash system did a --no-ff merge previously. Plus there's a lot of contributors so people often had to merge master into their branch to resolve merge conflicts so they could merge, as there's sometimes 10 or more merges to master a day. So one branch might have 2-3 commits of work, and 2 merges from master. So our commit history might look like:
* Merge foo into master
* Merge master into foo
* Update foo per PR comments
* Merge master into foo
* Second meaningful commit on foo
* First meaningful commit on foo
This is not a long running branch, it could have been one that was started and had a PR opened at 11am and was reviewed and merged at 4pm.
So they instead configured our Stash instance to do an autosquash and if your PR didn't squash required it done manually.
Now it looks like
* Merge foo into master
* Merge bar into master
* Meaningful message about foobar from someone that does a manual rebase to avoid the autosquash
* Merge quu into master
I'm think the new situation is even worse, but I'm outvoted there.
1) Don't introduce git right away. In the early stages of a project, at least for me, I'm usually making major changes across multiple files. The commits during this time have relatively little value for me. Only once the basic structure has stabilized do I run "git init".
2) The git index is your friend. Always be incrementally updating the index in preparation for a commit. For example, as soon as a new feature has been written or a bug fixed I'll immediately add the changes to the index. Then I'll go back through and look for ways to "clean up" whatever changes I made and then add those changes to the index. Finally, right before committing I'll run "git diff --cached" and ensure each changed line is the smallest possible change necessary to achieve what I was going for.