Git's commit workflow is backwards, and encourages bad habits
rory.bio
rory.bio
As for getting distracted and refactoring something else - you're context switching. When you do that, you can also context-switch your repo by stashing your current changes (git stash) while you work on that tangent. Or commit your current changes to their local branch (you're liberally using branches, right? Branches are cheap in Git) and switch to a new branch for refactoring.
If you really want to write your commit first (which is a great idea, because you should know what changes you're planning to make before you start working on them!), then you can create an empty commit when you switch to your new feature branch:
git commit --allow-empty
Write your "plan" as the first commit in that branch. Then do your work, committing frequently, and when you're done and squashing your changes into a single commit for review you now have that first commit available and waiting.Most issues I see from people getting confused with their changes and using Git is not using enough branches for their work, not being practiced at branch manipulation, and not being disciplined and intentional about their changes. People aren't thinking about "how am I going to release this?" or "how can this be reviewed most efficiently?", and they get carried away with making a lot of changes at once instead of many small changes.
My rule of thumb is that I'm only allowed to take one step off the path/plan, anything else gets put on a todo list to revisit later.
Plan the work, then work the plan.
[1] https://www.informit.com/articles/article.aspx?p=1235624&seq...
Different people have different standards and opinions on what's acceptable. Personally, I'm in the camp of people that would hate to see an unrelated typo fix being included in the same commit as some other change. I strive for semantically meaningful commits, and prefer the people I work with to do the same. It's fine to make such fixes, but they should be in separate commits (though they can be in the same code review request). The pain point the author of the article mentioned is very real (needing to do `git add --patch`).
> Write your "plan" as the first commit in that branch. Then do your work, committing frequently, and when you're done and squashing your changes into a single commit for review you now have that first commit available and waiting.
There are often times when multiple end commits make sense. Squashing/rebasing down to essentials is fine, but that still may be more than just one commit on occasion.
Because who wants to go through previous commit files to check for typos in other peoples code. People wont, either they fix it as they are noticing them or not at all.
They don't do zero harm though. The git blame history gets thrashed. If I see the last change one line 15 of a file is associated with a commit titled `Enable TCP connection reuse`, I'm going to wonder what a change to a line throwing an unrelated exception has to do with that. This means I'll waste time investigating a red herring. If I need to see the blame to find out who to ask about the feature, I also get garbage direction.
The same preference for semantic commits applies to things like optimizations. Don't sneak in keep alive timeout changes in the same commit as `Add email notification support`. If a bug turns up on the email feature, I don't want to waste time potentially looking at the wrong thing.
> However, forcing people to split it into two commits likely means it wont ever get done.
That's not been my experience. They either get corrected at code review time on first release, or someone will make a semantically isolated commit for the fixes and include them in a future code review/pull request.
> Because who wants to go through previous commit files to check for typos in other peoples code. People wont, either they fix it as they are noticing them or not at all.
Not sure what point you're making here. You don't need to go through previous commits to see typos if they're still there in the current commit.
That's exactly it; the keyword here is 'churn'. There was one OS project (I'm sure there's more) that basically said that small (typo, style) pull requests will not be accepted for that very reason.
But that's where gratuitous code review, tooling, and rebasing come in; ideally a feature is complete and reviewed and tidy before it gets merged into master and 'solidified' as part of a project's history. Ideally anyway.
The harm to history is marginal. It is quite unlikely that you fixed that many typos in all surrounding lines. I have yet to be confused by history like that.
Typos were just an example. There are other changes that are just as harmful to history that people might otherwise think are safe to squeeze in. Examples include whitespace changes and minor refactors. All of these can really mess with history and make things confusing. Refactors in particular really mess with your head and can even make the code review step confusing if the review wasn't sent with multiple commits.
Keeping semantic isolation of commits has value in its neatness. You may be able to get by without doing so, but you'll be better off if you did.
History is not that epically useful that we should avoid cleanups.
This is a subjective matter, so I don't think your point of view here is wrong or anything, but I do disagree with it. History is something I've consulted many, many times on large code bases. It's helped save precious minutes while triaging situations where an alarm went off and customer was ongoing.
Branches may be cheap in Git, but they're expensive in typical modern software eng. process. For every branch, a PR must be opened, code must be reviewed, full test suite must be ran, the testers must ok that the change does not require extending the test suite, and only finally after the whole ordeal we're ok to merge. That's why I prefer my branches big and meaty.
Surely this isn't true for local branches on your machine? I hope you're not living under such a dystopia.
Code reviews length takes exponentially longer and misses way more when branches are "big and meaty"
The approach I try to follow is below. The BLUF is I use a similar but more granular approach and agree with writing the commit before you do the work.
First, I write the commit message capturing the intended design in a file and then try to create a skeleton of API level changes with method names and stubs from my design as my first commit. I use the file as my commit message for this.
Next, as I work I commit freely locally with small commits but I don't push to remote. I do this rather than commit --amend so I can track and revert as needed if needed. I probably could just use --amend.
Last, once I have it working I squash those commits into one commit and push to remote, then merge in changes from master if any and repush, and then start cycle over.
I tend to do this at one level below the issue I am working, so I then still have multiple commits for an issue, one per sub-issue. I will sometimes leave these as is in the PR, or sometimes I will squash these into a single commit as well. Which I do depends on how active the branch I'm targeting with the PR is, how big the commits are, and the extent of the changes.
Now whoever is making this decision is cursing you since you could have just made the dummy commit in the first place and prevented this extra work that went into figuring out when the bug fix commit alone wasn't working/compiling.
Which absolutely should be a new branch and probably a different ticket.
How about calling it "camp grounding"?
https://biratkirat.medium.com/step-8-the-boy-scout-rule-robe...
And the word 'scouting' already means "moving ahead of a group to gather information", which makes sense in the software engineering context too.
One is free to commit any time they like, it's not "Git's commit workflow", Git itself does not have an opinion on when you commit, just how.
Bad habits are not encouraged by Git, people get lazy, or complacent, and learn bad habits.
There are GUIs that allow you to stage and commit hunks of code intuitively, but GUIs are often shunned by people who spend a majority of their time in the CLI - however I am sure Vim and the like have some kind of tool for doing something similar.
Git add -p reminds me of ed. And sure, you can technically write anything in ed. Still, there's a reason we improved on it.
You could create a commit and —amend work into it if that is the route you’d like to take.
Edit: I think a buffer of lines you navigate is just the right metaphor for working with a bunch of text. It's why line editors are generally considered obsolete.
I'm not, I've not seen anything else like magit's diff views: they provide a great overview of the entire state (of the working copy or staging area) and allow very fine and nonlinear staging/unstaging.
While technically `git add -p` provides the same capabilities, magit's allowance for jumping around and staging and unstaging (and reverting / killing) as you wish and go is a huge boon, it requires significantly less pre-planning and remembering that `git add -p`'s linear and blinkered (one hunk at a time) workflow
The only thing that I've not found perfect is reverting parts of an existing commit (e.g. re-reading your log you realise that you added debugging crap to a commit): magit will create a staged revert and an unstaged re-add, which is nice to move stuff out of a commit but a bit confusing for straight reverting it. And even that is still miles ahead of everything else I know of.
The only thing magit is lacking is fixing interactive rebases: it beats the pants off of doing interactive rebases manually, but i'd like it to compute "dependencies" between patches so it stops me from moving a patch above its dependency, or at least tells me beforehand (as that usually generates huge conflicts which would be better handled by not moving the patches that way because it's an error, or splitting the patch being moved beforehand).
Edit: I also dislike hunks, they never correspond to what I want reliably. Another annoyance is I think magit is quite good at the low level selection in general, but it's hard to get a high level overview because I'm always collapsing and un-collapsing things.
I wish it allowed me to stage a hunk, too. But I make do with staging all lines from that hunk. It is not a lot more effort.
I think a buffer of lines you navigate is just the right metaphor for working with a bunch of text. It's why line editors are generally considered obsolete.
Git add -p is still pretty useful for reviewing individual changes, but slicing finer than that, or editing patchsets directly as text, is definitely a dark art, but one that I'm glad that I have learned through repeated trial and error (e.g. recognizing that the first column must always be a space, plus, or minus). I would probably have added to my prior comment that a specialized diff editor paired with git add -p might be more useful than doing so with a usual text editor.
git add -i lets you get the status of files in the working directory, add entire files, revert entire files, or patch individual files, all from a single UI "session". You can do all of the same stuff with appropriate pathspecs to the individual commands such as git add -p just/one/file-i-need.md, but git add -i is an option for those that want an "almost GUI" for that.
If I'm wrong about the previous point on how one would carry out the post's vision with current Git, then I would attribute some of that to my opinion that Git CLI isn't very good UX-wise and is why people have made a mountain of GUI tools just for Git
Honestly, I know that Git supports the kind of workflow the post describes, but there is some psychological barrier to making seemingly "permanent" commits because my coding workflow doesn't work that way. It works more like the way the post describes. I like to toy with code, make mistakes, backtrack a ton of times, sometimes restart from zero, or even if there was somehow the "magic pill" solution, it would be awesome if it could somehow automatically associate the files you worked with for a specific feature and know that some subset of the modified files are irrelevant, or at least reduce the number of context switches each time you like a new addition to a giant feature
Our idealized next generation of voice assistants might be able to handle such tasks.
"I need the files I was just playing with added"
"Are the lines 3-10 from main.c related to these?"
"Yes but not lines 8-10"
"Ok, I've hidden the changes that you believe are irrelevant. Try testing again. Does it still work?"
Unfortunately our voice assistants are absolutely nowhere near the kind of fluidity that this kind of voice-based workflow demands and if attempted today, would be extremely frustrating or cause a lot of mistakes. I can't even tell Google or Siri to go backwards on a conversation if I misspoke. I need to restart the entire conversation like I'm speaking to a toddler
Stuck? Add a parenthetical note (TK - need X to happen to Y here - TK) and move on. (Why TK? Because it rarely occurs in English words and so makes a useful, searchable marker for changes or challenges to be added or addressed later.)
Coding is akin to writing. Get in and stay in the zone, let it flow. Don't worry about the other things.
If you want to commit to support rewind and blame, do so, just don't push (or don't push to the real remote, push to a working remote set up for the purpose, one you reset periodically). Commits are cheap and git provides plenty of tools for massaging them to make sense (and even appear to be in whatever order you think best!) after the fact before pushing them for real. Between patching and cherry-picking and squashing and other capabilities, the tools are there.
Committing is like editing. Well, more like publishing. It is the act of recording your work...
...and telling OTHERS what you did.
They are different activities. Both can be hard work.
Do the context switch and do the second in the way that makes the most sense to your workflow, to the workflow you share with others. It might be tedious and it might be a pain...
...for example, if all changes need a distinct issue, so you have to cherry-pick the typo, create an issue for it, commit the change referencing the issue, etc., etc. I'm glad I don't work there, wherever that there is. I know theres like exist, though.
I'm glad the author developed a tool and approach that helped them, but git isn't to blame for the conflation of two remarkably distinct activities.
This is pretty much the operative line of the post. Rather than invert my whole development flow, I just use a UI like Magit which makes patch commits trivial.
(and 'git push --force' to 'git yolo', then later to 'gpf', but don't tell anyone)
Get 90% of the way there and for the rest if you've got unrelated changes that are readable, just move on, commit them together and eat the pain.
That's my belief, but of course git permits others. Git is very low level, it's a tool to build a workflow for you.
My workflow involves tolerances for imperfection. I believe in developing like I move and act: my movement has easing, compensation, and feedback. I focus more on being quick to patch errors than on never making one.
But I don't write code for a Mars rover and I don't know if I would have a different workflow there.
The utility offered though doesn't address the fact that now non-related changes are still bundled together.
We've just managed to write the git commit message earlier?....
So you're working on "git plan --commit-message='important feature'" or whatever the syntax would be.
Then you encounter a place where some refactoring would be really helpful, so you do
"git plan --commit-message='refactor: ...'"
Any changes made during this plan won't show up in the original git plan.
I'm not sure if this is how it works, but that's how I understood it.
In reality, git-plan just keeps a directory of draft commit messages for you, and when you run `git plan commit` you can choose one of them to use as the template for your commit message. Essentially, it runs `git commit -t <YOUR_PLAN_FILE>`.
(I'm the author)
I suppose an ideal tool would be for me to highlight a bunch of lines in the IDE, say "These code changes are part of plan: [ABC/XYZ/Create a new plan...]".
It's even earlier in development than git-plan, but the `commit -m "some message"` command works.
My workflow is to edit code, figure out bottom-up how it really works, then look at all the mess, realize I could never get review approval, and stash it. Then I start over top-down, make the same changes again -- wiser this time -- add in comments which make sense to others, and commit to a new branch. Repeat with another new branch until everything in the stash has been incorporated.
With `git plan` I'd have to skip over the 1st exploration phase. Desirable, but irreducible.
> to facilitate the feature you're working on you decide to make a small architectural change elsewhere in the codebase
Imagine this happens _after_ you've already made a bunch of commits. So you write some code. You make some commits. And then you run into a wall where you suddenly realize that you really should change something else and do the code in an early commit differently.
Writing your commit messages in advance doesn't solve anything related to git. If you actually already knew all of the commits you'd need to make, you wouldn't ever suddenly decide half-way through to make an architectural change that you hadn't planned to make from the beginning. But you _don't_ already know all of the commits you're going to need to make. That's the problem, and this proposal doesn't address it.
So now you make a new commit for the architecture change, and then you make a dummy commit for the code that uses the new architecture change, and then you git rebase interactive to move the dummy commit after the earlier commit and mark it as a fixup, and you move the architecture change commit before the whole thing, and then you pray that nothing is broken in the process.
The problem was just that I was distracted or forced to solve a different problem than I had planned. Making a plan doesn't change that. If anything, it leads to ending up with a planned commit message that is still unrelated to the changes (which seems worse than having no commit message at all)?
If you find a typo or something you want to be in another commit, stage that line and commit it and carry on. Use rebase and cherry pick later to move commits around as you want.
However, the dangerous part is including a refactor in an unrelated change. Good luck backing that out if you need to. Personally, when I'm in this situation I just park what was supposed to be doing and then do the refactor separately.
It helps reviewers and is more transparent.
Bundling 12 unrelated changes does a disservice to your other team members.
I personally didn't like git GUIs and only ever used it from the command line until the release of Sublime Merge[1], which imo is hands down the best git client. Exactly what I need, not loads of extra cruft. Plus the performance you expect from the people that made Sublime Text.
I almost immediately paid for a license after I started using it and still actually use git from the command line for a variety of other things, but for staging changes Sublime Merge is invaluable.
Sometimes I do take the time to organise the changes into multiple commits, but that process is very prone to making mistakes.
The most annoying part is when I accidentally forget to add a crucial one-line chunk belonging to a particular feature I'm working on, so that I have one commit "refactor: ..." that accidentally has an important config change or a one-line change that actually belongs to the feature. So now the feature commit doesn't represent a fully working program. Very annoying.
I'll try this out, see if it works any better. Hope it does!
I consider small fixes unrelated to my main change acceptable if they're in the same file. If not in the same file, that means they're probably easy to put in a different commit or PR. Which one of those it is, depends on how big the change is, and whether it constitutes a meaningful change.
Having to think about this after I've done the work has been done, seems to be the right time to think about it for me. Though maybe not for the author. For people who prefer a different workflow, I guess it's nice to now have that option.
But here is a situation I find backwards. If I am doing work, and then I go to commit. I do not want git to do a merge or tell me there is a merge conflict. I mean... I do want to know there is a merge conflict, but I do not want git to try and merge and dump a bunch of merge conflict information in my folders/files.
What I want to do, but git does not seem to have a simple way of doing, is... forget I am trying to commit. Pull down the repo to the latest. Put my files back in. Then let me do a git status, git diff so I can see what my changes are clobbering. What is the correct set of git commands to get it to do that?
In fact, I think git outright encourages you to make lots of small diffs.
Feel free to squash groups of your commits before you merge in.
Makes git bisect debugging a hell of a lot easier as well.
In VS Code, you click the source control button, you instantly see a list of everything that has been modified.
Click on any particular file, you get a visual diff. Press the "+" button, you stage the selected file/files, add the commit message and check them in. Continue that until all the changes are in place, properly sectioned off with appropriate comments.
Yes, this doesn't apply to anyone using the raw git commands, and knowing those is important. But in getting stuff done and keeping everything nice and neat, developer tools to the rescue.
At first I bristled at the notion that the workflow is wrong, but this places the "todo list" front and centre. That's an integration of workflows that makes sense.
Even if that's not appropriate, gut add -p really isn't that hard (and git commit/add -a is a really bad habit you shouldn't use unless absolutely necessary, e.g. you are making the initial commit).
As many others have said there are already great UIs to help you manage these things retroactively, but there is probably no good substitute to properly planning your work and executing on that plan (if order is your top priority)
IntelliJ products allow you to select only relevant lines in commit dialog.
And, I really, really dislike the term Boy Scouting. It is anti agile. The only result is usually that the none senior developers are changing things that simple is not important and causing new bugs. Formatting code etc, I'm fine with.
Or is it "Agile" to never fix spelling errors because there's always a bug which is more serious?
Is it agile to not vaccum the house because the toilet upstairs is dripping? Or am I mixing up agile with Agile now?
What I'm saying is that all these issues goes into the same funnel as everything else and get prioritized. If you feel that all these small issues is a problem. Fine, fix them. Otherwise, your senior developers will spend more time reading unnecessary pull requests about garbage than producing real code.
You (and the op) is barking up the wrong tree. Surely you have to understand it is not the fault of git that ends you up in this stupid situation. It is a broken process.
Edit: grammar, not all :)
For example:
> Neither of those choices are good, but git doesn't give us much choice.
and
> Git forces you to write your commit message at the end of your coding session - after you've already turned your thoughts into code anyway.
No, git tends to not force anything, it gives you tools to stash, commit, rebase, push, etc. you are free to use them at any point of your coding session. There is little to none workflow made mandatory by git, which is how various workflows can be supported by it (github, patches over emails, git-flow, etc.)
> If you context-switch during your coding session, git is not aware of it. In-fact, git isn't even a part of the picture until after you've finished writing code.
maybe that's a problem of your workflow, not the tool, and again if git plan helps you with part, that's cool. This is however does not look like a typical git situation to me.
> The "right" thing to do is to spend time un-tangling your changes, staging them in groups and committing with well thought-out commit messages.
The encouraged thing is to commit early, commit often (http://www-cs-students.stanford.edu/~blynn/gitmagic/ch05.htm...):
> So commit early and commit often: you can tidy up later with rebase.
In your examples, my workflow would be to use git between each context switch.
> Picture the scene, you're in the zone working on a new feature in your codebase when you spot a typo and fix it.
magit-status, mark the line, commit it with a quick message "fix typo". If you find that magit is cheating, then (1) it feels like cheating, in a good way, and (2) git add -p file is still quite easy to do?
> Then, to facilitate the feature you're working on you decide to make a small architectural change elsewhere in the codebase.
oh, let's git stash first
> Finally, before finishing your session you decide to refactor some code in one of the files you were working on.
Ok but first, let's do a quick git commit, then work on that. Eventually I rebase interactively to write better messages, clean things up, then push.
I see the benefits for a backlog in Git by using git plan, though. I would have used Gitlab Issues for this...
Maybe that could be taken further - have autocommit every few seconds, and then have a `git add --patch`-like workflow, but one that selects time ranges instead of line ranges.
Later with "rebase -i" I move all such commits at the start of the branch, and in case deploy them with an earlier merge.
Context-switching is hard. Having the discipline to prioritise effectively is hard, using git branches effectively is hard, and building tools to magically make it easier is...hard.
Git-plan doesn't solve this problem outright, but I hope it moves in the right direction.
I do think that once you have draft commit messages in your workspace, it is possible to build new tools that consume them to help you context switch. For example, a tool could automatically stage hunks for you, based on the draft commit message you wrote.
There are some great thoughts and criticisms here and I'll be taking them all onboard for the next couple of versions of the tool. If anyone wants to contribute, PRs are welcome and I can help out as needed.
Every single commit and upload is immediately followed by `git checkout -B master origin/master`.
Use WIP tags for things you intend to keep working on and download the patch and upload new patch sets as you continue. Using the code review tool as a substitute for branches and continuously doing work off of a clean master has changed my life.
Git-gui (bundled with Git) doesn't look that nice, but is also pretty fast for making partial commits.
For some reason your website is blocked by Tulane university.
I have no idea why.
Proper use of "git gui" and "gitk --all" makes doing the right thing (independent semantic commits) become the easiest thing, remaining in the flow.
After your flow session, you can cherrypick or git rebase -i at will to create branches after the fact.
gitk looks dated but it's rock solid and dead fast even with huge repos. Its only actual drawbacks are some smart but not-well-discoverable features.
the missing gitk documentation https://gitolite.com/gitk.html