How to write good Git commit messages
altcampus.io
altcampus.io
I'm sure someone will say "but I use the history ALL THE TIME to source dive and paragraphs of context are super helpful". This is not the case for 95% of developers or projects so I can't really endorse spending time learning this "best practice".
It's fine to be aspirational, but it's such a shame if people see posts like this and think they are failures or "bad" developers or that this is a widespread practice.
If it helps you personally or you have an open source project and you want to help with a changelog, knock yourself out. But there are so many more impactful skills to be learning or spending your time on if you're a working developer in a typical environment.
Your argument that "most engineers don't care about commit messages, so they must not be important" is akin to "most people are overweight, so health is unimportant".
I would say, yes, health is unimportant to the average person because of their actions. I don't think that's good or fair, but it is the current state of the world.
If you want to continue the analogy: my argument is akin to telling someone trying to lose weight to research the most bioavailable supplements -- despite the fact they still eat a Snickers bar every afternoon. It's a micro-optimization that has been elevated to "table-stakes".
You can and should value practices based on your context. But I will be the asshole and ask if writing good commit messages is "so, so, so important" -- what things are less important? Is it more important than a good test suite? Well factored code? System documentation? Capacity for senior staff to answer questions? These things cannot all be so important and, in my experience, worrying about crafting amazing commit messages is way down the hierarchy.
In contrast leveling up from the terrible "WIP WIP do the thing" to something slightly less awful takes maybe an extra 1-2 minutes per commit. And every time you do it you're doing your future self and future co-workers a huge favor.
"WRITE BETTER MESSAGES".
but put WIP in the title so we know not to review it.
but ... do commit often, so we can get builds out. put WIP in the title so we can decide to build with tests or not. or something.
Tying automation steps to 'commit messages' is begging for "WIP JUST COMMITTING TO GET A BUILD OUT" messages, which people then complain about.
Will I disagree that good test suites or well factored code or systems docs are more important than commit messages? No.
However... most projects do not have those things and most developers on those projects are in a "I don't know what I'm missing" sorta state where they don't bother adding them. Exactly like you point out for good commit messages.
So if 95% of projects have inconsistent and incomplete test suites, never-refactored spaghetti code, and almost no system docs, that doesn't mean we should tell people not to try to do those things. In the same way that the existence of poor commit messages doesn't mean we shouldn't try harder.
The nice thing about commit messages, being a small and simple thing, is that it would be much easier for someone to learn how to write good commit messages overnight than to learn how to write a good test suite, or to refactor their code.
If you've reached the point when the next "optimization" you can do is to work on your commit messages, that is awesome.
I also think spending paragraphs on "commit message standards" is overkill. I don't care about full stops or capitalization or anything, I just want to know some basics: "What is this fixing or adding. Why is it done this way. Are there any special considerations that kept it from being done a different way?"
You bring up a good point about location, too; i don't particularly care if the info is in a ticket or commit message or a pull request, as long as I can get to it)
Having commentary built into the artifact that introduced the change ensures your commentary applies to the specific change introduced. It's then possible to tell why someone did a specific thing, and also to tell how things are today, and to decide if that justification still applies, or if things no longer match the description.
If you dig around a little bit, using e.g. the git pickaxe (-S), or github's excellent blame browsing tools, it's not that time consuming and you might find out some pretty surprising things.
Maintaining comments is part of the job of the software engineer, to keep code understandable and maintainable by others on your team. If your comments are constantly going out of date, then either:
1. The team is not putting effort into keeping comments relevant for each other.
2. The comments are too far away from the code that they are referring to.
There's definitely a balance to strike between too many and too few comments, but comments are extremely important.
Commit messages describe why the code _changed_, not why the code is currently implemented in the way that it is. Those are two different pieces of information, and the latter should not be hidden away in a commit message that you have to go searching for. And if that code came from another file, it's extremely difficult to trace it back to the first time it was written.
Not only that, but how would this even work with the "atomic commits" requirement? Most of the time you can't change a single line and create a commit message explaining why that line changed, because it's not atomic. That line change happens in coordination with a bunch of other changes. How do you explain why one of those lines was implemented a certain way in your commit? It just doesn't work.
I'm not saying you did this, but I think that most people who point out little issues with the mundane processes within software development haven't yet grokked the dev process in its totality. Commit messages, code reviews, comments, documentation, unit tests, design patterns and idioms, all these practices' strengths and use-cases compensate the others' weaknesses. I.e. to basically any gripe (again, not saying you griped) like "process X isn't worth it because of some maintenance issue" there's an answer along the lines of "well that's what process Y is for". Together all these processes produce high-quality codebases but when you lack one or more of them the whole thing falls apart quickly.
The exception is ticket numbers - it's sometimes useful to be able to link commits to tickets to see the context of a feature, especially since it's cheap to do.
So, status quo rules, right?
But as your predicted person who nearly always writes this style of messages and who frequently dives into old commit messages, PRs, and tickets: this shit is not hard. It does not consume much time. It is not a deep skill, it's a basic skill, almost at the level of correct indentation in terms of effort required.
It just takes a second or two to think of what your commit does and maybe why you didn't do it the other way, then drop in a ticket reference. A tool can even do that last part! As a bonus, there are times when I tried to explain why I did something and then thought of a better way or a bug I'd missed.
If everybody did this basic exercise that takes just a minute, the commit history might actually be useful. Then most people might use it, who knows!
I think what matters is the type of project. A lot of commercial software is all about moving fast and not caring too much about the past. In such a setup, writing good commit messages may be less useful or even seen as a waste.
For software built to last however, I think good commit messages are invaluable. Issue trackers are replaced, documentation is rewritten or lost, comments are modified; but commit messages stay the way they are. This helps you get a better understanding of why decisions were made, something rarely documented properly.
that's not at odds with what the GP posted.
I've worked in places where people spent a lot of time and effort 'improving' commit messages - reviews on those, rewriting, meetings, etc. after several months, the 'team' such as it was was writing 'better' messages per the few people who made the determination as to what 'better' was. they were the people who made it a focus and said this was necessary.
bug reports didn't go down. time to turn around functionality and fixes didn't go down. code review time didn't really go down. we didn't get more code coverage. code didn't execute faster. no one we delivered business value to was happier, or got more value. not in the immediate moment these changes took place, nor in the months that followed.
Last sensitive code I touched is a trading platform that must be responsible for a few billions of dollars a day (could be a trillion, not sure, didn't finish adding logging). It was initially created in 2006, followed by a few commits over the couple next years, then basically untouched for the last decade or so.
Just from the first 10 commits message and the author's name. I can tell you this was hacked-in grossly over a few weeks, because everything that guy did was hacked in quickly and filled with security vulnerabilities (he worked under very limited time constraints). Yet the software does the few things it was meant to do very well, so much so that it lasted this long mostly untouched and was built upon. The history shows the critical core code has less than one patch a year, mostly fixing trivial matters like a new path or syntax change from the language. I bet the project will be moderately easy to patch and upgrade because work from that guy at that time is usually decently organized and limited in scope to the essentials.
Lesson is. If you work in a long term industry (aerospace, defense, finance, healthcare) and long term projects, it's very likely that the product will still be running a decade from now and it is expected that future developers will be reading the history of commits.
Also, pointing to a JIRA issue, etc. if your company uses that is a nice way to complete the loop, the one link that explicitly ties code to planning.
At some point things are likely to change and the issue tracker that was being used will be changed. And now all those commit messages are useless because the underlying issue tracker is gone.
I'm not saying don't put a link to the issue tracker, but please make the commit message self-sufficient.
> I'm sure someone will say "but I use the history ALL THE TIME to source dive and paragraphs of context are super helpful". This is not the case for 95% of developers or projects so I can't really endorse spending time learning this "best practice".
> It's fine to be aspirational, but it's such a shame if people see posts like this and think they are failures or "bad" developers or that this is a widespread practice.
I'm that person who's often digging into the history.
It's often to understand other things that the people writing poor messages were sloppy about.
So sure, focus more on writing good code than writing good messages but the truth is that the context of the code and the feature will change over time so preserving the original context you had in your head when you wrote it is super valuable when the next poor sap has to come along a year later and change it.
I mostly work on relatively short term projects these days rather than long lived products and so my commit messages tend to be single sentences, with an extra paragraph if it was particularly involved. It takes little time and once in a while will be useful.
I probably lean more towards them not mattering that much relative to what anyone who talks about "best practice" would like to see. That said, I picked up a codebase from another team the other day where one commit message was simply a ";)" and the rest not much more in depth. That did make me wince a bit.
You're right that 99% of the time they will never get looked at again. And that's fine, but when deciding how much time to spend on your PR they're very helpful.
There are a lot of things in software that need to be possible, but you don't need (or want) the entire team doing it all the time. A team made entirely of 'people like me' would be a disaster. But a team without anyone like me will die a slow and lingering death and not even know why.
It's important when thinking of group behavior - be it work or personal - to differentiate between "I don't need this." and "nobody should have this."
The only method that consistently results in atomic commits is to squash commits when merging, in which case the commit message is your PR title/contents. The nice side effect is that these messages should absolutely be readable since your audience was your code reviewers.
Perhaps the solution is to train developers to look at commit messages then? Source control is amazingly powerful and nearly universally used at this point. Why not make the most of this?
[1] https://vincenttunru.com/img/Commits-are-documentation/annot...
[2] https://vincenttunru.com/Spend-effort-on-your-Git-commits/
Myself I use vim and the vim-fugitive plugin. Being able to do ":Gblame" is amazing. And then being able to step back in time to see what the code looked like before this commit is great too.
As someone who manages many codebases this combined with thoughtful squashed commits allows me to (1) understand at a high level what each commit does, (2) drill into each commit should I choose to in order to review what was done in more detail, and (3) review a JIRA ticket relevant to the change (which often times is more valuable than 1 and 2).
It's great to recommend and we should keep coaching and mentoring but ultimately this is an area where I think you'd be better to go up the chain of command until you reach a benevolent dictator who can mandate criteria. I'd love to convince people based on the merits alone, but to paraphrase a common observation in academia "the resistance is high because the stakes are so low".
I usually think a commit message should have a "what" changed and a "why" it changed portion.
Of course putting this information in code comments may also be a wise thing to do.
Now I don't know why don't you put one! I vaguely to remember reading this on the last "how to write good commit messages" blog and thinking it seemed fair, but now I can't think of a reason. Why follow grammar rules for the start of a sentence, but not the end?
edit: Another blog has a simple reason https://chris.beams.io/posts/git-commit/
> 4. Do not end the subject line with a period
> Trailing punctuation is unnecessary in subject lines. Besides, space is precious when you’re trying to keep them to 50 chars or less.
When I think of a patch it's almost always a patch that will "do the thing" rather than "did the thing".
In fact, and this is probably hard to believe, I don't think there were any comments anywhere in the commit history. Fun times.
Along the way I looked at some plugins for Pycharm and came across this commit templating plugin that adds easy workflow for adding scope and commit type. https://plugins.jetbrains.com/plugin/9861-git-commit-templat...
Does anyone use commit message templating or style like this?
The branches ideally gets squashed with proper commit message when merging so git history is clean. And no time is wasted writing commit messages that often say whatever the code should already say.
Last year it took me around 3 months of internal discussions between us devs and the project manager to get us to use a PR template in Github (The PM was more or less sold on the idea).
Commit messages (Not mine, obviously) rarely amount to more than "Fixed <this>".
And don't get me started on code comments (Non-existent, obviously).
The other coder is responsible for around 90% of the new code we ship these days (The product is "frozen" because there is a better product developed by a different and larger team, but we still have clients on this one) and I also had to strong arm him into adding type hints for vars and functions (In PHP) because he is very fond of patterns (Think, SomeObjectFactoryFactoryFactory) so it would make development with PhpStorm easier. So now the code is polluted with
/**
* var $sometint integer
*/
that explain absolutely nothing. * write a short title, starting with a verb
* include additional information from most important to least important
It just doesn't matter if you write "Adds" or "Add" or "Added", use capital or lower case letters, if you use full-stops, if you include tags (like BUGFIX or FEATURE) or issue numbers. Obviously, if you or your company decides to require anything specific, that's fine, but they are just details with respect to "good commit messages".And extra info is not very useful if it isn't well-organized. In particular, a narrative is not suitable for technical writing.
Personally, though I was in the "comments are running code and should never be used" camp for a long time, these days I'm more inclined to want a concise and clear comment on a particular thorny bit of code than a commit message
Tying a PR to an issue ticket where the conversations actually happened to come to consensus on how a feature would work or what tradeoffs were made is also helpful
A commit message in and of itself just doesn't have the bandwidth to give much useful context, and now with all the squashes happening, hard to see what commit did what
Also: Rob Pike's Google internal rant on the very same topic https://groups.google.com/forum/#!topic/golang-dev/6M4dmZWpF...
I just want to easily be able to roll your changes back if there's a problem. That's all. With good tests, git bisect will find the issue regardless of how fancy and well written your commit messages are, or how squashed or un-squashed your commits are.
So I think there's a tooling reason for squashing - more effective use of `git bisect`.
#issueid - summary of issue
- additional details
- more details
Without a blank line anywhere.
This makes it really easy to review in graphical git clients because the hyphens collapse into a single line like this:
#issueid - summary of issue - additional details - more details
Putting a blank line makes this rollup not happen.
Updating the foobar threshold Updating the foobar threshold to improve performance.