Commit Messages Don’t Matter
matt-rickard.com
matt-rickard.com
Correct, those don't matter much. On the other hand, describing what the change is about matters a lot.
It is 10x easier to review a merge request when you can see discrete steps that make sense (ex: update method X, update route Y that calls X, update all code paths that use route Y, update client to handle new data) vs one big jumble of "stuff", "fix it", "update getCustomers" that touches 27 places at once.
Breaking this rule is one of my bad habits and it's annoying for my colleagues to review and increases the risk of including some unintentional change ("I'll just edit this config file for some local testing...") as it gets lost in the noise.
I make the commit messages also for myself, how many time you find some code and wonder why is it like that, then you search git and find if lucky a nice message with a ticket id and if you are lucky again the ticket has clear details on what was the issue that was fixed and what behavior was requested.
Of course you still need an overarching description in the merge request - should be really easy to write, when you have all the steps taken right in front of you.
For a more concrete example, say you're adding a new screen to a mobile app, I'll probably split into chunks like these:
- add new `/foo` endpoints to API
- add `NewScreen` (can be reviewed as a feature, does the screen contain / do what it's suppose to)
- add `NewScreen` to navigation stack
- link to `NewScreen` from settings page
- add tracking for `NewScreen`
Each of these have very different purposes, and I find being able to focus on the different aspects (business logic, infrastructure stuff, analytics) helps a lot. They can also be reverted cleanly if needed without destroying all of the work.You can still review the entire PR at once, with the option to drill down into individual commits if desired.
Bug tracker? It can change. Reviews discussion? The tool can change. You get the point.
As an added pain point, I don't want to juggle different tools to get the whole picture about changes. I don't want to copy/paste a Jira ticket reference in my browser from my terminal to get the reasoning about the change. I don't want to copy/paste a bugtracker ID to get the reasoning about the change.
Commit messages are here for a reason, and it is to document the change. I find them invaluable when digging through a codebase with plain git tools, tig, or the git integration of an IDE. Conversely, this is why I hate squashes of PRs of more than X commits. Just no.
Anything occasionally gets lost. (some languages have surprisingly useful decompiles..) It's probably best if all approaches are taken in parallel. And taken reasonably well (as opposed to gold plating), enough to be workable alone in a pinch but decidedly not fully redundant. Comments, commits and out-of-band tickets all have their place and neither is a good substitute for the others. Allow them to complement each other, and strive for just enough overlap to make losses somewhat tolerable.
I like them because of the linear history, but the commit message should keep the explanation for _all_ the changes.
Speaking for myself, I also find that writing a good commit message is useful in helping me think about what work I've just completed and whether the changes I'm committing form a logical group.
Granted, there are all kinds of workflows and I can vaguely imagine some in which commit messages are redundant, but still, I don't think it makes sense to generalize the statement.
I don't want to see the local commit history of individual developers in the commit history of any shared branches (be it mainline branches or shared development branches).
What I do want to see is a history of pull request merges. I don't care how they are formatted (and this is often done automatically) but it matters that if I want to check out a historical changeset that I can. This is particularly useful when there are regression bugs that managed to slip through the QA and automated testing processes and it is unclear what caused them and when. It is nice to be able to go through the commit log and identify suspect work items, check out the commits during that time and see what they changed.
How you format the commit messages isn't that import as long as your pull-requests are squashed-merged (to suppress local dev workflow commits that aren't relevant to the team) and the commit messages facilitate doing that with something that links back to a work item.
Saved me hours of time and headaches. This is my anecdotal example from the last month for why this article is wrong.
Write good commit messages, people can and do read them. When it matters, it matters a lot.
The reality is that developers convey intent through pull requests and rarely go through a commit log.
Most people use commits as "ctrl s", which means commit messages are just some dumb shit in the way of me backing up my files.
Sorry but I feel like you have to be just in denial then.
> But even if that were the case, I don't think the majority of commit messages being bad would mean they don't matter.
OK but I think if it's the case it's fair to say hte majority of developers don't agree, whereas you stated it was an industry wide opinion that they are. I think commit messages are mostly stupid and the git UX is absolute trush. Theoretically, if the git UX weren't trash, they could matter, but it is so they don't.
I never claimed to be stating an industry-widr opinion. I was disagreeing based on my own personal experience. Based on upvotes and other comments, it seems I am not alone in this opinion.
Note I'm not arguing for that as a prescriptive state. I think a good commit log is a good thing and can see that there are different ways of getting there.
If comments are often considered harmful because of their lacking quality, why would commit messages be considered helpful when they suffer from similar issues?
Anyway, software development as a whole decided on a pretty stable set of comments that are considered helpful. Commit messages are clearly in there, as well as API elements description.
I have no idea what you're talking about.
Yeah the Linux Kernel tree is _littered_ with those, it's a real pain /s. To be more serious, any workflow that's a tiny bit serious about technical debt will plain reject a PR that's littered with those. I have no problem with rejecting a PR with obvious commits like those. In a professional setting it's vital to keep codebases manageable, especially if those are in the tens.
> The reality is that developers convey intent through pull requests and rarely go through a commit log.
Not really. Pull request are usually a 1:1 mapping with an user story / ticket id / bug id, etc. They might require multiple changes in multiple places with their own reasoning.
They really don’t.
That doesn’t mean they convey it through commit messages either, mind.
But the latter is definitely the worse sin. Having the information in the commit message is much easier to access (especially with a good editor), and it’s way less likely to be lost.
PR discussions are often a trawling mess of historical baggage, if something has been discussed back and forth for 6 months I’m generally a lot more interested in a final summary of the whats and the whys, than having to go through the entire thing every time.
That’s why the NTSB produces reports, they don’t just publish thousands of hours of raw investigation footwork.
Please don't. As much as it is useful to fix buggy code at a later point, it is the same useful to prevent introducing buggy code at a later point!
Never saw some code that looked funny at first sight, and that you considered 'fixing', then checked the commit messages and understood why the code was written the way it is? Fixing it, would have meant introducing a bug.
My firm went through 3 different ticket systems in 5 years, including a round trip to Jira. Teams/apps migrated to different Jira projects, repos got split apart/merged, tickets were archived, etc.
So the commit blocking requirement of having a link to the story ends up pointing to dead ends or stuff you may not have access to quite often.
This is the kind of stuff that causes senior devs to commit less often, or to have an always-open "misc" ticket to book everything to, with a copy-paste ready commit message in the proper format to pass checks.
Also... why would a company go through 3 different ticket systems in 5 years? That must cost a lot and definitely is not gonna help convincing people to use it...
They're a bad place for documenting the repository for users. They're about the sole place to document what's in the _commit_...
> and a common source of bikeshedding.
You can bikeshed anything if you like. But I'll grant the author's claim that bikeshedding the exact format convention of commit messages is excessive. If someone wrote "Added feature X" as opposed to "Add feature X" - yeah, that doesn't really matter.
On the other hand, if the first line of a commit message is extremely broad and vague, this makes it difficult for the developer browsing the commit history in a one-line-per-commit view.
> Adding in information about what files have changed doesn't make sense
It can makes sense if it's essential to the commit, e.g. "Moved `foo` from `file1.h` to `file2.h`". But TBH - this is bikeshedding.
> this information is more easily stored by looking at the diff
Ah, but the whole point is that you want to _avoid_ looking at the diffs of many commits.
> Information about why and how often have better homes.
Nope.
> There's the code review history, the bug tracker
These are not part of the VCS commit history, and may not even exist at all.
> the code comments
So, we want to put commit comments as actual comments now? No.
> and the actual documentation.
No. The actual documentation is:
1. Intended for users of the repository rather than developers/maintainers of the repository.
2. Documents all commits up some release (or up to some commit).
> Your commit message essays will probably be lost or munged anyway – squashed, merge committed, or rebased.
So will our code... but the point is that, after such squashes etc., there is still a commit history for which it is important and useful to have commit messages.
You've neglected to state what commit messages do matter, unless you think none do.
In my experience, commit messages only don't matter if the code is never used beyond the first creation.
Once you have to maintain the code, commit messages absolutely matter.
Does it need to include the granular details mentioned in the article? No. But commit messages help determine what commits should be reviewed during maintenance (including adding new functionality).
(I also experienced a touch of cosmic fear while reading the article, thinking about someone who might rebase or squash every commit, leaving only one commit for an entire repo.)
The author has a seemingly naive notion of the lifecycle code; commit messages matter a lot less if you have deep thorough documentation in the general case.
But obviously not all code is being written in those types of environments. In fact I’d venture the vast majority of commits have no associated code review thread.
Sometimes my little project I’m working on myself , which has no code reviews, eventually gets grows and has others involved; but the early history remains.
* more changes < feature/PD456
* more changes
* more changes
* changes < origin/feature/PD456
A week later: "what did I do here? Let me check the source code" (10 minutes of checking changes trying to understand what it does)No thanks, I prefer to at least write some minimal indications like
* cleanup and tweaks < feature/PD456
* implemented feature, new test pass
* added new test (now fails)
* updated version to 1.2.3 < origin/feature/PD456Better commit messages would be more helpful with PR review process if it was easier to review commit by commit which is possible but rather cumbersome.
Sure, there is also code comments, but have a good PR message preps you to read the code., And that good intro comes from commit messages.
There'll be cases where the following don't match, so let me explain what this is based on:
- In our flow, a PR is small and is the unit of work that is guaranteed to compile/work (tested against master before merge)
- Each individual commit inside the PR can be helpful for a reviewer, but the code should stand on its own and your commits (inside the PR) should not be a source of documentation
- I have never been able to pick out a commit at random from non-squashed PRs, and be sure it will compile. Unless you're running CI locally before committing, you simply do not know
- Spending a ton of time to create a nice history inside a PR is a waste of time, since throughout review or experimentation you will need to adjust work that you previously committed as if you had everything thought out - the alternative is aggressive history rewriting and force pushing which is very quickly a huge mess
- Finally, it tremendously lowers the bar of what's required from each developer, which scales much better, when all you have to do is take extra care in the final title and description
Now, most importantly, squashing PRs is tremendously helpful when digging around the code and having something like Gitlens explaining the changes:
- Commits from a merge commit are on average a lie of the final reason, and very often something like adjusting formatting, fixing linting, which makes the original reason disappear as this was the latest change
- Viewing the squashed PR title and description gives me so much more information and if I want to dive into the specifics, the PR itself is worth so much more than any individual commit message will be
Again, this is something that fits with a specific workflow. If you're developing directly against master or similar styles, then obviously this won't suit you.
---
The review and history part is actually something that is interesting to dive into for the importance of commit messages and what you and your teams expectations are. I'd recommend explicitly talking about this with your team :)
The funny thing is that I've been through many of these systems in the past years, and each time we mostly lost all the previous PRs and comments, but all my commit messages are still there preserved in the history.
These ideas are all a "better place" until it's suddenly all gone because it's not actually a part of your code history and can't easily be moved.
As we've transitioned to git, we don't have a universal culture of writing commit messages, and I wish we did.
Without messages, commits are only identifiable by arbitrary and meaningless strings of letters and digits. This would make git bisect needlessly more difficult, if nothing else.
It's actually quite good and if one takes 30s to write a little more descriptive message of what has been done, it'll help the next person.
I suppose it depends on your use case, one example is that I use these titles written by myself and co-workers to write beta release notes each week, which folks seem to really appreciate.
I tend not to worry about that and generally just squash the commits when I merge, so that, on main, there's a single commit for a particular feature. That's an overall piece of work I did. If someone wants to dig into a feature commit by commit, they can typically check the PR out and do it that way, but in practice, I've rarely seen that happen, and if it did, it was more for the overall feature rather than any one commit.
So, overall, it's not that commit messages don't matter, but rather that, while a pristine history is always nice to see, it's not always possible without unnecessary head scratching that generally just costs more dev time for very little benefit.
> If someone wants to dig into a feature commit by commit, they can typically check the PR out and do it that way, but in practice, I've rarely seen that happen
Of course not, it’s a pain in the ass having to go back to the PR and trawl through a commits sequence you describe as essentially broken.
Github, sadly, very much is not designed for that.
The only instance where it paid off was to include the jira task code in the second line of the commit, but this also has been short sighted and rarely useful because most discussion about PRs happens on slack and teams, not on Jira.
Yes, because commit messages do matter and rebasing/squashing is how you clean up your initial draft commits. (Merging doesn't lose or munge the commits.)
If there is, the PR discussion might be the right place for explanation, but if not, the commit message is still valuable.
Assuming people are religious about putting the PR ID in the commit message, you might be able to relatively easily jump from the commit you found normally (e.g. blame-diving or via the log) to the PR. Whatever your tool might even be able to do the jump in one shot, though I think it jetbrains stuff it usually takes two (“open in > Github” from the commit, and from there click the reference to the PR).
But obviously that means you did go through the commit message, and having had a good summary in the commit message would likely have saved you the trouble.
Commit messages are like timed notes, or artifacts opportunities to explain the path of evolution of a code base, while explicitly pointing at it.
To increase concision to the point where bitrot is unlikely loses useful contemporaneous detail.
So, I prefer concise comments that are highly salient and unlikely to turn into misinformation, usually about key details on a mechanism where subtle mistake is more likely than the norm, and detailed commit messages on contemporary justification.
Also, the author mentions the commit messages being "probably lost." I'm not sure what the author is doing to get themselves into that situation.
I work on code bases with good commit messages. I use emacs's vc-annotate with regularity.