GPT based tool that writes the commit message for you
github.com
github.com
A summary of what the change does is still useful, and I'm glad this can now be automated, but it's not the main benefit of a well written commit message.
Starting point. After that it should up to the committer to make modifications and upstream the message. For those cases, seems good enough for me.
What you're asking for would result in extremely bloated messages.
No. It's best to assume that all information in your ticketing system will be totally lost.
The commit message should have the "what/why" directly in them, because that's where you'll first look when you want to find that out (why add an extra hop), and it would require extreme incompetence (which is not unknown) to lose information stored there.
Integrating with another tool is fine, but not at the expense of your commit history!
This was learned from hard experience.
During my tenure, we've migrated corporate wikis/"collaboration systems" four times, and we've used at least as many ticket systems. We also migrated from webex to Teams a few years ago. Each migration resulted in at best a severe partial loss, and in some cases a total loss. The most stable resource is the codebase.
Basically, whenever anyone does a migration for you, they will lose things, and they might even think losing them is a "clever" shortcut that will make them more efficient. Our recent documentation system migration only covered documents modified in the last year as a corporate cost-saving measure, and we were on our own for everything else. Weak/understaffed teams did a bad job. They retroactively implemented a retention policy on Webex that wiped out a lot of valuable recordings, and then the coup-de-grace was the migration to MS Teams, which wiped almost everything else (basically everything that the original recorder didn't conscientiously download and save). We recently migrated to MS ADO as our work tracking system, and that was the only time I've seen old stories/bugs migrated at all, and even then we still lost all embedded images.
Here's a list of things in order of least to most likely to be lost:
1. Head/tip of your project in source control
2. Source history
3. Documentation stored in a central location as files you can easily copy
4. Documentation stored in some system (like a wiki)
5. Tickets/stories in your work tracking system
Please read a few commit messages for the git project itself.
That is how git was designed to be used. Maybe that provides some perspective of just how bloated commit messages can be, but also how useful they then become.
There is a great blog post about it: https://dev.to/jacobherrington/how-to-write-useful-commit-me...
It boils down to the following commit message template:
<what you changed>
because <why it was needed>Most developers seem to be poor "commit historians" by default, I'll admit. If you have the energy and consistency to enforce good history, it eventually becomes self-sustaining.
(Those are not the only situations, of course.)
My experience here varies wildly; I’ve worked on projects where commit history was carefully maintained and documented, as well as on projects with random squashing or random commit messages. Both had their own benefits and came with a cost too.
For larger, more complex or more mature projects, going through git history can be a common task. So here I’d prefer meaningful commit messages; git blame/bisect/etc gets easier, and sometimes any additional piece of docs has immense value. For smaller, simpler projects or prototypes, where history isn’t used that often, I don’t care that much.
Source backup is one of my primary uses for SCC, and when I step out for lunch, or need to knock off for the day, I commit where I am, and push with a "wip" (work in progress) message.
The feature branch name I merge is probably the most useful. It's an aggregate of all my commits and what I'm working on.
In fact, even the hairiest long-lived branches can get cherry-picked apart into several "layers" of branches to get reviewed and merged piecemeal. It's really easy if you're starting from too many commits. It gets harder (not impossible) if individual commits commingle said layers.
I find the latter quite common.
Hopefully as a side benefit others also find them helpful, but my present self is always grateful to past-self's conscientiousness!
Out of three (commits, comments, docs), the commits are least approachable, just few changes in the code will erase original commit in the blame.
By all means write good commit messages, but crystallize the info in them into changelog or other docs every bigger release
It's all in addition to the changelog and other documentation if following a defined release methodology.
Personally I like having the "why" in a comment (with a pointer to the bug tracker item) if adding code. That doesn't work as well when deleting code, and for that, the commit message works as well as anything else.
If a what looks like this: Store user full name in session table
The why can look something like this:
The full name is authoritatively stored in the user database and should be looked in using the user name. However, during deploys of the Raticate system, the cache gets invalidated and we need to quickly refresh it. This leads to a thundering herd system, as in incident 5244. By caching the full name in the session table we sidestep this problem entirely.
The reason why the Redis cache was not chosen for this data is because the session object is handed to the checkout system for processing, which can for security reasons, as described in document 7745, not access Redis.
* What problem were you solving?
* Why this way and not other ways you might have solved the problem?
Or do you think those aren't usually needed for a very good commit message?
That the ticket was open
> * Why this way and not other ways you might have solved the problem?
This is first one that worked.
There, that's those answers covered for like 80% of the developers
The “why” goes into the PR and more importantly, engineering documentation and inline comments.
Plus, a commit message seems woefully plain and too short to properly communicate these things.
The beauty is that this isn’t prescribed. We can all do what makes sense for our contexts.
This just ensures that the “why” is lost when someone comes looking years later.
From experience, SCM history is far more durable than just about any other work product we produce.
Over five decades later, commit history was still available for the Unix sources and could be reconstructed: https://github.com/dspinellis/unix-history-repo
I’ve used 35-year-old commit messages to help understand a long-standing issue, decades after all other related organization tooling and data had disappeared.
This kind of stupid.
1. At best, all this can do is look at what you actually committed, but that's not going to tell you the most useful stuff that should be included in a commit message, like why you're making the change in the first place and what your intent was (which may deviate from what you actually did).
2. If you wrote the damn commits, you should be able to describe them yourself. If you can't, something's wrong.
The only case where I can see a tool like this being useful, instead of a bad habit, is summarizing old commits, where the committer didn't put anything useful in the commit message.
It can't work, unless it's hooked up to a brain scanner and can read your thoughts.
It seems more like dumb thing that people who don't know what they're doing will think is smart.
I see a useful use case for using a tool like this after the commit messages are written (e.g. auto-generate these summaries on the fly when browsing history), but using it to write commit messages themselves is a terrible idea.
As such, no amount of “gpt magic” will solve this.
It’s correct often enough that you can get complacent.
Communication is part of the job. How many well-paid and creative professions get to just "nope" out if their responsibilities like that?
Also if an engineer has troubles describe change than maybe you don't want that change to end up in you repo.
My experience with this hitherto relates to terrible scaffolded/generated code. I’m scratching my head wondering “why did s/he do it this way?” only to realize the nonsense code was lazily churned out by some generator. I wish this stuff was labeled clearly.
Like, in this case, I would guess that the vast majority of commit messages are something like "foo", "fixed", "argh", "stupid CI thing try 10", etc. Pretty sure GPT can improve on those.
`npx add-gpt-summarizer@latest`
The CLI will walk you through all the setup steps. source code: https://github.com/soof-golan/add-gpt-summarizer
Writing good commit messages is thinking, not toil. 90% of the time is even easy, one line description, good to go. The other 10% is when you gotta do the thinking, though. Extracting the relevant changes from the incidental ones, summarizing the whys, the pros and cons, the weird interactions that lead to the specific choices made in the code. This is not toil, it's thinking. It's not even a matter of whether AI can do it well (for this 10%, I'm *very* skeptical it can), the issue is that this is the kind of thinking you should *want* to do, that you should want to get better at.
TL;DR: this feels like asking the computer to jog for you and expect to loose weight.