The merge commits everywhere argument falls apart if you use log --no-merges. The shitty commit messages argument is solved by not allowing shitty commit messages. The "fixed typo", "oops left TODO", "another typo" chain of commits argument is solved by telling people to not do a million commits like that and IMO those should be squashed. The clean history argument is solved when you don't allow the useless "fixing typo" commits in feature-xyz branches. Your history should be clean and contain a history of the progress that was made.
But ultimately the easiest way to solve all these problems is to force everyone to squash. You only have to police people on a single commit and people don't have to learn about options like --no-merges.
For those of us on maintenance teams, who actually have to dig in to the history to figure out what happened, not squashing matters a lot.
Other comments I made on another recent git post:
>One of those small typo fixes could introduce a bug and multiple commits makes it easier to track down. Blame would pinpoint it exactly - you'll see Bob's "Fix a typo" commit instead of being buried in Bob's "Add trucks to the game" commit. That saves you from having to go look at the PR or diff and to figure out why Bob renamed cares to cars.
Over the past few months we've been trying to formalize our git usage since we're split across a bunch of teams, and on this particular issue I've found a very hard split: People on the new-feature-development teams love squashing and squash-merge, people on the maintenance teams who have done work in git repos for a while are almost all against it (of the ones with no opinion, several are still only working in svn repos so haven't had a chance to form an opinion), but most people on the maintenance teams don't really speak up about it so our company-wide guidelines remain heavily in the squash-merge camp.
> don't allow the useless "fixing typo" commits
This sounds like a much much more heavy handed approach than: I don't care what you do on your feature branches, just squash your commits to master and write a nice commit message explaining what you did.
IMHO, all your rules do is increase inertia of people to fix typos.
If you're fixing typos in new code I see no reason to have the history of introducing the typo and then fixing it. If you've come across a typo in a file you're editing then by all means make that typo fix in its own commit.
It gets worse when you start fixing typos in files unrelated to your fix. One of those small typo fixes could introduce a bug and multiple commits makes it easier to track down. Blame would pinpoint it exactly - you'll see Bob's "Fix a typo" commit instead of being buried in Bob's "Add trucks to the game" commit. That saves you from having to go look at the PR or diff and to figure out why Bob renamed cares to cars.
That does not go into the PR and someone should make you split that out into a code-cleanup PR and it shouldn't pass review.
The first is to keep a log of what you are working on. For me, that's lots of small and dumb commits.
The second is to provide a story for review.
Most of the time, when the change you are putting out for review is small and simple, you can just put it all into a single commit.
Sometimes, your change is more complicated and it makes sense to break it into a series of related commits. Along the line of 'first do the refactor that makes the complicated change simple, then make the simple change'.
Your job as an author is to hide the ugly reality of how you actually came up with the change, and present the reviewer a sanitised view of reality. Reviewing code is hard, so it's best to make it as easy as possible.
I am doing a lot of exploratory programming and way point commits.
I don't need the reviewer to understand all the mistakes I made and bad designs I considered. It's enough work to understand the finished design.
As a rule of thumb, every commit that lands in master's history should build and pass tests. So that eg 'git bisect' works.
But it's not a good idea to put that same requirement on waypoint commits I make along the way, when exploring.
Doing hacky and messy WIP commits is totally fine. Just clean them up before having others read them.
In case you're not aware:
git rebase --interactive [remote/]ref- History is sacred and can't be changed. It typically goes with a merge-based workflow. The good thing is that nothing is lost, what you see is what really happened and you won't make a mess with inappropriate push --force. The downside is that it looks messy since you are going to see every typo and you will have some non functional code in your history. People using that workflow usually don't squash since squashing rewrites history.
- History is like documentation and have to be nice and clean. It typically goes with a rebase-based workflow. This is what the article is about. The good thing is that every commit has working code, and git log can effectively replace a more formal change log. The downside is that you lose information and the git log doesn't represent what really happened, furthermore, since push --force is often used, your local branch may not be the branch you think you are on, you may even end up destroying other people commits. People using this workflow usually squash to make their commits nicer.
I prefer the "history is sacred" workflow myself, but both options are valid.
It doesn't matter to me at all 3 months later to know that you initially had a typo in your first commit and then fixed it the subsequent commit. That should just be a single commit.
There are some very rare edge cases where having the whole messy history can offer some insight of how a certain mistake ended up being made, but IMO in 999/1000 cases that's not the case.
All to often in rebase+squash heavy repos I start going through the history and find the mega-commit that introduced a change and gain zero insight. In comparison, merge heavy repos have a lot of merge this, or fix this typo commits, that are fairly easy to skip right over when not pertaining to your actual change, but prove invaluable when the problem was introduced by some weird bad conflict resolution including a typo, or countless other problems where just all context is otherwise lost.
I think it’s unfair to characterize one group as not using the history, but instead represent differences in how groups use them. As an archaeologist vs daily change log perhaps is a better characterization?
When you look at the master branches you will mostly see merge commits, all the mess will be on the side, for example in feature branches. Merge commits can be clean, and when you are blaming, you will see the merge commit, not the dozen of typo commits. If you do "git log --first-parent" you will not even see the messy commits, but they will be there if you need them.
The history should be treated as the perfect object to develop. We should continuously revise how this program should have been developed to its current state, by what sequence of changes.
Git is essentially a blockchain, rewriting history is like rolling back a bitcoin transaction, it is impossible without breaking cryptography. Instead, when you rewrite history in git, you actually create a new project to replace the old one. You then have to tell everyone to switch to your "new project", otherwise you will create a lot of confusion and painful merges, if they don't decide to say "fuck you" and keep working with the original history.
Git is decentralized, and to work it needs some consensus, and the consensus is in the history, if you break that, you break the decentralized nature of git. Now, if your don't open your repository to the general public and use git like you would use svn or other centralized systems, it is fine, but don't do that for public repositories.
Upstream says: "master is now this SHA, deal with it".
Actively developing downstreams can easily rebase their stuff (if any) across non-fastforward changes, and life goes on.
For a pure consumer of a repo, it makes no difference.
There is only the cultural idea that it's a no-no, not a technical idea.
yet another typo
test fix
aaah, why is this test failing?
merge foobar/narf
revert foo
foo
Is just not a good history to preserve
i.e. the anti-rebase people are usually of the "never rebase" flavor. Because they're also the "history is sacred even though I never look at history because if I did I would get dizzy and throw up" crowd :P
Sometimes, the flow for a single feature A: - commit change AA - commit change AB - fix change in AA - extend AB
In this case, until merged into main/master and to reduce noise for the reviewer (especially if using a code review each commit type process), its best to squash all commits into a single "feature A".