How to Squash and Rebase in Git
jenweber.dev
jenweber.dev
https://fle.github.io/git-tip-keep-your-branch-clean-with-fi...
There are two exceptions.
1. branch developer suffered from several "thinkos" along the way, the branch doesn't contain that many changes, and there's simply no benefit to seeing the contrast between the initial (mistaken) changes and the final result.
2. the branch commits create a situation where git bisect can't be properly used because intermediate commits will not compile. Squash is one option here, though it is often preferable to just re-branch and re-work commits to avoid both the bisect problem and the squash.
With a 21 year development history, we have found it invaluable to be able to trace the "thought processes" behind a series of commits, and squash would rob us of that.
Similarly, I don't really care how many times you merged main in the course of developing on a branch. Those commits are just noise to me.
As long as only a single developer is working on a branch at once, I'm very comfortable with them treating commits as a work in progress, then doing an interactive rebase before merging to main to clean things up.
Even in the case where multiple devs are working on a feature branch, as long as they're coordinating properly (and using safeguards such as `git push --force-with-lease`), I find more value in having tidy, organized commits at merge time than a full history of the development process.
However I typically squash by just doing a soft `git reset` to the merge-base (the last commit in common with my branch and the one I'm building my work from) and making a new commit. It avoids touching the working tree at all, which avoids breaking my build cache and confusing my IDE.
If I instead want to do a squash-rebase onto the _latest_ of the upstream branch, I do the same as the above, except I `git merge` first (to resolve any merge conflicts in one shot.) Then the merge-base is just the upstream head. This is way easier than doing `rebase -i` and squashing, because you don't have to fix the merge conflicts over and over again for every commit you're trying to squash... just fix the merge conflicts once during the merge, and the "reset and commit" trick is going to create one clean commit in the end anyway.
The commands will look like that if you want to just squash your changes: git reset --soft $(git merge-base HEAD origin/master) git commit -m 'squashed commit` and like this if you want to squash and rebase on "upstream" (most likely "master" on "origin" remote), andrewmackrodt already wrote it in another comment: git fetch git merge origin/main git reset --soft origin/main git commit -m 'HN-1337 Awesome new feature'
I like this approach (haven't tested it yet!) because as author wrote it doesn't unnecessary change much files in you current working directory, and IDEs sometimes get crazy when you do rebase.
Currently, I usually try to write my individual commits to be logical changes grouped together and if I see there are some WIP commits or commits that don't look like one logical change, then I just squash this commits (not whole branch). This helps a little in avoiding unnecessary merge conflicts in rebase, but it's not always perfect solution.
I still need to figure out what this "git rerere" magic does though...
I also like the “reset and recommit” approach to squashing in place. It greatly benefits from being easier to explain to coworkers who are less experienced with Git. Reset and especially commit are common operations. Interactive rebase and the text editor that pops up and requires manual editing can be overwhelming for less experienced team members.
For anyone wondering what that command would look like:
git fetch
git merge origin/main
git reset --soft origin/main
git commit -m 'HN-1337 Awesome new feature'
Last year, I finally merged an 18 month "long" dev branch that I had kept up to date with rebasing many times during work on that branch. Why 18 months? Some things just take a long time. git rerere made this easy and almost painless.
Also it'll be harder to go back and review changes when troubleshooting issues because formatting changes were squashed together with business logic changes. With merges at least there are still original commits available, and it'll be possible to either rollback everything or a single commit.
One argument, and which I hear often, is that they want to keep the git history clean.
The other argument is that keeping things rebased & squashed allows you to do reverts easily.
I personally don't mind the merge commits cluttering the history, and especially when working on a longer lived branch I will always prefer merging-in changes rather than constantly rebasing and fixing same conflicts ad nauseam.
git revert -m 1 $commit
You can also visualize a clean history with merge commits.
git log —-first-parent
The only argument I’ve heard that holds weight is that GitHub’s UI squashes the merge commit history into a linear view.
In my experience, devs rarely pay attention to this, so commit messages end up as a big list of:
* Ticket-1234 feature
* fix
* cleanups
* fix
* another fix
* now the real fix
Which I wouldn't qualify as a contributor to a "clean" history. It'd be great if GitHub's UX around crafting the commit messages would be more considerate and foster more meaningful commit messages.(edit: formatting)
There’s no way to automate version control.
What’s more, I also want only commits that relate to the feature implemented on the branch. Preferably no merges from the parent branch.
To get zero oops commits requires squashing in the branch. Note that this is independent of whether the branch is merged or squashed when the PR is completed. Squashing 10 branch commits to 3 is the important step. If those 3 are squashed to a single on the target branch is less important and is a tradeoff.
As a developer I’ll not care to maintain build/testable code on feature branches, which is my main reason for usually squashing to the target branch. For larger changes I merge (after cleaning up and interactively rebasing), but git bisect will have to use —first-parent to work.
+100. Though every time this comes up on HN you would be surprised by how much some people just vehemently disagree.
I will also add another requirement - all unit tests should pass at each commit in the PR. That way, you can use the rebase strategy to merge and you would still be able to bisect.
That's a fair point, but I would say that if you encourage your devs to make each commit logical, meaningful and self contained, this is worth the extra cost.
Also, you don't need to run ALL unit tests - just the ones affected by commit you are testing. This of course requires a build system with good dependency analysis.
What's more important is that the commit history of the main branch remain sane. That's why you might want to squash-merge, since it mostly solves the issue of having too many extraneously-named commits in a place where one indeed may want to examine individual commits.
Honestly, I couldn't give less of a crap what sort of commits someone has on their branch as long as it's not so many that I can't possibly review all of the code.
Don't get me wrong, because I think it's great if you want to maintain that sort of standard in your own work. I don't find it particularly reasonable to expect this of anyone else because it's more of an aesthetic choice than a practical one, not that either are exclusive of the other in this case.
> Preferably no merges from the parent branch.
Yes. Always rebase. Rebase frequently.
I aliased `diff` to open up my browser showing me the full diff of what I'm working on - super useful!
Git history management can be tricky business, and most teams have their own norms, as this discussion shows. What matters most of all is that everyone on your team learns and feels supported with your chosen git processes, including junior devs, and that you have some consensus on it.
For anyone who is overwhelmed by the options and syntax, one of my friends swears by Fork, for those who prefer a GUI over doing git by hand. It has its own learning curve, but works better for some devs.
Fundamentally, this is my fault - I don’t have a good mental model of git, and even when I do have a solid understanding of what I’m attempting to achieve the actual repos/situations I find myself in rarely line up with the examples to any useful degree.
The way I teach programmers to do this is to make use of cheap branches and to instead work on a temporary branch, and eventually use ‘git reset --hard branchName’ to turn your original branch into a mirror of the temporary one when ready.
This is a helpful mindset to get into; you can just throw it away in a worst case.
Those “backup” branches really aren’t that for me. They’re just named commits which makes it easier to figure out where I was. But they are sometimes necessary since the reflog isn’t always nice and easy to read.
And pushing that branch to a remote? That’s unnecessary unless you want to delete the whole repo and clone it again in case you “mess it up”. Well, I guess that could be useful in a tutorial.
`git rebase --abort` can be helpful too, if you run into some unexpected merge conflicts along the way.
I’m pretty sure it just creates a single commit on the target branch rather than a merge commit.
Work on a temporary branch, get stuff all working there, keep merging changes elsewhere into the temporary branch to stay up to date with your PR target branch.
Push the temporary branch early and often and don't worry too much about its commit messages (will never be part of official history) except as they pertain to your understanding of the evolution.
Then create a new branch (for PR) off the one you'll eventually PR to, and `merge --squash` your temporary branch into the branch for PR with an awesome commit message. Submit the PR.
Kill your temporary branch when no longer useful.
Modify the strategy a bit if you want to break up the PR into multiple coarse commits ... 100 temporary-branch commits, 3 new-material commits in for-PR branch, and 4 back-merges of PR target branch to for-PR branch into temporary branch. This keeps a full context of your changes vs upstream changes.
Not really. The secret is to rebase fairly often, at the very least daily (depending on the size of the organization you work in, this could be more often). I often see complaints like this from junior devs. They work in isolation on their branch for a week, and then they complain about how hard it is when they try to do a rebase on top of the 100+ commits that happened while they were out there doing their own thing.
Another common issue comes from not understanding how rebase works well enough. You are supposed to adjust just the changes of this commit, not take anything else into account. If you start making unrelated changes when rebasing, your commits will soon conflict with each other, which means you'll run into a lot of very confusing conflicts.
It's fine to squash before a rebase if you are not interested in the history of your branch. But if you are, it's just not an acceptable solution.