Fortunately, I don't squash my commits
blog.ploeh.dk
blog.ploeh.dk
The problematic commit is described "Extract CreateTokenValidationParameters method", without an explanation of why the refactoring is necessary or what problem is it solving. It looks like it is improving code readability, but it doesn't go as far as fixing the more glaring issue with global variables. Other commit messages seem to follow the same pattern.
In other words, the commit:
- has minor code readability improvements
- contains no useful message for the future programmer (why is it needed?)
- provides no new functional feature/improvement/business value
- is later found to contain a bug
I find squash/rebase/cherry-pick useful when reviewing my work and deciding what should go in the current pull-request. For example, a refactoring might be postponed for later if it is deemed too time-consuming or irrelevant to the current PR. Or, one can squash logically related commits together, add a useful message and merge them separately. The resulting commit log will still be bisect-able.
For beginners: there's a very good article with tips for Git Commit Messages[0] that helped me have git histories I enjoy reading.
He claims to be an expert on dependency injection with two decades of automated testing experience. I wonder what was so hard about writing a test to cover this scenario?
Given that this issue is within ASP.NET that's not likely to happen in one commit of this author's project.
^: https://docs.microsoft.com/en-us/dotnet/api/system.identitym...
It takes someone that's, at the same time, making 'messy' commits with horrible messages, but also diligent enough to never commit any breaking changes.
If I rollout a rewrite of an endpoint and run into a weird issue in the QA environment, I'm a simple git revert away from fixing the issue. If I had spread that endpoint across 25 commits, I'd have to actually debug the issue in QA and figure out what I broke. Reverting quickly lets me debug the issue on my time instead of keeping our test suite broken.
It's not always feasible, and I try not to be a stickler when people on my team don't do it but the fact is, if you're working in a world when you're delivery code quickly into real environments, having a back-out strategy is paramount.
No, you could just revert the entire lot in one go.
last_prod_tag < the_lot < current_prod_tagThis gets a bit more murkey with mono-repos, but even microservices can combine to create complex production issues.
That way one gets clean shared history while preserving local work history.
Why not just rely on a merge commit instead?
If there's a way that I don't know to show merge commits in blame rather than the actual source change commit, then I'd be all over it. Until then, single (whole) units of change per commit.
1. The large majority of PRs I've reviewed have a single contributor. Additional contributors are rare. When they do happen, they're often a minority contributor or simply consulting on a PR. It's net neutral when all PRs are squashed in the same pattern.
2. Even with multiple contributors, most features have one leader. It's much easier to talk to that person (and have them delegate) than it is to piece together multiple contributions.
Your use-case is actually the cause of a common rule I've seen at work of requiring a ticket reference in each commit message, which allows looking up the original ticket and associated PRs, along with any commentary & discussion at the time the commit was merged.
On a big code-archeological dig, I often follow a path like run blame -> look at the diff -> pull the ticket reference -> find ticket in issue tracker -> read its description & comments -> find linked PR #'s in the ticket tracker -> open PRs & read diffs and comments -> repeat for linked issues if needed (and then as often as not still end up baffled)
One team actually kept an old redmine VM instance running mostly based on my personal use long after we'd migrated to JIRA, so... I think my approach may be a little unusual! At the least, doing better sized commits would a huge step for every case involving blame.
It also defaults to causing the main to have a ton of commits with "Merged from XXXX branch" as the summary lines when that's not nearly descriptive enough to quickly find what type of commit I may be looking for.
Using merge commits doesn't have to mean that you don't squash at all. At one extreme, you include every single commit ever made on the branch, and at the other extreme, you squash the entire branch down to a single commit. Using merge commits, you can go for any option inbetween.
> It also defaults to causing the main to have a ton of commits with "Merged from XXXX branch" as the summary lines when that's not nearly descriptive enough to quickly find what type of commit I may be looking for.
Do you often read the log linearly? 99% of the time when I investigate history in Git, it's either through "git blame" or through "git log -S somestring" (search for commits that introduced or removed "somestring"). I rarely, if ever, just read the log as-is.
If I want to merge my hotfix topic branch into both the release and the master branch, their commits won't match so I can't check if it's present in both automatically.
If a topic branch is left up instead of deleted after a squash merge, I can't even see that master is ahead of it!
Squash is an ugly hack that creates as many problems as it solves.
I desperately wish git had a "group commits" feature that let me manage a cluster of related commits as a single commit for the purposes of history-viewing, reverting, and cherry-picking.
alexmingoia proposed:
> That way one gets clean shared history while preserving local work history.
But if you keep on building upon the squashed versions each time you do a merge, you won't have convenient access to the local work history any more.
If you wanted to have access to the local work history you'd need to keep each branch alive still, each time based on the squashed history plus your individual commits up until the next squash.
You have the local work history already, and that work is done to the point of it being squashed and pushed. Why would you need to keep referencing it to the point that it's inconvenient to have it in another branch?
> you'd need to keep each branch alive still, each time based on the squashed history plus your individual commits up until the next squash.
exactly what I do, and there are maybe just three or four branches I maintain for reference, very few squashed branches are useful a couple of weeks after they are released.
Maybe you can do something similar to that with git-replace?
Otherwise, you could just tag your commits by putting something in the commit message and use `git grep` to search for those commits. Then you could just build a small helper script that 1. greps for a tag 2. loops over the found commits 3. rewinds them or whatever and 4. squashes the resulting commits.
Dunno, there's a bunch of other ways depending on how exactly you want it. But I agree that grouping commits would be pretty cool :D
I usually deal with this by `git rebase`ing your feature branch on top of the shared branch as soon as the PR is merged. You sometimes still get merge conflicts with this approach, but they're always in the code you've just written so they're usually pretty easy to fix.
Isn't this what git does by default when you merge your changes? The merge commits group small commits together.
Merge commits work fine for most of that, you just have to adapt to merge UX/commands:
--first-parent (to git log, git annotate, etc) gives you clean history viewing of just your "groups" (your merge commits).
--first-parent even works for bisect allowing you start by figuring out which merge commit brought in a change (and then dig into the merge branch itself if needed as a second bisect).
You can revert or cherry pick merges if you provide the -m (mainline) flag to tell it which parent to consider the mainline (usually the first parent, but not always depending on your intended revert/cherry-pick; it complicates what you need to know about the revert/cherry-pick, but if you are in the process of revert/cherry-picking you should already be figuring out what your mainline is and expecting some possible complications).
I think sometimes the only "problem" with Merge commits is too few pretty UX tools default to a --first-parent view of the git graph and don't themselves provide good tool for picking that -m (mainline) for revert/cherry-pick.
But yes, generally history is much cleaner when squashing commits.
Of course, there are caveats for reverting a merge - finding the right parent adds another failure-prone step, and your team can get in all sorts of trouble if they try to work with the branch without reverting the revert.
As for cleanliness of history, everything should be as simple as possible, but not simpler. Squashing and rebasing is destroying history, which often could be valuable, as OP shows.
I'm firmly of the belief that any benefits that squashing brings would be better achieved with better tooling, rather than by re-writing history and throwing away potential debugging information. In this case, what you need is for git to make it easier to revert 25 commits in one go
(there are pros and cons of both approaches - this is a con of the git approach, I'm not knowledge about enough about esoteric details to comment on if git actually made a bad choice or just a compromise)
The problem mentioned in this thread is rolling back a feature that was merged. The only solution I know is navigating the log to find the first commit on the branch from which to revert. Don't forget there may have been several merges from and to master, as well as commits shared with other git branches that should not be reverted.
In a Mercurial branch, each commit is tagged with the branch name. Unamed branches, à la Git, are called "bookmarks".
The only step necessary to revert the changes from a merged branch is to revert the merge commit. See:
https://stackoverflow.com/questions/7099833/how-to-revert-a-...
https://github.com/git/git/blob/master/Documentation/howto/r...
If you revert one specific commit it just does another commit with the exact inverse of what happened in the reverted commit, so all other changes are preserved.
git revert -m 1 <commit>
Or something
I don’t think this is true /just/ for squashed commits as you can simply revert the merge. Squashing is also bad, because you’ll lose the history of the bug fix that prevented the deployment. Or at least someone will have a hell of a time deciphering the PR to reintegrate the original change and your bug fix.
After a few days, weeks, or months, the argument loses even more water because code will likely depend on the commit in question.
War story: there was a time someone accidentally deleted a multi-gb table in production. The table would take hours to delete and replicate globally, so the entire company spent at least an hour deleting the feature from production to stop the database errors. There wasn’t any reverting of original commits. No one had time for that.
It's better to always have merge commits. They can be reverted just as easily. I don't understand why merge commits aren't the default in git.
That's why Mercurial (esp. with evolve and topics) has forever my preference over git.
I can't figure out how to explain to Junior Devs (who have only ever known git), that they have a concept of a branch in their head that doesn't match the concept their tool of choice is giving them.
We talk about "branches" as logical sets of changes. We give them meaningful names, we construct the concept of Pull Requests and code reviews around the concept of a branch. We later refer to Feature X as having landed in master from branch Y. But git doesn't have any of those semantics. It has lots of ways of dealing with commits, and a facade of a branching model is just one more way of dealing with commits. Branches are not a first-class concept in git. And certainly not like they are in our minds.
However, git is amazing at what it does! And if I was running the world's most popular OS kernel development team and was expecting to receive hundreds of patches a day via email from developers in whom I have limited trust, I would definitely start with git's model and change the way my brain works to match its semantics.
Instead, I find myself on a small team of high-trust coworkers who all talk about branches as if they really exist in our git history, and somehow I'm the crazy one for pointing out that every time we hit a problem with this mismatch the fact that we're using git is the reason that we can't have nice things.
Git commit is always a standalone item that linked to parent commits. The actual content can be completely unrelated to parent commmits if you must do.(but what is the point of doing this?)
A merge commit is just a commit that has more than one parent.
Bisect is the killer feature of git, for me. Squashing releases takes that superpower away.
I really like making tiny, continuous commits as I work. It's a great flow. git-revert becomes a Ctrl-Z on steroids. I don't what to clutter up the "official" history, with all these tiny changes, many of which don't even compile. That breaks git-bissect and all kinds of other flows.
So the only option is to squash commits. But there's something deeply uncomfortable and unsettling about permanently re-writing history. Plus, it's nice to have a history of those working commits as an artifact. If I'm trying to unpack the reason that I did something 9 months ago, then seeing a replay of the code changes is super-useful.
I do see your point though.
Edit: the only reason to merge “main” into “history” is to enforce convergence.
I've been wishing for a git GUI that lets me drag hunks around such a stack so I don't have to keep moving them by hand.
??
git tag
??Surely this is why we have things like release tags and snapshots of previous versions?
Unpicking individual features is rarely simple even if you have got your git repo into an immaculate state, as things are interdependent.
This question always boils down to this metric. Which happen more often...
1. A circumstance arises where the dynamics of fine grained commits make the problem more obvious.
2. I have to interact with the git history
Now(and this is just me) I interact with the git history for one reason or another multiple times a day every day. I have been using git for 10+ years and I haven't yet encountered the first circumstance yet. Its not to say I won't encounter that circumstance and when I do I'll probably pine for the fine grained commits that would make it stand out.
For me the clean, easily browsable history benefits me every day. Furthermore because of how often I have to use it I should optimize it heavily at the expense of just about anything else unless the benefits of that item can be realized with similar frequency. The access to fine-grained commits hasn't helped me once in the 10 years I've been using git. With that calculus in mind I must conclude that I should squash the commits and pay the piper on the other thing when/if that bill comes due.
EDIT: Imagine this. If clean browsable history saves me 5 minutes a day then it has saved me ~10,000 minutes since I started using git. That equates to 1 full time working month. I really can't think of a circumstance where having access to fine grained commits would deliver a similar net savings.
FWIW, the way I use git results in me committing a lot. I treat it as, essentially a save point that I can undo to if I need to. The history for my personal branches is a mess of broken tests and false-starts on code paths that just aren't going to work. Useless for anyone but me (but fantastic for giving me an hour-by-hour breakdown of what I've been working on and what attempts I've made).
I definitely advocate for this. But I've worked some at some spots with "never squashers" that deliver every PR as 119 commits and tell me this yarn about "the time having all those commits totally saved their bacon". I've just never regretted squashing my commits into something manageable.
The pull request workflow popularized by Github and adopted by others (like Gitlab) is just bad for code reviews. They force you to choose between squash and history. Making several, smaller PRs are not the universal solution either. In a lot of cases those PRs have a dependency over each other and dependent PRs is also a pain in the ass in Github's PR workflow.
In a better code review system you don't have to make that choice. For example, in Gerrit (I _think_ this is also similar in phabricator, but I'm less familiar with that), every code review will end up a single commit when it's merged (if you configured it so, and that's the default configuration), and you don't lose the history during the workflow, all the intermediate states are stored as git refs on the server (you can't use bisect on them though, but they are there and you could write your own custom script similar to bisect to loop through them). Code reviews with dependencies are also handled naturally.
I and several others I'm aware of spoke with a product team at GitHub about this over a year ago, but nothing seems to have come of it. In the meantime, my default is for any repository I maintain to have atomic commits with fully descriptive commit messages only.
The kind of use case described in the article is not, to my mind, a good reason for the purpose of a commit message - describing for future maintainers _why_ a change has been made - to be defeated.
https://euroquis.nl/blabla/2019/08/09/git-alligator.html first parent log isn't mentioned, though.
It’s better for people to get comfortable dealing with commit history than kicking the can down the road by getting the abridged edition. You’re infantilizing your coworkers, and hamstringing the people who do deep root cause analysis at the same time.
>Maintaining safety equipment takes time and effort we could be using for something else
You haven't established the relationship between"safety equipment" and your preferred style commit history. You haven't shown one to be safer than another.
>It’s better for people to get comfortable dealing with commit history than kicking the can down the road by getting the abridged edition.
I do deal with the commit history just a simpler, more meaningful version. If you could write code perfectly with as little effort as possible wouldn't you do it? Well we're not perfect but we can go back and change the history so it looks like we were :)
There's a middle ground though: rebase your commits, but not necessarily into a single one. Before I submit a PR, I rebase my branch (which has lots of small commits, some of which undo previous work or are a work-in-progress), and make sure that every commit is as small as it can be without including work that is halfway done (so all tests still succeed), and have a clear description of what they're doing.
I then have both a nice history, and I can bisect to find a problematic commit, or inspect the commit history of a single line to get more context about it.
Especially "git rebase -i -r" is an incredible tool.
Do you have a good way to deal with this?
My one liner for viewing the git log is: git log --oneline --all --decorate [--graph]
I haven't tried, but I bet HEAD~n works for --fixup, or some quick rev-parse magic (which I don’t know much of).
For context, my career has been in web development start ups, which generally reward the cavalier and tolerate the careful.
Any one person can muddy the history up (by ruining bisect, say), but any one person can also improve it (by leaving good notes and right-sized commits whenever possible).
If a strategy demands perfection, it sounds like wishful thinking honestly.
The story told in the article mainly shows to keep your PR focused. You can squash as long as you don't do 2-week-long 37-commits branches.
A lot of higher level tools don't expose this functionality but it's not like it's not there.
git fetch
git merge origin/master
git reset origin/master
git commitIn contrast, if you instead merge, you only have to resolve conflicts once, for the merge commit.
That being said, I tend to be very fastidious about cleaning up my history before pushing, even if it means tedious, repetitive conflict resolution. But I do understand that other people might not think it's worth the trouble.
That said, I must say that I've personally experienced as a pain, so I don't actually do that.
The commit history is performance art. You don’t need to know that I left out a comma. Or that I tried to make a bit of code easier to read but introduced regressions in the process (that also breaks bisect). You do need to know that I changed this code over here because I changed the code over there and broke something, because I either missed something or you’ll have the same issue the next time.
This will leave all of the final changes in the index, "squashing" any churn. From there, individually stage lines/chunks into a small handful of atomic commits for the PR.
Whether they squashed their commits or not would make a far smaller difference to the time taken to spot the bug than simply not assuming that it's everywhere else than their own code first.
Because 19 times out of 20, that "wild goose chase" finds the bug faster.
As there is no perfect debugging methodology, sometimes whatever method you use will go wrong and you'll end up on the bottom of your list of things to check, or worse, right off the bottom of the list. (Those are some bad days.) That is not, itself, proof that your list is broken. After all, the author found this an unusual enough experience to write a blog post about.
This can be done with break points, of course, but in my setup adding a print is usually just faster.
This particular goose chase involved digging into everything except the developer's own code. 19 times out of 20, the issue is with your own code. OP even acknowledged this in the blog post, but his actions didn't align with that acknowledgement.
Well scoped and well sized commits, squashed, and in my personal preference rebased, provide a commit history that's segregated on actual tickets/stories/features/fixes that I find are navigatable very far back.
I'd say squash usage is very much case by case. But yeah equating a feature like squash to large commits is a fallacy. It's a useful tool imo that can be misused, but that doesn't mean the tool is "bad".
If you're simply talking about removing "fix typo" commits, then just don't. Just ignore them. You don't need them, but someone might. They're not hurting you?
Phabricator’s preferred model, which is heavily influenced by Facebook, is to forgo feature branches entirely and just stack many small changes on top of each other, landing as and when you want (this doesn’t preclude you working on feature branches locally, of course, because Phabricator doesn’t care what your local checkout looks like).
Because of this, Phabricator considers each diff to be discrete, and if you have multiple changes making up a single feature they should in turn be broken down into separate diffs.
Personally, I think stacked diffs are the killer feature of Phabricator. Unfortunately I haven’t been able to find a similar flow with PRs (recently we migrated from Phabricator to GitHub for one of my projects), you end up fighting against the tool a lot.
It's called ghstack (https://github.com/ezyang/ghstack)
If you want to learn more you can email me at ericyu3@gmail.com. Also happy to help you get it up and running - just put some time on my calendar at https://calendly.com/ericyu3/15min
> they really would like a complete story, not 20 gargantuan commits that contain 3 years of development
That sounds like maybe we split up work differently. 3 years of development for me or my team would likely have hundreds+ of merge/squash commits, not 20 large ones.
With the people I've worked with, I'd say most don't commit at all until they think the code is "ready" and they commit all at once. In the teams I've worked with, squashing vs not squashing isn't the question. I just want them to commit/push as soon as they've hit a stopping point or at least once a day. Maybe the people you've worked with are good with git but I am not that good with git.
I'm still stuck on 6: Resolve a merge conflict on git exercises because I made one too many commits and now the exercise says I have too many commits.
https://gitexercises.fracz.com/
Previously on HN: https://news.ycombinator.com/item?id=24671638
You’ll have a ton of noisy commits that together make up one full feature. In this case having all that noise squashed into one commit with a proper description is much nicer.
I’m asked “how long has this been broken and what caused it?” and answer it via git blame or bisect probably every few weeks. This is life on legacy projects.
But other than that? Small commits please.
And some features are just monsters, either through the nature of the feature or architectural choices that were made before it was conceived.
That way, you can have "nice" development history where tags for the old versions are redundant since they match commits one for one:
$ git branch
main
$ git log --oneline
abcd000 (HEAD -> main) version 5.2
abcd111 version 5.1
abcd222 version 5.0
ef01234 version 4.7
9876543 version 3.4
fedb123 version 2.12
dedc456 version 1.128There's nothing wrong with squashing a bunch of commits that really should have been a single commit from the start:
- Update X to do Y
- Fix typo in X
- Add "foo" option to X for when Y is a bar
But most of the time "features" consist of many changes: we add one function we're going to need, then another, then yet another... Then we change an API to expose the new functions, then extend the UI to make room for it and finally make put it into the application and pull all the strings together.
Squashing all these into one commit is just a bad idea. You can always do git rebase -i before merging to do minor fixups or even reorder your commits (like when you notice a typo after several other commits), but completely removing all granularity... just no.
Commits so small that they're nonfunctional also break bisect.
If you do make them by mistake - and everyone does sometimes, certainly including me - then sure, commit --amend or squash them before pushing them up.
The squashing the article is talking about is collapsing functional, distinct commits into one when merging to the trunk. It would be useful if we had distinct terms for the two kinds of squashing; the git command commit --fixup and its associated interactive rebase operation suggest the name "fixups".
Both are (interactive) rebase operations. Both are established terms in git, and it is probably not very useful to change that nomenclature now.
"Hmm, I had a half working X509 chain resolver there that turned out to be unnecessary at the time but would save me a day's work now..."
I don't think there's an easy answer here. I think there are good things to be said for a readable and bisect-able version history, but also that if you're not preserving your real commit history in a way that's backed up remotely then something's probably wrong.
I'm beginning to think this area, this schism into two schools of thought, is really a signpost that there's something lacking in git's branching. Or something everyone is missing, including me :)
Depedends on the git flow being adopted by the commiters.
HNers constantly espouse how clarity is more important than cleverness. It may be the case that a single commit offers more concision than multiple commits which could be perceived as just more noise.
I'm sure that's barely any compared to large companies, so I really question the value of commit history unless message guidelines are well enforced.
I don't often (ever?) browse the history, but I _do_ regularly search the history using tooling, (git log | grep <something>), and more is more in that case, even if the history isn't perfect.
And that gets squashed into `Change X to 50 by Baz`, you lose the context as to why Foo and Bar changed it, and the values. If I need to go investigate an issue with X, I'd rather thave the history of all the changes.
In my case, it's generally because I'm preparing training material.
When I do that, I have two branches: dev and master. master is the one that is exposed to students, and dev is the one I use for prepping, testing and staging.
Sometimes, I may even have the dev branch in a private repo, or one that is associated with a different GH ID, like so: https://24ways.org/2013/keeping-parts-of-your-codebase-priva...
The main "gotcha" for me, is to make sure that I merge the master back into dev, after doing the squash. Makes life a lot easier, for the next squash.
If I need to bisect a bug that was introduced by the PR I still can do this in the original branch.
With some discipline this makes the commit history actually worth reading and makes git blame a useful tool.
So I end up with "one commit per feature" rather than "one commit per logical change". So I actively using squash during interactive rebase when I prepare to merge completed branch since my commit history sometimes looks like this:
Backend: new set of APIs for XYZ
Backend: implement feature X (+ some long comment)
Frontend: implement feature X
fix for feature backend
fix for backend API XYZ
front fix
backend fix
Yeah again I know I could do better, but yeah I use squashing for this reason. A lot.I'm sort of agnostic on this issue, but I do feel the article's author kind of overstated it here. Say 5 to 10 commits of this size were squashed together. Git bisect would've taken him 90+% as far and he'd have to read code or manually trial and error changes just slightly more. The binary searchable problem space would be slightly smaller, and the linear manual effort space slightly bigger. Less good, but really not that big of a deal.
If i am making a commit that makes a complex but important change to some significant application logic, i want that commit to contain that change and only that change, so that when i have to re-read it a year later, it's completely obvious what i did and why. Bundling a load of refactoring and cleanup in there is a significant speedbump for my understanding.
Years ago, a sage pointed out the argument for squashing is really an argument for better tools. Imagine if you could flag commits as being of two types - major/minor, significant/insignificant, feature/refactoring, foreground/background, melody/rhythm, etc. Then imagine if the tools would by default hide, roll up, or otherwise de-emphasise the commits of the latter kind. This whole apparent dichotomy would go away in a flash.
This idea is floating around in the Wiki world. I believe it was Ward's Wiki that introduced a 'minor edit' checkbox in the editor; if a change was marked as a minor edit, it wouldn't be show on the recent changes feed.
You can imagine other ways to get somewhere similar. For example, you could have a special kind of commit that just groups a previous run of commits, and the tools could show that and hide the members of the group by default. There are probably many other ways to do this.
I seems like madness any other way to me... why have a bigger granularity than a single merge commit?
If you want to continously deploy to production (with real users using it) you need a process to very quickly revert bad commits from it. At my last company with hundreds of developers, we would push to prod several times a day. Each deployment would have several devs changes included. A clean linear history is essential to making that process work.What do I miss?
Keeping the history of each incremental change, even in a branch, is IMO too useful to give up.
Note I'm not advocating keeping those "fixed typo in previous commit" fix-up commits: those should be properly fixed up _before_ merge, by judicious use of git rebase.
I had one job in the past that used Gerrit[0] as Git tool. One of it's features is that it creates a "pullrequest" for every commit in the branch you push. Which needs to be reviewed individually. This is really anoying if you're used to organising your work in lot of commits to record each step of your developerment. But from a project's Git history perspective it makes a lot of sense. As every commit is 1 change, one feature, one contained unit, that is added to the main branch. So instead of the main branch now containing countless commits with each developers complete history on a specific feature (where code is added in one commit to be removed in the next) it contains the features as distict commits, making them easy to bisect and revert if needed. This looks a lot like squashing but because you do it before you push your code you learn to put much more thought into that single commit and the commit message.
But the second best thing is to think and work messily, and use rebasing to fake a history of small, logical, incremental commits!
The alternative is making bisecting harder which is not something that I want. I want every commit to compile and be testable individually. That definitely doesn't mean that I think it's a good idea to have "gargantuan" commits.
An ideal commit should contain a single, atomic change to the codebase. Not more, but also not less.
But doing this means I have many commits that don't mean anything or are unfinished and in an uncompilable state.
So I git rebase and massage the history to make more sense.
https://www.mail-archive.com/dri-devel@lists.sourceforge.net...
https://www.linux.com/news/why-linuxs-biggest-ever-kernel-re...
If I make 10 commits I can squash it into two logical commits (e.g. refactor, add feature) then I merge the branch with those two.
If the branch is small and has 3 commits that are one logical change, then I might as well squash to main instead of merging.
Commit history is readable regardless (I'd never use anything but --first-parent ever).
There are also people who insist you always need to rebase your commits before pushing, in order to get a nice, linear history, and again I disagree. It's fine when you're rebasing a very short (single commit) history, but for a long history, it not only gets very tedious, but when it introduces new bugs halfway through that history, you may not notice. You will notice later, and then tracking that bug you end up in the middle of your own (rebased) history and you may wonder why you ever did something so stupid, when the actual cause of the bug was the merge of two histories at the end of your work.
I do rebase sometimes, but only when the history is short and I can easily see what I'm rebasing. A history longer than 2 commits should not be changed.
I never squash.
It's really quite easy to get into the 50s or close to 100 commits in a branch, which can lead to a horrendously messy history.
There are also people that like to commit before the try something locally and their branches are effectively a mess.
All of the above is magnified in a monorepo environment, where you might have thousands of commits a day against master if it weren't for enforcement policies that force a squash.
As mentioned elsewhere in the comments here, that also makes it a lot easier to revert said feature if something goes wrong.
* Commit 1: Fix a bug
* Commit 2: Fix linting issues with the fix discovered through CI
* Commit 3: Remove dead/commented code introduced thrugh commit 1
* Commit 4: Update documents as required by change
I'd always squash those into a single commit before merging into upstream.It's a good idea to isolate the bugfix so that reading the diff is clear. Stylistic changes such as #2 and #3 can be separated and called out as "should produce no behavior change" sort of updates.
I once had to enforce to an developer that wanted to make 20 commits per PR that were titled like "wip" "wip" "wip" "fix bug" "fix mistake" "format code" in a PR that only changed like 10 lines of code in the end.
In this case, the main problem for me was that git blame became useless because this developer was touching way more lines than necessary and then undoing it using the code formatter.
Also, we auto-stamp the PR on the commits so we can get the context. I don't understand how anyone could make sense of a large codebase when all the commits are added with the standard inane commit messages (e.g. "fix it", "typo", "name changes", "tests") that people do when building a feature. I routinely have to look at a piece of code, do a git blame, get the PR in the commit and figure out what was being done.
I also aggressively reorder my commits. When it gets to yak shaving you basically develop in a stack, so you end up with a half-functional commit to system A at the top, then a complete commit to system B, then the finishing commit to system A, so it's better to just rearrange it so that you have one system B commit followed by one system A commit.
But I do routinely read the commit history (using git blame) and no I don't want to see the "complete story" of my past self having "got to this point"--I want the documented MR commentary.
# Implement feature
# Whoops fixed issue in feature i just implemented
# Add in whitespace
# Remove whitespace
# Forgot place to add in whitespace
# Fix variable name for feature
Vs # Implemented feature "x"
Which ones easier to rollback and read.I would never send out a patch set for review that includes all the "oops" "typo" "iteration 50" etc. commits I make as I work on code. Those are pure noise.
However, git is a tool for development, and during development I should be able to use git in whichever way is convenient for me: commit, rewrite and do whatever the hell I want with my local history. Not having that freedom is the primary reason why working with most non-distributed version control systems is such a pain.
When it comes to actually merging patch series to master, I like not doing fast-forward merges, since you maintain a natural grouping of the applied changes.
Yes, squashing parts of a discrete piece of work together, where it makes sense, then makes it trivial to git bisect any future problems.
Code committed to the mainline should always compile, and should be made of discrete changes. However, do what you like on your unpublished local branch.
> What's even the point of it? It's not like people routinely read the commit history, and when they do, they really would like a complete story, not 20 gargantuan commits that contain 3 years of development.
My company routinely reads the commit history. The commit is usually the 'why', and the code is the 'how'.
Of course, that only works after you leave proper commits. The reason people don't do that is because you only look at your commit history if you've kept it useful, but if you've never looked at your commit history because it's not, you don't know the benefits of keeping it useful.
I suspect there is a pattern in the comments here. People who work on small teams with granular repos that, individually, don't see a lot of daily activity think squashing is bad and erases valuable history. People who work with large repos that see high commit velocity (like a monorepo) think squashing (at least merge squashing) is beneficial and don't see the loss of information as problematic because it's hard to access it in the first place. Maybe I'm just projecting my own opinions on this; I'd like to hear perspectives that conflict with my assumptions.
A commit is a change. A change has a ticket. A ticket is a small piece of work that does not result in 3k changed lines.
This ensures that the change rationale is fully documented and easily identifiable.
Nobody needs those "fixed typo" commits. Nor the "implemented function A" commits. What IS the change that you're doing? What functionality? That's the most comfortable commit granularity to debug imho.
I could obsessively fiddle with every commit like an artesenal snowflake, or I could click the squash checkbox on the request.
You take bug or enhancement of a couple of points, work on it on its own branch, then merge req+squash the mostly noise commits into a develop/master at once.
A chunk is approximately 3 days, not 3 years. For forward-looking projects, it saves time and works well. "Enterprisey" projects maintaining legacy branches are less well served.
And then after the PR has gone through review, most people tack on extra commits with equally-useless commit messages like "addressing feedback".
It's infrequent that I see PRs with commit histories that actually chronicle the history of the change itself; it's more a chronicle of the developer's changing thought processes as they try different things, go down blind alleys, change approaches several times.
Most of the individual commits have test suite failures and some of them don't even compile, so it's impossible to bisect across them if an issue is later found.
In those cases, I wish people would just squash (and some do, where I work). Yes, you lose information and separation, but I'd rather have one large working commit than 15 small broken commits followed by one working commit. Ideally people would curate things before submitting their PR, but I've found that most people just don't care, or don't understand git well enough to even attempt to do it.
Sometimes I toy around with the idea of trying to teach people (what I consider) better practices, and insist they are followed, but we all have a limited amount of social capital in our workplaces, and I'm not convinced this is something worth spending it on.
It seems like a small price to pay for a clean history, given the rarity of occurrences like this.
The author has only used bisect for a head-scratcher once, but I've used it often, sometimes to pin down bugs in code I'd never even looked at before.
That's not feasible if the commits are huge.
Also, the method I described has nothing to do with commits. PRs and small commits aren't mutually exclusive.
That was the price I was thinking of.
You're certainly right that PRs and large commits are orthogonal.
How is this annoying or time consuming? It takes less effort than adding, committing and pushing changes.
I don't like having to go fiddle manually for that to work.
Ideally you have CI set up so the average test failure email points you to the small patch that broke it.
For bugs from the wild that doesn't work, obviously, but I still prefer less friction.
I don't have this set up in my current gig, because we have more pressing priorities.
It is incredibly useful when you do have it, though.
Generally it's more important for really large projects - IIRC Android is really strict about not merging branches containing any commits that break tests, for exactly this reason.
As that implies, you wind up rewriting branches in this approach, too, but for different reasons.
However, I try to only squash into logical units, I don't try to cram the whole PR into one commit
You don't remove PR's branches after merging them?
The ability to use git bisect effectively is one of the more important reasons to enforce a clean and readable commit history by squashing and rebasing before merge.
A history where the majority of commits doesn't even compile ("sorry, updated test value was wrong", "oops syntax error", "forgot to update these references in the last commit", "big refactor wasn't complete") is a major headache not only to readers but to anyone using automated tools such as bisect. The author here was lucky the commit history was in a good shape. The size of the offending commit was reasonable too, so any trivial commits had been squashed away here.
My personal issue with this particular commit is the useless commit message. "Extract CreateTokenValidationParameters method". Well, obviously. But why? What was this intended to result in? Why was particular change made and not something else?
A more suitable commit message would have included something along the lines of "JwtSecurityTokenHandler methods belong conceptually with other code that configures JWT parameters. Break them out to CreateTokenValidationParameters because ..", that would have make much more sense and made the change easier to understand for someone else!
After all, this is how the author describes the patch, when taking the time to do so in order to write a blog post. It isn't that hard.
Git bisect can only narrow you down to the scale of your commits. If you do infrequent, large commits, it's not very useful. If you do frequent, small commits, it's great. If you do frequent, small commits and then squash them into infrequent, large commits, it's not very useful.
There are plenty of reasons to squash (and, personally, I generally think they outweigh the arguments against), but this is not one of them.
When you do frequent small commits, unless you are superhuman the majority of them false starts or contains errors. Squashing these useless commits makes the history understandable, and enables the use of tools such as bisect.
That's what "please squash before merge" means.
It does not mean you should do rebase instead of merge. There may be good reasons for that too, sometimes, but that's not the point here. Should you wish to enforce a linear history, it is still just as important to squash undesired commits.
This is a misunderstanding of what proponents of commit squashing advocate. Often, when working, you end up with multiple commits which represent a single change:
remove debug logging and fix bug
more logging
debugging
typo
Squashing these commits together into one coherent change makes for a history that is much easier to understand.Squashing unrelated commits, however, doesn't help anyone.
(I prefer the --first-parent approach: https://web.archive.org/web/20180710234754/http://www.davidc...)
Those are four commits. That's what happened. It does not matter that it's untidy. It's your history.
Just leave it. It will save you a lot of work some day when you need to find the error you introduced when you accidentally removed one too many lines on that "remove debug logging" commit.
I already spend a fair amount of time on other people's problems and they aren't doing much to save me, so what is the point of being obsessive with my own problems such that I may avoid spending a little bit of time on them later?
But you wouldn't commit every keystroke, because it's not informative. It doesn't tell you anything about why you did something.
We're not advocating throwing away history specifically, we're advocating writing helpful history, leaving helpful traces.
That is only true if you are working in a personal git repo
If you are committing against a repo with many others though, having a tidy git history is important. If I discover a bug and bisect to the "remove debug logging" commit, it is hard to revert that - or to even understand what it means. I then have to go in and try to determine which feature that debug logging is for and how many commits I need to revert to get back to a sane place.
My general rule of thumb is that I squash commits (when merging to a public repo) if the commit doesn't warrant time into creating a proper commit message (explaining what I am fixing, why, etc). I have no problem creating a set of 10 commits to merge in one feature, but each commit needs to be individually reviewable and understandable.
I definitely leave in commit A and a revert of commit A if I think it'll be useful.
The way I see it is I'm leaving breadcrumbs for myself for later. It's obvious that a sequence of keystrokes is useless (and most IDEs will undo a large number at once - correct behaviour) and that no history is useless so you have to find the thing that works for yourself or your team wherever.
I understood the parent as "remove debug logging" removing what was introduced by the previous two commits. In this case, squashing the commits will actually make the issue _much_ more obvious, since the accidentally removed line now stands out by itself without the clutter of the back-and-forth changes.
This however is something I've seen people do. Worst case I've encountered was a group of people that squashed all commits belonging to a whole product release - we are talking a team of 6 people working full time for half a year. Absolutely atrocious behavior, but there are lots of weird people out there.
Which can be perfectly fine, as it's hard to ensure that all developers use Git in an optimal way (I'm thinking of intentful use of interactive rebasing), uniformly.
The only problem I find is when squash proponents claim their choice is superior. It's not; it's only superior if you aren't willing to maintain great history as you work on a given branch.
Of course there are also added benefits to a fine-grained history, as the article mentions.
My experience having maintained a production app with a fine-grained semantic history for 5 years is overwhelmingly positive. Understanding root causes, confidently reverting things etc becomes a much quicker job - particularly important when production is red.
- a historical, fine grained log of changes
- and a log of merged features.
There's value in keeping both of these data sets. But the commit log as it stands can't easily serve both masters.
Once you can see the problem for what it is, the solution is simple. Instead of conflating these two use cases into the concept of a 'commit', we need separate tooling for each of these use cases. The commit log should probably house the historical record. And then we need a way to mark a set of commits as belonging to a particular feature's development. That could be achieved either by adding special support in git or via convention using commit messages.
Either way, I want to be able to see the commits in my repository grouped by the feature that they belong to. And I want that data set browsable on github, referenced against the corresponding github issues when thats appropriate.
git log --max-parents=1
git log --min-parents=2
E.g. try this on the git.git repository (not perfect there, since Junio doesn't use this pattern all the time), but it's good enough for a quick demo.The most useful convention in a DAG like git is to have the topology of your commits reflect your workflow.
Straw poll - do you enforce this? If so, for literally every commit, or do you use partial squashing to maintain this property?
While that's certainly a _desirable_ property, I've never really been concerned if, say, the penultimate commit on a PR failed CI. It feels like it would be a hassle.
TIL about `git bisect` and it's really vindicating.
If the validation isn't extremely fast (seconds or minutes) then maintaining it for every commit is almost impossible. It would just get other side effects like people avoiding commits because they don't want to run an hous-long test suite more than once.
As I work on a feature branch, I'll check in WIP commits as checkpoints, especially at EOD. I don't expect these to pass the full CI suite.
But as the code starts to shape up, I'll unstage all those WIPs and start to group the changes into logical commits. As work progresses, I'll use `git add --patch` to split new lines of code into those existing logical commits. Sometimes I'll split one up, sometimes I'll group two together; it's still flexible and amorphous at this point.
By the time I'm ready to merge upstream, these commits tend to be neat, focused, and functional, and I do check to make sure they pass the relevant tests (though I don't enforce a full CI build here).
Then a rebase from master, push to CI, and then a no-ff merge commit into master to retain both the low-level commits and the ability to easily revert the whole lot.
It might seem like a ton of busywork, but I find that staging atomic commits like this doubles as an excellent line-by-line review of the code I've written. It also forces that review step to happen throughout the process rather than all the way at the end when I've forgotten all that deep context.
IMO squashing commits is something that you should do locally, not remotely. For example I'm currently debugging a fairly large C application on an embedded system. Each bug has it's own branch. When debugging / testing / trying to fix it I tend to do a lot of commits. When the bug is fixed I do an interactive rebase to pick / drop / squash commits in order to have only a clean one at the end (and then push it).
My point of view is that if work is being done on an individual feature or bug fix, having a view of the individual commits might help give any reviewers important context on how a final solution was arrived at. It can also be helpful to be able to see the specific changes that have been made since a prior review.
While in most cases I favor squashing commits when it comes time to merge into the parent branch, it seems like doing it manually just creates extra work and potentially throws away information that may have been useful to reviewers.
- "Squash your commits" folks - yes, it's good to be able to revert a commit and remove entire features and have a readable commit history.
- "Make small granular commits" folks - yes it's good to be able to bisect and see where exactly some behaviour changed.
Rather than repeat these points (which are both true), there's a better question to be asked:
Should there be a way to have a readable history and keep individual commits? Eg, 'subcommits' etc?
- git may not have these features now, but revision control systems come and go (rcs, cvs, svn, bk, hg, git)
- Or maybe it does. Or we can implement something similar on top of git as it stands.
Discuss.
git revert -m 1 $sha_of_the_merge_commit
That provides the "revert a commit and remove entire features" use case, assuming the original feature was all encapsulated within the merge (branch) in question. You can get the "readable commit history" with: git log --no-merges
You can have the best of both worlds. git has had these features for as long as I can remember, but devs have had the "NO MERGES ONLY CLEAN HISTORY" platitude repeated so often they must think git lacks an alternative.In my opinion, squash merges make the most sense when the following are true:
1. Developers cannot push directly to shared branches; all commits into shared branches are done via pull requests with mandatory code review.
2. Shared branches must build and pass tests at every commit.
3. The team builds features incrementally and uses the "Branch by Abstraction" approach (feature flags/experiments) to ensure functionality can be merged before it is ready to be enabled in production.
4. Changes that merge into the shared ("trunk") branch are released frequently.
Now, you may argue against some of those practices, and if you end up in a different place on one of those, squashes may not make sense to you. But then the actual difference of opinion is elsewhere. It's not really about squashing, it's about how much code is a reasonable granularity to code review at one time.
If you can keep the code review cycle tight, then a pull request branch basically becomes almost like a virtual pair-programming session. There may be some back and forth as the author and reviewer settle on a final form of the commit. But I don't see the commits that happen along the way as valuable in the least. (In my experience, the vast majority of intermediate commits are things like "fix typo", "make linter happy", "rename this function because code reviewer pointed out it was named inconsistently", etc.) Instead, my mental model is that a pull request is like a mutable commit: you keep mutating it until it's the commit you want, and then you merge it. Since my mental model is that a pull request is a kind of commit, it makes sense that when it merges in, it does so as a single commit.
Again, I'm not saying that this is the only valid way to work. But neither is it an invalid way to work.
As a meta-comment: When you see someone else making a technical decision that doesn't make any sense to you, rather than instantly assuming that they are a clueless incompetent, it's usually more instructive to assume that the decision does make sense to them, and to try to figure out what about their circumstance is different from yours to make that the case.
I think part of the problem is that modern repository software has built an 'additional layer' on top of git (IE: pull requests) that most of us have become accustomed to using in our day to day workflows. Git itself doesn't really have a 'lossless' way to persist that extra information.
Squashing commits at merge gets us close, but ultimately we’re throwing away potentially useful information by doing so. The main time I’ve seen it be problematic is when multiple people are working on a complex feature where PRs are getting merged into a long-lived feature branch rather than directly to the default branch.
Say I need to base my work on someone else’s WIP branch for said feature. If they squash their commits when merging their work into the feature branch, I’m going to end up with merge conflicts that need to be manually resolved as the individual commits that I had based my work on no longer exist, as they’ve all been squashed into a single new commit.
What if git supported a sort of 'soft squash' concept, where associated commits could (optionally) be grouped together with additional metadata. That would let the application consuming the git history (IDEs, CLI tools, etc) make the decision in how commit histories are presented to the user.
While that may be true, I'd say in this case, order-dependent tests are evil. This case study pretty much fits the definition -- the test result changes depending on what other tests are run before it (due to shared state across tests). It certainly is possible to write automated tests for this -- reset that state to its default before each test!
One thing that Google recently (in the past year or so) added to its CI setup is a job that goes through periodically and randomizes the order in which tests are run, thereby increasing the chance order-dependent tests are caught. (Early on, I remember being annoyed by all the triggers due to floating point error accumulation in a monitoring library -- do I really care about an error of 1e-10 when my noise is >10?)
That's a long-winded way of saying that it's also possible to catch order-dependent tests by shuffling the order in which they run, although built-in support may vary depending on your testing framework. (from a brief search, it looks like rspec and junit both support randomization; xunit forces it).
That said, I absolutely agree with the other conclusions, especially that binary search for debugging is immensely useful (even when dealing with Google-scale tens of thousands of unrelated commits due to monorepo), and definitely empathize with debugging often taking much longer than the bugfix.
But IMO it's not a big deal, you just do "git bisect skip" for these comments and at the end of bisecting session instead of one wrong commit you'll get couple of, for example one commits that builds and two that you skipped and doesn't build. It's very likely that looking only at this 3 commits will help you find the issue because they might be part of much bigger branch which would otherwise be squashed to one huge commit with enormous diff.
I much more prefer to require that PR's branches are always merged without fast forwarding. So there is always merge commit for PR. Then you can actually display list of this commit with "git log --first-parent" and they should always build because build server verifies it.
Unfortunately "--first-parent" doesn't work for "git bisect" now, but it finally will in the next git release! [1]
[1]https://github.com/git/git/blob/ab4691b67bc1a2cd8d9068fb03e3...
- initial prototype of feature x, mostly working
- feature x working as intended in requirements
- fixed edge cases not originally identified when thinking about feature x
- rearchitect feature x a bit now that it's better understood
- write tests for feature x, most pass
- get all feature x unit tests pass
- a few more test cases for feature x
- and a couple of integration tests for feature x
- code cleanup for feature x
- fix typo from last code cleanup resulting in bug
- better docs for feature x
- fix typos in feature x docs
- update main readme to include notes about feature x
Can anyone explain to me the value in preserving this history, because I'm just not seeing it, and it would completely muddy up the overall git history. There's tons of intermediate work we're not concerned about preserving, like the notes I scribble on paper, so why are we so concerned about preserving all these commits? I admit there's a small chance it's useful in rare circumstances, but I'd rather optimize for the common scenario.
BTW if such a thing really does happen in production and you already squashed and merged, you can always fall back on on git reflog on the developer's machine. Commits are never really destroyed in git, the branch simply shifts focus to another set of commits. The old commits are still there, dangling and not referenced by any other branch but they are still reachable.
I usually prefer through space, because 1. it actively helps narrow the buggy line in the code you're working on, instead of reflecting over different snapshots of code 2. sometimes you have nothing to bisect on, if it's code you're actively writing.
The author seems to have started bisecting through space using the debugger. Unfortunately they had the wrong understanding of the execution flow and quickly stopped after one step at the start of the route. Had they realized the route wasn't triggered they could have checked what happened in the authentication code.
Squashing also turns cherry picked commits into merge conflicts.
I'm not a fan. I can always go back to the feature PR and read the ticket details and diff commentary when I want the history with a full diff. And of course very old history isn't very relevant because less of it remains in the present, so rot of ancillary systems isn't a huge concern.
> Squashing also turns cherry picked commits into merge conflicts.
1) To provide a code-level record of what has changed in individual files and why.
2) To provide a high-level record of what features were introduced and what bugs were fixed over a period of time.
Often people will forget about one of them when arguing for a particular approach.
Squashing can make 2 easier but annihilates 1. Rebasing gives you 1 but makes 2 difficult, or requires that you track high level changes in an external system. In theory, approaches using merge commits can give you both but they are often difficult to apply in practice.
A question I ask, is that if I reverted/rolled back to this commit, would the system still work? Commits that require other commits to work should never be in isolation in case a rollback is required. You should be able to check out any commit in the entire system and have a (hopefully) working program.
It got crystal clear when the breakpoint was not hit or when removing the Authorize attribute worked.
The repeated saying of my tests were passing and it must be the framework kind of annoyed me as the Authorize attribute is not magic and there need to be wiring stuff which need to be written.
I've asked this question in job interviews. It's amazing what people say.
With Mercurial's evolve, the hidden commits are always there. When you push/clone, etc they get sent around.
In Mercurial, hidden changesets are kept locally indefinitely, but they are not exchanged; only their obsolescence makers are. So you always know the meta-history of a changeset, but not necessarily their original content.
git log --no-merges git log --no-merges --first-parent
I don't get the "no merge commits" argument - git log has dozens of options allowing you to bend it to your use case.To address the grand parent that "a clean history is has no merge commits" I would argue that a clean history is also a lie if you're working in a branching workflow (hint: you should be working in that most of the time).
If you want to avoid (note: not eliminate) merge commits then make sure to rebase against master before merging, and then ensure merges are fast-forward merges.
* merge commit
|\
| * branch work
| |
| * branch work
|/
*
If you squash to merge, typically you're also going to rebase it (equivalently, if it's more familiar, cherry-pick the squash onto the branch your 'merging' it into). * squash cherry-picked / rebased
| * both branch works squashed
| * | branch work
| | |
| * | branch work
|/_/
*
(In this case the target branch could have been fast-forwarded, but this also works if there's some other work on the mean time:) * squash cherry-picked / rebased
|
* something unrelated
| * both branch works squashed
| * | branch work
| | |
| * | branch work
|/_/
* git log --first-parent
This only shows the top level commits (either a direct commit or a merge commit) and doesn't show any of the subcommits in the branches.This gives us a very clean history, something like:
commit ce29331f7da82ce528ca6e437b8893248a842169
Merge: a14522cbf 936ef9f90
Author: Joe Sample <joe.sample@example.com
Date: Tue Jul 7 14:34:20 2020 -0500
ISS-2047 - Add ids to user invitation or creation
commit a14522cbf32487dc590c8b6f3332d3fc9371a640
Merge: d08342170 cb94a86b1
Author: Jane Dow <jane.doe@example.com
Date: Tue Jul 7 14:28:28 2020 -0500
ISS-2032 - If an enterprise is disabled, their api keys should be disabled
commit d083421702e3ff50bc4c62e85b687e172e4bfe76
Merge: 0d167d25e ba3900a4a
Author: Joe Sample <joe.sample@example.com>
Date: Tue Jul 7 14:25:41 2020 -0500
ISS-2045 - Return a reason of why the invitation was auto-cancelled
Then if I want to see everything that happened in the branch that was merged in, I can just run: git log d083421702e3ff50bc4c62e85b687e172e4bfe76^..d083421702e3ff50bc4c62e85b687e172e4bfe76
and see all the commits that were in that feature branch. No need to squash things or hide them.I really wish git, by default, showed commits under a merge as a tree, rather than as a flat list. This is what bazaar (bzr now breezy) did, and it made a lot more sense when looking at the history.
Reason for the feature squash: most of the context of the original commit(s) are meaningful only to the original dev who is free to keep that history locally (or share with others). There are time's it's convenient to structure the commits in a certain way, but this is often most useful at PR review time, and less so after merge.
Granted, you may want to be able to re0review code at a later date, and hence keep that structure, but hopefully this is a rare case, and PRs and not often so huge.
## Squash
I often start in a feature branch with a new proof of concept, or other code which is persisted in several commits. Sometimes I'll iterate on a few things, like testing a change in a GitLab CI yaml, and checking whether it works. Or a different compiler flag to enable faster package builds. Or a refactored function which needs to run all the e2e tests to prove the performance gain.
These changes may, or may not work. When they do not work, I'll reset the commits - either soft to keep the changes, or hard to throw away the attempt. This follows a changed history and force push into the remote branch. In case of a shared branch, message colleagues with "git fetch && git reset --hard origin/branchname".
Within pair programming sessions, we often left with commits like "Add REST API HTTP server, WIP 2" in branches and depending on the availability, either one of us continued. At a certain point in development time, we decided to squash and amend the commits. Sometimes not all of them, as rebase/squash also allows you to do the following:
c1 s \
c2 s /
c3 p
c4 s \
c5 s /
Which squashes c1+c2, leaves c3, and squashes c4+c5 again. You can navigate into this on the CLI with "git rebase -i HEAD ~5".
## Rebase/Merge
Short-lived branches which are quickly merged back to the main branch shouldn't cause problems with a broken deployment. In case you get a task assigned where the main branch is far beyond (say, 100 commits or more), it may be the case that
- Branch and merge request works fine, CI/CD pipelines are green - Changes in the main branch which affect your feature.
These changes can be
- Function interfaces renamed, or not existing. Easy to fix upon rebase, build/run does not work anymore. - Runtime changes, for example, queries take longer roundtrip due to a refactor. Your feature only takes the old behaviour into account, and increases the runtime complexity. Or it consumes 10x memory resulting in OOM crashes later.
The last change may not be immediately visible, as it involves staging environments and application performance monitoring results.
### Merge without Rebase
If said changes occur, and the rebase did not happen, the green CI/CD MR is merged back to the main branch. Depending on the releases, you either roll into production, or after days/weeks/months, a new release is cut.
At that point, the regression may be seen in the main branch, and cause delayed analysis and debugging. Often times on-call alerts and all the debug fun which may lead to burnout (been there myself).
### Merge Request with Rebase
During the final review, and prior the merge, the changes are rebased against the latest base in the main branch, to see if they compile or any other influences.
A rebase puts the existing commits onto a new commit base, which influences the calculated checksums. Therefore all commits are newly generated, the author date is preserved with changing the commit date.
### Merge Commits
There are different opinions on them. One of them is to always rebase the MR and then do a merge with a commit. Rationale: Even without GitLab/GitHub/etc. you can reliably see the git graph on the CLI or with other visualization tools.
https://gitlab.com/dnsmichi/dotfiles/-/blob/main/.gitconfig#...
I've recently seen the possibility to reference a PR/MR to a commit as the merge-from-branch reference, without the dedicated merge commit. This can be handy to avoid it, with using GitLab/GitHub/etc. to store this detail in their database. It also is a vendor lock-in in a way, that the native "git clone" does not provide this information for you in Git's database.
That being said, I used to dislike merge commits. With enriched details, and CLI work, I now prefer them again. Git commits as datasource are valuable, and they can be shown/parsed in any environment.
### Rebase, Merge, ... large environments?
This can of course get more complex, with fast moving main branches and lots of merges which depend on each other, and should not reach the main branch. Instead, you'd want them to be queued and tested. We experience that at GitLab quite often, and have created so-called "Merge Trains" which ensure that all MRs in such a queue/train are taken into account: https://docs.gitlab.com/ee/ci/merge_request_pipelines/pipeli...
## A personal note: The best merge/rebase strategy is nothing without tests
I've been working on a monitoring tool in the past which includes distributed environments, and often needed to fix bugs with memory leaks or other performance issues in multi-threaded scenarios. Things you do not see immediately when the MR/PR is green. There was one commit which caused a OOM crash after 3 days of runtime, but only in cloud environments with >100 satellite nodes.
The turnaround was to bisect all the commits, and run each of them in production for 3 days until the crash occurred. IIRC we had 1,200 of them to do in a binary search.
Now the question is:
- Fewer squashed commits - More development history
In this case, fewer commits would have unveiled the error sooner. The resulting commit would be larger and harder to debug & fix though.
In the end, it did not really matter. The thing which would have helped: There was no reliable test environment coupled to CI, CD which ensured to run specific commits & MR/PR in dedicated scenarios and alert of breaking changes soon enough - before the release happens.
## Conclusion
One thing which greatly helped: Looking how others do it, Open Source projects and customer success stories and webinars, online training sessions. Even though you may not adopt the workflows, trying them out is a good way to learn. For instance, "trunk based development with feature flags" is something different to well-known branching models for me, I needed to try them out first to change my opinion. They are indeed useful for certain scenarios.
While committing changes, and keeping them throughout un-squashed MRs, I always remember that I will be highly likely debugging the changes later on. Or someone who finds the MR reviewed by myself, documenting every thought or idea in a commit or MR comment can help.
Some more tips and exercises are discussed in an OSS training I created in the past: https://github.com/NETWAYS/gitlab-training/releases/tag/v2.5...
I say this because as a low level/system developer, I often have to solve problems for the people integrating my platforms from higher level languages. When something doesn't work they come to me, and it often goes like this (with quotes from the article):
First stage: someone else's fault.
"the problem is environmental: a network topology issue, a bad or missing connection string"
Second stage: magical things, but not my code.
"Not my code ... the problem occurs somewhere in the framework"
Third stage: acceptance.
"Global variables are evil. Who knew? ... It took me hours to find the bug, and ten seconds to fix it."