Resistance Against Git Merge Hell (2015)
tugberkugurlu.com
tugberkugurlu.com
People seem to have very strong opinions about this. I don't much, either way; I have to use the gerrit flow at work, which mandates a lot of rebasing, and in general I prefer not having a commit which only records a merge. Some people seem to want highly detailed tracking of what an individual developer has done in their personal checkout. I will only note that the act of observing changes what is observed.
There is some benefit to my tendency to do experimental work in a branch where I've started to work on something else. In that mudpit, I find serendipitous solutions to multiple problems, and I find incompatibilities between desired changes. And when multiple issues have overlapping changes, there's usually an optimal ordering of which to address first -- and that's not always obvious.
I greatly prefer to do all the forensic work up-front, to make a clean history with a cogent story in the commit message. When I see other people's messy histories with every merge/revert, I need to do that forensic work every time I go back in history; without the benefit of recent first-person experience.
But there's a difference between a few separate things here:
- assembling your own chaotic actual work into cogent commits using interactive rebase and friends (but not moving the base, e.g. git rebase -i --keep-base): good and necessary. people doing reviews should also review commit history, not just a megadiff of all changes.
- rebasing in the sense of "moving the base upon which your work is based", which for unfortunate historical reasons uses the same command: this is rarely ever useful, despite how much people seem to enjoy doing it. replace this with a no-op in nearly all cases; and in the few cases where you might want to (the upstream you want to integrate with changed incompatibly) merging a recent version tag into your topic is better.
- rebasing in the sense of "as a maintainer adding completed work to an integration branch like master, rebasing the commits atop the integration branch rather than simply merging": this is actively harmful (as is "squash merging", which is this _plus_ destroying the commits crafted in point 1) and is also unfortunately what most people seem to mean.
my "reroll" alias is "reroll = rebase --interactive --keep-base --rebase-merges"; but the --rebase-merges is less useful to most people (I do maintainer operations as well, and it's sometimes useful there)
autosquash is on in the config, though i should really add it to the alias for redundancy, it's an alias and length doesn't matter.
and for the last one, how is the maintainers rebasing my branch onto master any different from me doing it before submitting my changes? and why is it actively harmful? (i agree with squash merge, that's not useful. in a project that does that i have resorted to carefully crafting my commits such that each commit can be a separate PR so they can't be squashed)
(When I mention it above, it's for a case like "the branch is several months old, two versions have been released in that time, and I'm revisiting it to finish it up now", and in that case you just merge the latest version tag (tag, not branch) before continuing. It's a very specific thing.)
The main reason I see in practice that people do something, rather than nothing, is because someone is insisting on clicking the button on Github's web interface to merge a topic, which insists the merge be trivial. Faced with this sort of (bad) maintainership, the topic author has to choose between rebasing the topic, losing some useful graph properties (independence of unrelated branches) in the process, or merging upstream just before upstream merges them, which is incredibly stupid-looking (and very annoying when reading history, nothing is in the place you look for it).
Usually they choose the former, or have it chosen for them. And now the branch, topologically, includes a bunch of unrelated things, and if you have any sort of triaging integration, QA, maintenance versions, or just want to read the graph to see what topic depends on what other topic, this is broken on this no-longer-independent topic. (In very bad cases, you might need to acquire a lock on a human, precisely what the whole system is meant to avoid.)
The best thing is just to do nothing. If you're using Git as originally intended, your branch became patches anyway, it doesn't have a history. (As the maintainer receiving them, you just apply the patches to the latest version tag, pretty much, unless it's in series with something else, or it's a bugfix (always apply a bugfix commit directly onto the commit that introduced the bug. it can then be merged into any maintenance branch or other branch containing the bug. the word "cherry-pick" need never cross your keyboard.)) But a bad webapp broke things for a lot of people, and it was worked around in various ways.
as a maintainer, why should i have to make the effort of merging a contribution myself, instead of asking the contributor to make sure that their changes can be cleanly merged?
i agree that merging upstream before they merge me seems odd, but that suggests that rebasing is the only way to go. or how else can i ensure that my code can be merged cleanly without creating extra work for the maintainer?
As a topic author, you have no such responsibility to deal with arbitrary other topics you had nothing to do with; this holds in an opportunistic open source contribution, in a workplace, anywhere. There's someone who reads all the patches and has a working knowledge of them: the maintainer. And if they ask you to do their job for them (which you will likely do poorly, not being the maintainer), they are a _bad_ maintainer. (And if software encourages this it's bad software; I'm not attributing this to maintainer malice or something, most people are just taking their cues from common bad software like Github's website.)
> why should i merge a tag and not the branch head?
Because a branch head is an arbitrary collection of things and a version is a specific supported collection of things.
as a FOSS contributor, making the work of a maintainer easier can mean the difference between my submission being accepted or rejected. if i want to push a change upstream it is therefore in my interest to not rely on the maintainer to do that work. especially not if that maintainer is a volunteer. not every contributor is going to do that, nor should every contributor do that. only those that have sufficient experience to work at the level of the maintainer, but simply are not the designated maintainer because it is not their project.
the same holds true at work. i expect everyone in a team to be familiar with the whole project and be capable of doing the work of the maintainer, at least everyone at senior level. whether they end up doing it or not. having a single person be responsible for all merges is an idea that i do not agree with. among other reasons it causes the designated maintainer to become a gatekeeper preventing others from developing that skill and step up to do that kind of work. whether it is practical for the maintainer to do all the integration then is a matter of their workload.
the suggestion that only the designated maintainer can do a good job at merging is even somewhat patronizing.
on tag vs branch, this seems to me only makes sense if we are developing against a release branch. a development branch often does not have release tags. the "releases" in a development branch are the merges of new feature branches. if each feature is merged on completion then there should not be any work in progress commits anyways. even if rebasing is used, those commits will be pushed all at once, so HEAD is always at a completed feature.
i think that these are simply different approaches to development, and it can't be said that one is better than the other. it comes down to preference. i favor an environment where everyone can step up to contribute to the extent of their experience and knowledge over an environment where my capacity is limited by my role.
being told that i am not capable of doing something because it is not part of my role is not something i want to hear ever.
The information is all there. You can grep it to find specific commits.
This isn't a hill worth dying on with your team if it's their MO already.
This is putting lipstick on a pig
Which can be fine, but let’s not pretend there is a perfect approach, needs differ
I think it's a natural instinct to want to protect your production environment so I'm sure others have made the same mistake as I did. Main is my default branch, I don't want anything pushed into the default branch to result in a deployment to production.
So naturally I create a new branch called production. The relationship makes sense to me because we develop in main, might even deploy to staging from main, but it's not until we feel we're ready that we merge main with production. And to a user of git this requires a manual step where you explicitly specify the word production, git checkout production.
But that resulted in merge hell until I learned I could just rebase main onto production.
And the command structure is even exactly the same, standing in production branch I either do git merge main, or git rebase main.
This comment was a message to my younger self.
Did you mean rebase production onto main rather? Or do I have my branches and trees mixed up.
git checkout production
git rebase main
that gets everything from main into production.
git rebase (current branch onto) main
However, if there are conflicts, I've always found them to be far easier to resolve during a merge than a rebase. That's not to say it's always easy, but at least with merge you only need to resolve them once. With rebase, you resolve the conflict, then have to resolve it again when git tries to apply the next change, and so on. Sometimes you even need to resolve conflicts bacwards if you're reordering commits. Much of this can be mitigated with git rerere, but if you're averse to merge, you probably don't want to learn about rerere either.
git checkout production
git pull --ff-only main production
git push
(at which point CI auto-deploys 'production')(edit: have you got "rebase main onto production" the wrong way round?)
Does hacker News fudge timestamps on comments when it boosts a post?
When you're in the release process, the branch won't match what's currently deployed, but that's ok. The point of a production branch is not to indicate what is on production at this instant, but to be a record of what was deployed to production or at least was intended to be, for changes that get canceled before deployment.
Also removal of code smells. And actual smells.
/I love the Tube. So damned useful.
I honestly hate Git compared to SVN. The Git tool seems technically better. However, the easier aspects of the tool seems to have promoted less planning, more stepping on each other's toes, and what I see as non-value added work (merges, rebaseing, retesting, etc).
But I probably just work for a really shitty company.
I also recommend it to git newbies, as a tool to understand the state of their repo when something has gone wrong (e.g. they did a bad merge or rebase).
If I make a variable in commit A and think of a better name for it in commit C, why wouldn't I use rebase to squash C into A? Some sense of purity of history?
Or more directly related to the rebase vs. merge debate, if I fix an issue across three functions, and in the meantime someone has removed one of those functions on the main branch, rebasing eliminates the "history" of me fixing that removed function and I think that's good. It makes my commit simpler.
Rebase can certainly be used to simplify the history too much, but that will always be a judgement call. That shouldn't keep us from editing our branches in ways that are clarifying instead of confusing.
We never capture the full complexity of what we went through writing code in our source control. It would be bad if we did.
The issue with is with rebasing multiple commits onto master instead of merging - the intermediate commits were never code that anyone ever actually wrote or tested, so any issues with them stem purely from the fakeness of the history.
If you have commit A then you write a branch A -> B -> C while someone else writes A -> D and they merge first, rebasing to get A -> D -> B' -> C' means B' was never something you wrote or tested. This code never existed on anyone's machine or had CI run on it before the rebase.
Does it really make sense to run CI on all of the new intermediate commits that rebase invented? What if some of those fail, are you really going to go through and fix tests for fake intermediate commits?
The solution for me has always been to squash. Inventing fake history is totally pointless and counterproductive. If you want cleanliness and bisectability via destroying the "real" history, just go ahead and really destroy it, don't invent a fake, possibly broken history.
Once it's satisfied that your branch can be merged, it runs a subset of the tests, and throws an error if they fail. This way, even if you do rebase your branch, its latest commit will still be tested. (Having intermediate commits pass tests is encouraged but not required.) Finally, it regularly takes groups of 8 or so accepted PRs, tries merging them all in sequence, and runs the full test suite on the result. If it succeeds, the merge commits are pushed to master; if not, a human operator gets it to try again without the offending PR.
By your terminology, I suppose this would count as running CI on all the "invented" commits, and forcing PR authors to fix all their tests. But in practice, it's not too odious, since most PRs don't conflict (unless you're touching half the codebase), and any test failures from a non-conflicting change will get caught by the merge step.
article's proposed solution: simplify the git history by destroying information that are irrelevant. That's what rebase is.
The problem is what is or isn't relevant depends on context. I think the right way to go about it is to simplify the presentation and otherwise improve tooling to get information out of the git history. git already has some ways to filter the history, but it lacks very feature rich query language, like mercurial. git guis should also step their game up.
git log --first-parent gets you there.
> Then be able to zoom in futher if necessary to see how an MR was arrived at.
Yeah, an interactive UI would be nice for this, maybe there are some, but I really only use the CLI.
git log --oneline # `gl`, to see "linear" history with branch annotations
git log --oneline --merges # `glm`, to see merge commits only
git log --oneline --graph # `glg` see the train tracks too, if I need to
And I know that git log has me covered, if/when I need to narrow / slice history more.- in your git client: enable "only follow first parent" for the base branch which only shows merge commits and only open merge commit history as needed/ when relevant, i'm only aware of fork.app doing this nicely but should be an option in more git clients
- hide all branches by default and only show your base branch and your own work, just let go of caring what others are doing, the only way to interact with other branches should be a review system at which point you can selectively pull in the relevant commits as needed to try locally. (exception is maybe a technical manager or team lead who should of course care somewhat if branches get abandoned or not cleaned up properly, but this is a completely separate workflow)
- Look at saplings notion of public vs draft commits, this is a game changer and works just as well for git, just not supported by the tooling in the same way. Don't try to argue and adopt workflows that are the same for both, they are completely different phases of work with completely different needs. In a nutshell: a) public commits are the ones that were pushed to a branch that is shared with anyone else. This maybe main, a feature branch worked on by multiple colleagues etc. These commits are considered immutable, never amended or rebased so the only option to bring them up to date with another branch is a merge. b) BUT every commit that is not part of one of these branches is considered a "draft" these are never merged into, always amended or modified and mutable draft state. It is considered rude to expose your internal working struggles like merging in master 100 times a day to avoid big merge conflict buildup or reverting changes into your work submitted to review, you submit a clean stack of pull requests, one per reviewable batch of changes
- rebasing and amending has a few other big advantages. One of them is that syncing to the base branch does not pollute history and can be done as often as possible without a down side. This allows keeping all work in sync without buildup of big hard to resolve conflicts. There is also tooling to do this for all your open work at once and to resolve conflicts smarter (.eg git rerere)
- merges should be set to allow only squash + rebase as this works the same for devs working with clean commits as well as devs keeping their commit history in their branches. this way at least the base branch is mostly clean
A lot of the merge commit was triggered when you pull and merge instead of rebasing locally.
I have found an almost security hole with it in the dropping of a BC layer which just no one would've expected to do that. And because it was all deleted code tracking down the origin bug any other way is fair impossible.
It seems to me that, with respect to bisecting, a merge workflow with git bisect --first-parent is equivalent to a squash workflow with a bare git bisect. Am I missing some way in which that's not the case?
One issue I've seen a few times is that some commit in the middle of a topic branch is the problem, but if you didn't rebase it then that commit itself would look fine on top of the topic branch's parent. However, after merging it's now also on top of other commits, and the interaction with those was the problem. That makes it very hard to find such a problem. Rebasing the history of the topic branch before merging will make finding it much easier.
I just find that "always squash the entire branch" is a common reaction to "history is messy" and I wanted to surface that (per my understanding) it doesn't actually improve the situation (vis-a-vis bisect in particular, assuming you're passing the correct arguments for your situation) over merging (no-ff, I neglected to specify...) branches where some of the commits do not build.
It could be a bit more visible somehow though, I get the sentiment. Maybe it's more of an add-on to git's role though, at least without plenty else also becoming more visible/GUI-like too.
The bigger deal is that things are never shared simply for being in the reflog - which is probably correct for its intended use but doesn't really fit what's asked for up thread.
That's exactly what rebase does.
(OJFord said that too, but buried the lede slightly, so I thought it worth saying in a single sentence.)
First, merge only ever allows you to arrive at a single commit, so it's strictly less powerful than rebase. With rebase, you could start with a sequence of three commits, "A", "B", and "squash! A", and turn that into two commits, "A'" and "B'". Merge doesn't let you do that.
Second, sometimes a merge commit really is semantically useful and you want it to be shown as a merge with two parents. There is no canonical way to distinguish this kind of merge from the kind of merge you seem to be thinking of.
Personally, I think the way to resolve this would be to have optional "squash/cherry-pick parent" metadata on a commit, so that commits that result from a cherry-pick or a rebase can point back to the original commit(s) in a more structured way (remember that a rebase is really just a sequence of cherry-picks). This metadata could also be used to preserve `git commit --amend` version history. Augment it with a "reverse diff" bit and it can be used to track reverts as well.
Merge is useful when there are multiple "canonical" repos which occasionally merge changes in between each other. Think web of trust vs central authority.
I think a lot of the disagreement about this is really people talking about different things. Some people say "always squash; nobody cares about the trivial typo fix commits and whatnot" and other people say "never squash; you lose important history" and really they're both right... you should squash when it's not important to preserve the history. Obviously people are going to disagree about when that is but in my experience if a PR is big enough that you think it should be more than one commit then it's too big, unless it's a big feature branch that has been worked on by multiple authors.
Similarly with rebase vs merge, if it's a small single author PR then definitely rebase. For big feature branches you may want to use merges though I would still suggest rebase is better. You just need to make sure everyone is using the safety flags when they force push.
Rebase workflows are awful and unintuitive. Leave the rebasing to the git wizards who actually know what they're doing, in no circumstances should this be part of your day-to-day work.
(Ironically, I've found that this style of development makes it less likely for bugs to be introduced in the first place.)
If the extra “noise” bothers you you can use —-first parent with git log.
I absolutely do. Every lead I've ever worked with absolutely does. If you merge an unnecessarily big commit to master, you are potentially making life very difficult in the future.
If you make a meandering series of commits during development, squash them for the PR. But please, please do not squash the entire feature upon merge.
Also, I'm not sure how rebase could be confusing except in the case where there are multiple commits with big conflicts, but that's rare, and you can make exceptions if the developer is really that unsure with git in those rare cases (or ask for help).
> If you make a meandering series of commits during development, squash them for the PR. But please, please do not squash the entire feature upon merge.
I typically turn my PRs into a series of buildable and testable commits. I also put a lot of work into making those commits tell a useful story to the reviewer and to anyone doing git blame later, and squashing them all into one commit undoes that work.
Does your CI test each individual commit? Afaik most of them only test the top commit. How do you know/enforce that all the inbetween commits also build/pass tests?
How do you know how many commits to revert, if you need to revert the feature? Instead of reverting 1, now you have to revert N where N is not recorded anywhere.
Because they can explain their individual rationales, while still making the most sense to merge all together.
> What can you possibly learn from that that you can't get from looking at the entire diff the PR is introducing?
Ease of review (both before merge, and in the future when wondering why something was done a certain way). Saves me as a reviewer from having to guess which parts of the commit are meant to do what.
This, of course, means that every commit needs to be a reasonable change in itself -- fixup commits done while developing should be squashed into the original change with a local rebase (these are the "meandering" commits your parent post mentioned).
> How do you know/enforce that all the inbetween commits also build/pass tests?
I'm no CI expert, but I would hope that most systems allow this as an option.
> How do you know how many commits to revert, if you need to revert the feature? Instead of reverting 1, now you have to revert N where N is not recorded anywhere.
It's recorded in the merge, assuming you always make a merge commit.
Otherwise, since each commit is actually its own logical change, you figure it out the same way as you would figure it out in the "squash PR" model -- bisect to find it, then see if reverting it helps.
Hard disagree. I hardly consider myself a "rebase wizard," but I've been a near-exclusive practitioner of rebase workflows since I can remember. I find rebasing much more intuitive than workflows with merge commits. Squash merges are fine, but with proper intuition, they appear like a special case of rebase.
In my experience, the resistance to rebasing comes down to fears about "rewriting history" and false intuitions about how git works. I usually allay the former by pointing out that squash merges - which almost everyone approves of - also rewrite history. The latter issue seems to arise from arrows in popular git visualizations pointing in the wrong direction, e.g. in Gitflow.[0] In git, the child commit points to its parent, because each node is immutable. The git data structures are extremely simple (hence why git is so named), consisting of blobs, trees, commits, tags and references. Once you understand how these work in practice, rebase becomes intuitive.[1][2]
IMO, the only thing unintuitive about git is the CLI. Translating the graph operation I want into the commands is sometimes a challenge. Maybe that makes me a (frustrated) wizard after all?
[0] https://nvie.com/img/git-model@2x.png
[1] https://speakerdeck.com/pbhogan/power-your-workflow-with-git...
[2] https://eagain.net/articles/git-for-computer-scientists/
There is 'git bisect' which only gives you detailed answers if your commit history is as fine-grained as possible while still being compileable and testable. Also, changing commit IDs on branches that are visible to others are a problem.
Which is why my approach is as follows: Work on private branches, one per feature/fix/..., then do interactive rebase to create a compilable and testable patch series out of those onto a for-review branch. While developing, rebasing onto whatever public branch you want to merge with next is of course OK and necessary. When you think the feature in the for-review branch is ready, do a merge-request into the public target (mostly devel) branch as usual, and either merge or rebase, whatever you like best.
But: Whatever branches are public are only ever merged. Never squash-merged, never rebased, never force-pushed, never filter-branched (except maybe after a court order). Because all commits on those branches are necessary to trace what people were doing with the code. All those commits and their relations are necessary for git-blame, bisect, sloccount and other things. Any commit ID there could wind up in some test binary, release ID or stuff, and you absolutely truly need those later on. And while a simple rebase might keep some of the necessary details intact (the other branch will still be there after all), only a plain merge will also preserve all the relations between the branches properly.
This sounds like the refrain of someone who doesn't want to actually learn how one of the most fundamental tools of their profession works. I realize that git wasn't the optimal choice for the industry to settle on, but it's what we picked, and simply avoiding a feature that a majority (65%) of your peers use to at least some degree [0] will hamper your professional development.
Learn git. It's not pretty, but it's what we've got, and it's not going anywhere anytime soon.
The reality is that most people work on their private branches, alone. In that case it makes almost no difference if you rebase or merge. In almost any other scenario, merging still works as expected, while rebasing without understanding git will almost definitely lead to losing work and spending an absurd amount of time resolving conflicts. Why would you want to inflict that on yourself? Just use the approach that always works instead.
Learning Git really isn't high on the priority list for most developers, as they know what they need to use to get stuff done. The complexity gets really high really fast, so it's quite understandable why most people treat Git like DNS or other infrastructure - it's there, I know the basics to get stuff done and if anything goes wrong I ask an expert to take a look. And guess what? There is NOTHING wrong with that.
I'm well aware about the differing meaning of HEAD in merge and rebase, and if I had to think about it, would probably realize that it makes sense to always display HEAD first. And that as a consequence, the order would be swapped.
But I would definitely have answered "no" to the question as written.
I've saved you a click. TFA has nothing to do with TFL beyond this line.
- a public change history (e.g. commit history of master, dev-1.0, whatever)
- a personal change history, especially useful when prototyping a fix or feature on some independent branch.
you might even call this independent branch "master" but it's your personal version of master in a clone on your laptop. it's a time machine, you can go backward/forward. you can make a change and then reverse course. you can commit every work-in-progress that compiles as you work get it running or passing tests. whatever you like. you can pull or merge or rebase as you wish.
when submitting it to the public "master" branch, you are effectively saying "i want to commit this package of changes on top of the master branch.
do you really want to see "fhgiry pulled from master into fix-splork-feature on date x" and https://xkcd.com/1597/ in your public history of master? every branch you merged from or rebased on and every failed attempt and spelling mistake and WIP and addressing of review comments?
or do we want to see a series of commits that just record the end result in a series of nice little patches that simply add a finished changes like: "splork: fix canoe paddle explosion issue by decreasing default gamma"
the common ubiquitous guidance to "never rewrite history" is nearly always valid for the public history.
converting a messy personal history into a neat series of patches/changes/commits that apply cleanly onto the latest public master should not fall under that same guidance. i'd say it's closer to "never rewrite public history"