Git rebase, what can go wrong
jvns.ca
jvns.ca
> The golden rule of rebasing
> Once you understand what rebasing is, the most important thing to learn is when not to do it. The golden rule of git rebase is to never use it on public branches.
https://www.atlassian.com/git/tutorials/merging-vs-rebasing#...
For me, even though rebasing comes with some trappings, I still greatly prefer it to the alternative, which is to have merge commits cluttering up the commit history.
> Don't rewrite history on shared branches with proper communication.
I don't teach "never", I don't teach that `main` is special, I don't teach that force pushing is forbidden, because I don't believe in those things.
I highly prefer a rebase-heavy workflow. In addition to not "cluttering" the history, it's an invaluable tool to keep commits focused on "the right level" of atomic changes.
It's only becoming tricky if the MR has been rebased onto a different base in the process, but it's not very hard to deal with that too if needed (just annoying).
Git history should tell a simple, understandable story of each change. For example: 1) refactor existing code, 2) add feature. Or 1) add missing tests, 2) refactor existing code, 3) add feature.
But since you're working on the fly with imperfect knowledge, it doesn't happen in such neat steps. Refactorings and behavior changes end up interleaved in your raw git history, so you need to do a little bit of cleanup by hand in order to present a simple story in the commit log.
Of course if you have developers that don't do that and instead merge dozens of commits that just say wip, wip, wip, lol, fml, wip, wip, lol, yolo and you can't fire them or get them to change, then squash merges ftw.
I prefer one commit to main per feature, a long with a good description on the GitHub PR.
Sometimes I’ll branch out from a feature branch for the occasional and infamous ‘get CI working’ round of 10 one-line commits though, to not make it too muddy.
Several reasons:
* facilitates much better code review discussions
* enables use of git bisect to locate bugs
* allows for informative commit messages associated with the changes
* communicates clearly to future self about why changes were madeAgreed, it has been standard at most shops I've worked in the past 8-10 years.
Squash-merge is a scourge. I've seen squash merged commits 30 lines long ("try 15", empty line, "try 14", empty line...). I'm not even sure if you can do anything about such commits because squash-merge is a github/gitlab thing. So, I'm not sure if there are hooks to block it via a commit message linter.
And I've seen people going through some intense mental gymnastics to justify avoiding squashing locally, writing a proper commit message and then merging.
> Of course if you have developers that don't do that and instead merge dozens of commits that just say wip, wip, wip, lol, fml, wip, wip, lol, yolo and you can't fire them or get them to change, then squash merges ftw.
Yes, any large organization has plenty of devs who all have their own style and preferences, for better or worse.
Whoever demands they all bend to the one true way is a fascist (lol not really but you know).
Just set up your CI/CD in such a way that PRs with weird git logs get squashed into one pretty message, preferably the PR description since other devs have to review the PR it's often given more effort. Set it up so that if things weren't formatted the "right way" they get auto-formatted or a test fails and the dev says "ah, I have to run that one task and then update the PR".
I don't think a big organization is going to scale with developer "evangelists" demanding people write their commits a certain way either.
conventionalcommits.org was the worst. I worked at one "big co" that tried to get devs to do this. Even after we had been doing it for a while, nobody ever went back to look at the history in such a way that it was worth it. We ended up throwing in the towel rather than the company trying to get all other teams to do it.
Ain't nobody got time for that shit.
This is too much thought put into a VCS. I don’t want to have to think about my VCS at all beyond the commit message. For all of Git’s popularity, I’ve never seen benefits that justify the absurd amount of work and knowledge it takes to perform simple actions. It’s the VCS equivalent of Scheme or emacs.
If you need to do a workaround, or a complicated feature sometime it's nice to explain it as a comment in the code, but sometime it's better to put it as a comment inside the commit message. But if it's all merged i the end, along with lots of template changes, README changes, refactor irrelevant to the current changes, then you're losing an important way of navigating a codebase.
Fine as an opinion.
> For all of Git’s popularity, I’ve never seen benefits that justify the absurd amount of work and knowledge it takes to perform simple actions. It’s the VCS equivalent of Scheme or emacs.
This is just wrong. When people talk about this, it's not about git at all.
You write some code, and you make it into commits. Part of that is choosing if/how to organize it with multiple commits, and how much effort you want to put into that. This is fundamental to using a VCS, any VCS.
Or by analogy, if a lot of emacs users complain about your spelling, that's not because emacs is overly demanding.
Got fired. Kinda. I was laid off.
I'm a big proponent of rebase and squash if it helps to make a commit more coherent, but we use squash merges by default in the current project I'm working on, and I die a little bit each time I try to understand what changes were related to a line when tracking down a bug.
The commit information I see when telling teams to squash their branches on merge is not valuable.
* "fixing whitespace" * "incorporate review comments" * "fix broken test" * "fix other broken test"
(note, the broken tests were broken by the changes in the PR)
As soon as that PR is merged those commits are worthless. And there are branches with dozens of those "fixing X" commits that would otherwise pollute the commit graph.
Things like this should not be standalone commits though, they should be incorporated into the previous branch by amending the original work. It takes some effort to have a useful git history, it does not just happen on its own.
A change/feature/bug is a branch, which is squashed into a commit on your main branch, right? So your main branch should be a linear history of changes, one change per commit.
How does that impact the ability to git blame?
As a simple example, I recently needed to update a json document that was a list of objects. I needed to add a new key/value to each object. The document had been hand edited over the years and had never been auto-formatted. My PR ended up being three commits:
1. Reformat the document with jq. Commit title explains it's a simple reformat of the document and that the next commit will add `.git-blame-ignore-revs` so that the history of the document isn't lost in `git blame` view.
2. Add `.git-blame-ignore-revs` with the commit ID of (1).
3. Finally, add the new key/value to each object.
The PR then explains that a new key/value has been added, mentions that the document was reformatted through `jq` as part of the work, a recommends that the reviewer step through the commits to ignore the mechanical change made by (1).
A followup PR added a pre-commit CI step to keep the document properly linted in the future.
I would argue that those are by far the minority of PRs that I see. As I mentioned in another comment, _most_ PRs that I see have a ton of intermediary commits that are only useful for that branch/PR/review process (fixing tests, whitespace, etc). Generally the advice I give teams is, "squash by default" and then figure out where the exceptions to that rule are. That's mainly because, in my opinion, the downsides of a noisy commit graph filled with "addressing review comments" (or whatever) commits are a much bigger/frequent issue than the benefits you talk about. It really depends on the team.
EDIT: On top of that, there's usually a bit of 'related' work you need for a task, by example when you find an edge case related to your feature, and now you also needed to fix a bug, or you did a bit of refactoring on a related service, or needed to change the data on a badly formatted JSON file.
Unbeknownst to you, you added a bug when refactoring the related service, a bug that is spotted a few months after, only on a very specific edge case. If the cause is not obvious, you might want to reach for git bisect, but that won't be very useful now that everything I've talked about is squashed into a single commit.
I agree that's related work, but I'd argue that work doesn't belong in that branch. If you find a bug in the process of implementing a feature, create a bugfix branch that is merged separately. If you need to refactor a service, that's also a separate branch/PR.
That's actually the most common pushback I get from people when I talk about squashing. They say "but then a bunch of unrelated changes will be lumped together in the same commit", to which I respond, "why are a bunch of unrelated changes in the same branch/PR?"
When my branch is up to date with `main` I can build an artifact, fast forward merge that branch into `main` and RETAIN the artifact, and merely update its tags to mark it as `merged` in.
With a squash I lose that information.
Now, GitHub does not allow me to do a fast-forward merge but I can still trace the 2 commits that are the parent of the resultant merge, and find the artifact based on that, and retag.
I've seen people forget to go back more than one commit and then blame the person who last indented a file instead of going back to the commit that actually wrote the code many times.
I default my `git blame` to `git blame -w` which ignores whitespace commits. Though knowing how to jump back commits should be required knowledge.
That’s done by an automated tool. Correction of indentation is just a byproduct.
I don’t consider that “tampering”.
$ git log --merges
Now you can see your features in a nice history and also have added benefit of seeing intermediary commits. Pro tip: merge commits aren't required to use the canned "Merge branch into..." message, you can give it any message you want, such as "feat: ..." or whatever your convention is.I hate that branch squashing has become something of a defacto. I actually do rewrite my history and often add context to my commits. `git blame` can be an incredibly useful tool to get context about a given small change. Getting a massive diff for a whole feature is much less so, especially since you can just look at the diff of the merge commit.
can you help me understand this? It is the exact opposite of my experience. The flow I see is: bug reported, write a git bisect test, identify the feature that introduced it, reach out to that developer/team.
This is allowed by squash merges. When I've seen these more "clean" histories, they have commit points that wont even compile or have runnable tests causing git bisect to fail.
> branches that are squash merged were big
it must be this - how big are your merges? All the projects I've worked on strive for smaller PRs. Large PRs are usually broken up into smaller pieces. Large PRs are an anti-pattern.
> how big are your merges? […] Large PRs are an anti-pattern.
Depends, but they sometimes on occasion can get pretty big, if there’s a bit refactor and/or multiple people in the branch. Small enough PRs are a nice goal - it’s a goal that might agree with and exist in part because squash merges on large PRs lose too much. It’s just the real world routinely gets in the way. It’s very easy for someone who needs to do an ‘atomic’ refactor to touch a ton of files. It’s very easy for a planned feature to end up way bigger than intended. You can’t always keep PRs small or enforce it on other people. Sometimes stuff happens, and when it does, sometimes squash merging feels less good than merging a branch with multiple commits. The good news is that it’s always optional. The bad news is that I can’t necessarily babysit or dictate what others do, and some people prefer squash-merging to spending any time doing cleanup on a messy branch.
And for that matter, you'd manually do a squash with the interactive rebase tool anyway ("git rebase -i").
There's a few special cases that have their own names, a common one is when you amend a commit - to do that manually you'd make a new commit, then use interactive rebase to squash the two commits together into a new one (or, use the "fixup" command available in that tool, which is a squash that automatically picks the first commit message instead of asking for a new one).
Squash merges will squash a whole branch into a single commit, rebasing it onto the target in the process, and then fast-forward the target to the new commit. It's a tightly controlled use of rebase, and can be thought of a bit like how "for", "foreach", and "while" loops are a tightly controlled use of "goto", an abstraction built on top of a far more flexible tool.
They do but they have their own issues. e.g. having to delete local branches using git branch -D instead of git branch -d and getting the protection from deleting unmerged work.
I still agree that on balance annoyances like that might still be worth putting up with for larger teams with mixed skill levels.
That leaves you prone to losing work if you have a false start that you need to back out of. I prefer to commit early and often on my private branches, then before submitting a pull request I clean up the history to where there are a few good commits that form useful, standalone chunks (ideally the test suite fully passes on each commit).
Hasn't happened to me in over 20 years of using version control. I always keep moving forward, there's really never been a need to go back to a previous commit that hitting crtl-Z wouldn't accomplish just the same. If I wanted to try a new direction I'd just clone the repo again and do the work there. Littering the git history with dozens of superfluous commits just seems pointless. Having to stop and think about writing a commit comment is also just a waste of time - in aggregate it wastes a lot of time. It adds a lot of churn to a workflow for something that may never really be of any value.
This is where the final rebase comes in—you should be combining all the small commits into one.
> Having to stop and think about writing a commit comment is also just a waste of time
Most of my commits when I'm working like this are named "draft". The names don't matter when you're going to redo the history later.
> I always keep moving forward, there's really never been a need to go back to a previous commit that hitting crtl-Z wouldn't accomplish just the same.
You've never started down one path for solving a subproblem only to realize 30 minutes in that it's not going to work?
Sorry, but I'm a software engineer, not a git engineer, and the less I have to do with git, the better. KISS applies to git, too. A simple thing like not creating a commit for every stupid thing keeps the history clean, doesn't bog down the developer by requiring to think about writing a commit message every 2 minutes, and keeps git simple.
>Most of my commits when I'm working like this are named "draft". The names don't matter when you're going to redo the history later.
But then what value have you added by naming everything "draft" and creating a commit? There is no value in doing this.
>You've never started down one path for solving a subproblem only to realize 30 minutes in that it's not going to work?
Sure I have, but I don't need to enter it into the git logs. I'll either start over in a clone of the repo if I want to save the bad work for whatever reason (which is very unlikely), or I'll just stash the work, or whatever. The thing I don't need to do is commit the bad work.
The purpose of history is to remember. Rewriting history, whether git or in life, is bad; outside of the context of don't use it on public repos. Such advice is similar to saying, only point the shotgun away from you when firing. If you have to remember such a rule, it's best to avoid it.
I've heard this many times before, but haven't been able to figure out why this is a problem. In your workflow is it a problem to have a cluttered commit history? If so, could you explain how?
GitHub recently added a feature that prompts people to update their branches via merge. It's frustrating because every PR now had dozens of merge commits polluting the history.
What I want is for GitHub to track changes between sets of commits in a PR so that you can do most of the review with merges and "address review comments" commits, and then rebase into well organized, logical commits and review that those have the same diff as the messy history after a force push.
It wouldn't be a problem if people took the time to organize the history prior to merging as you said, but most people don't do this.
What matters is that you end up with working systems. That a lot of change happened is just, well, what happened. It doesn't need to be prettied up and made to look like your development occurred in a clockwork march of cleanliness. It literally does not matter unless you spend a lot of time doing git-bisect.
Let it go. Accept that coding is not a smooth, robotic, endeavour, where everything is always tidy. And that's just fine.
I've been working on dozens of projects since, and probably did thousands of commits. Some of the teams of those projects included dozens of developers working concurrently on the same codebases. We always merged the upstream branches into our development branches and never did any rebases.
I have NEVER ended up in a situation where I thought rebases would have been better. The git tools and IDE integrations of our current age allow me to find any information I need from the history without pain.
Instead of letting it go, maybe we should have more discipline and organization in our lives and not less.
The pro-revisionists (squash, rebase) say they do what they do so the history looks clean (no intermediate commits breaking stuff, a "straight line" graph, etc)
The anti-revisionists say they do what they do so the history looks clean (can see the actual development, can safely diff different commits to see what changed in between, see the log in chronological order, etc).
> Instead of letting it go, maybe we should have more discipline and organization in our lives and not less.
Again, both sides could argue that they're the ones with more discipline.
> The point is to make it possible to debug later, via bisect, or show, or even just a diff.
This sounds anti-revisionist.
> The point is to make the workspace clean for the next guy.
This is one of the most common pro-revisionist arguments.
> This sounds anti-revisionist.
That’s not how I see it. What makes debugging via bisecting easier is self-contained changes, not exactly chronological changes where you temporarily broke stuff and then fixed it before submitting your PR.
And git blame. And git checkout to a past state. It "doesn't matter" only if ease of understanding your project history doesn't matter.
Frequently, for any long and complex project. Large amounts were written by people no longer working on it, and the history of how things came to be can help fill in documentation gaps and make intent clear.
By "frequently" I mean something like "I check history for about 2/3rds of bug fixes, and 1/4 of adding features" to understand the surroundings better, when writing or reviewing. Anything that makes that better saves me hours per week.
It catches and prevents more than enough subtle issues to be worth the effort.
It's others history that I'm usually interested in. I can easy follow the small diffs of individual commits, but have a much harder time grokking a wall of red and green.
I've used git bisect the few times I've had to diagnose an issue that wasn't detected immediately, which gets you down to the exact granular commit.
That being said, after reading this stuff, I may start using it on my local branches to clean up multiple commits into one tidy one, but that's about it.
And a well-organized commit can also tell you the “why.”
Every time I try to review a PR and the bookmark resets because they decided to force push I curse the Git maintainers that don't have the backbone to just get rid of rebase.
The amount of time that has been saved in my life by someone leaving an explanation in their commit (for some weird edge case or context I’d have no way of gleaning because they’ve since left the company) is SO much more than the extra time I’ve put in to make sure the history has this extra info in it.
If I had a bad day and introduced something stupid, I want a bisect to point me a the code I wrote on that bad day. If you squash liberally, perhaps because you want each commit to correspond with a release-note, you're going to lose that debugging granulariry.
But it's simply intangible. My instinct tells me that it's helpful and that's okay. I don't owe anyone a justification for how I organize things, and there's nothing controversial about this. (Or maybe I could even come up with a logical example of a benefit, but that's a trap I'm not going to fall into) And a lot of people agree, and they know what I mean, so it's not merely an individual preference. If I have to work with someone who has strong preference against it I'll worry at that point about negotiating.
Very genuinely: I do not care at all whether you clean your room starting from left and continuing to right. Or, starting from doors and continuing toward window. Or whether you clean it in a random order. I also do not care about whether you clean every Friday or whenever you feel like. That is the equivalent of git history. Because this excessive care about git history is just that - insisting that room is cleaned from left to right as if any other order was an issue.
The reason why it is hard to defend the tangible benefits of this or that git history strategy is that there are very little benefits.
But you didn't say that you don't want it clean. It sounds like you're talking about how it's organized rather than whether it's organized.
> The reason why it is hard to defend
I'm talking about intangible benefits and no that's not the reason. Intangible benefits are inherently difficult to defend in words. Citing this as evidence of anything is akin to a debater's trick.
In the rare situation when I have to read it, I am perfectly ok looking at previous commit too or whatever. It is still less overall work then what people describe in here.
Even with room, I do not want my room infinitely clean. I am ok when books are not ordered by height and color for example. I do not need t-shirst ordered by color either.
I prefer not to have squash commits in our team for this reason. It makes master look good, but usually nobody ever looks at the master commit history first, they look at the merged pull requests. However, everybody must look at the commits you made in a pull request. If you have squash commits, you are encouraged to have messy commit history in your pull requests, leading to meaningless commit messages and even large commits (causing other problems...).
IMO the only advantage of squashing is that it makes it easy to roll forward when you accidentally deploy something that causes problems.
* use filtering commands like "git log -S"
* press the "annotate" button in my IDE and can see which commit introduced each line
* run "git bisect"
* use "tig" to drill down through the history of a file (shortcut "," is "move to commit preceding current line's blame commit")
...every step of the way, I get a meaningful description of why a change was made and what other diffs were necessary to achieve that change. And not just "fix", "bug", "PR commments".
* `git blame --first-parent`
* `git bisect --first-parent`
* At least one "tig-like" with a --first-parent first UI: https://github.com/kalkin/git-log-viewer
In PyCharm, I can see which commit introduced each line, regardless of branching. Same with drilling down through a files history. Is this an IDE limitation you're seeing?
> every step of the way, I get a meaningful description of why
Isn't this more about commit messages, than anything else?
> Is this an IDE limitation you're seeing?
I'm using Jetbrains too.
Makes reviewing a set of changes prior to a merge much easier. It's nice if there's a 1:1 correlation between a commit message and the actual patch contents.
Im sure you've dealt with the case of reviewing a colleague's changes with a commit message like "Enable logging in foobar module" and the patch is actually enabling foobar logging and a bunch of other stuff.
This makes bisecting your git history to identify and fix bugs much more difficult.
If the git history is clean, you can just read the commit messages and implicitly trust the developer if clean git hygiene is in place (as opposed to actually needing to read the whole diff on a per-commit basis to find out what _actually_ happen at commit XYZ, despite it's message).
Every time you do a git rebase, you are literally asking your source control system to lie about history. If you mess up, and you eventually will, you're then forced to manually figure out what the history really was despite being lied to. If you mess it up, well, good luck.
I used to work at a company where someone (we never figured out who) in another group would rebase every few weeks. We didn't find out about it until their stuff was pushed then released. The result was that features which we'd written, QAed, and released to production would simply disappear a few weeks later. With no history suggesting that it ever existed.
Have you ever been pulled off of a project to go fix a project from a month ago which has disappeared from source control? You don't know what happened, you no longer have context, you've just got complaints because your stuff no longer works.
Is your desire for a "clean history" worth potentially creating THAT disaster for other developers on your team???
The true history is not recorded in your normal commits either. Every time you modify your source buffer, that is the true sequence of events. This truth is lost already as you undo/rework things before you commit. You're ALWAYS manipulating and telling a false story of history whether you realize it or not.
Commits are a tool that give stronger backup/undo protections over simple file saves and in-memory editor undo lists. Just because you happened to save your work in a commit doesn't mean it should be instantly be regarded as holy history. Not anymore so than if you simply saved the file.
I think the bar for "holy" history should be whether it is published to a shared branch.
Every single point in the commit history represents an actual state of a repository at a specific point of time, along with the information of which point or points were next before it. This is all part of the true history.
This is not a full history - you don't have every keystroke, abandoned commit, switch between branches and so on. But nothing that you're being told is wrong.
As soon as you do a rebase, you're rewriting history. You're claiming that there were specific points of time with specific states that never actually existed. You're losing information about points of time and specific states that actually existed, which someone once considered important enough to do a git commit over.
The difference becomes important if that someone, which at a previous job was me far more often than I would like, tries to go back to the historical commit. And finds that it is gone without a trace.
Agreed. But the rewrite occurs in your private branch. It's history is just as private as the undo list in your editor. No one cares about what's going on in your editors undo list. And by the same logic they shouldn't care about commits in a private branch.
> You're losing information about points of time and specific states that actually existed
If you avoid rebase, then you end up "rebasing" without rebasing. You "squash" intermediate states by never recording them to begin with.
Failing to record history is not superior to squashing it.
> And finds that it is gone without a trace.
I don't have the details, but it sounds like someone rebased a public branch. Yes that is bad. But it's sort of like saying we shouldn't drive cars because someone chose to drive the wrong direction down a 1 way road.
Committers always have a choice of which of the changes present in their working tree they stage and then commit. The commit history is always a flat approximation of the real evolution of the files in the repo.
BTW rebase produces a new commit ordering, but does not modify the old one.
> You’re claiming that there were specific points of time with specific states that never actually existed.
No. You are asserting intent on the part of git users and git that has never existed, you have misunderstood what git history is. The git history is not a claim that the state at that point existed during development, you are projecting your own goals that are not shared by git or git users.
> You're losing information about points of time and specific states that actually existed, which someone once considered important enough to do a git commit over.
Hehe this is so full of assumption. You write it like I’m rebasing someone else’s work, but you already know I’m only rebasing my own commits, and I’m the one who decides what’s important enough to do a commit over.
I like commit early, commit often. I want to make small incremental commits that don’t display to others that way and I expect to put small commits and fix ups together later into a single useful commit with only one commit message.
This argument reminds me of a scene from Yes Prime Minister [0]:
> Humphrey: The minutes do not record everything that was said at a meeting do they?
> Bernard: Well of course not.
> Humphrey: And people change their minds during a meeting don't they?
> Bernard: Well, yes.
> Humphrey: The actual meeting is a mass of ingredients for you to choose from.
> Bernard: Oh, like cooking?
> Humphrey: No, not like cooking. Better not to use that word in connection with books or minutes. You choose, from a jumble of ill-digested ideas, a version which represents the Prime Minister's views, as he would, on reflection, have liked them to emerge.
> Bernard: But if it's not a true record...
> Humphrey: The purpose of minutes is not to record events it is to protect people. You do not take notes if the Prime Minister say something he did not mean to say, particularly if it contradicts something he has said publicly. You try to improve on what has been said, to put it in a better order. You are tactful.
> Bernard: But how do I justify that?
> Humphrey: You are his servant
> Bernard: Oh, yes.
> Humphrey: A minute is a note for the records and a statement of action if any that was agreed upon.
I think the analogy is pretty clear. A pull request does not record every single little change you made when writing it. You choose, from a jumble of ill-digested ideas, a version which better reflects your intent as you would, on reflection, have liked it to emerge. It doesn't matter that it's not a true record, since its purpose is not to record events but to communicate ideas. You try to improve on the commits as they have been written, to put it in a better order. You are tactful.
But I have never seen a commit disappear while rebasing, ever. That workflow is busted somehow. They were doing it wrong.
You get clean history by not merging branches with 50 intermediary "fiddling with X" commits in them.
Wouldn't this be trivially solvable by git bisecting your deploy branch?
I mean that's the equivalent of reversing your JCB through a house on a building site because "the house was not there a week ago when I last moved the JCB".
or am I missing something?
Alternative DVCSes which support this workflow include: fossil
I recall when I first switched to git at work and the team was insisting on a "linear history", I was bemused: Could these developers really not handle a merge graph? It was bizarre how something straightforward in other VCS is suddenly "messy" amongst git folks.
It is like moving to one's favorite part of town.
As someone who has had to put up with rebase for several years now, my life is definitely not better. And things that were trivial in mercurial are now complicated (seeing the actual chronology - both in the log and in the graph). That graph with multiple branches that git developers find messy can actually be really useful.
I'm a fan of rebase myself, but understand the point made above. For me, the biggest pro of a clean history is when doing `git blame`. If the history is clean and the commits are good, it might solve my issue. On the other hand, if the commit in question is a huge mess of unrelated things it doesn't help me at all. I also find it way easier to review a PR with a clean, well-described history.
But because GitHub and other tool’s version of rendering history just flatten merge commits into spaghetti we’re stuck with squash merge. Thanks GitHub.
Also sometimes you decide you want to backport some change to other releases, and if commits are in a good state, it is much easier to do this.
Whether this is feasible at all depends largely on the care developers put in structuring their commits.
A free form textual interface to document everything about why you made the changes you just made? Why not maximize the value of this resource!
I’m not necessarily on Team Rebase, but isn't this just as likely with merging gone wrong?
A badly done merge can indeed ruin code. But you'll always have the versions that went into the merge, and the merge itself. Your history has all of the information to recreate exactly what happened, find what changed, and then figure out how to fix it.
A badly done rebase not only ruins your work, it also removes from the branch any record of your work having been done. Unless you can find the right stray old commit which is not yet cleaned up, there is no choice but to start doing it again from scratch.
I find it easier to run git binary search with it like this too.
This being a principal reason for VCS, I very much understand the motivation.
If your goal is to be able to revert the codebase to a previous version, then you want your history to a series of well prepared, atomic changes where at each point the software is actually functional.
At that point I have a much better idea of the scope of my changes, and I can revise them into a few coherent commits, rather than a mess of "WIP" commits that are not a useful history to keep.
What exactly the word "version" means depends on context. If you put your essay under git, you probably want to track how it was changing with time. When you hack on some codebase and throw things at wall to see what sticks, you want to be able to go back to the previous attempt should your next one turn out to be useless. You may even want to commit things just because it's a convenient way to send things to be built by CI - so you basically produce "versions" to test.
But when you collaborate with others to develop a project, nobody cares about whether you made typos during your hacking session and had to come back to fix them. It's not a useful information and it never becomes a "version" of a shared project, because why would it? It's just confusing and wastes people's time on review, and makes things like blame and bisect harder to use.
Or when you put a tutorial under a git repo, with each commit representing the next step to achieve a certain outcome. You may have tweaked each step gazillion of times to perfect it, but that "history" is completely irrelevant for the resulting repo. It's meant to store "versions", not "history". Those may correlate, but don't have to.
A git repo is a data structure that you operate on. Treat it as such and use to achieve your goals.
What's funny is that "squash on merge" strategy gets you worst of both worlds. You don't get nicely curated versions in your project because if someone actually cared to fine-tune their MR you just throw that information away, and the rest doesn't care in the first place anyway. Rebasing and squashing is an incredibly useful tool for everyday use by developers in their work, but it often gets used as a band-aid for lazy developers instead.
In other words, the history.
> It's not a useful information and it never becomes a "version" of a shared project, because why would it?
Because the real world doesn't match your ideal of how software development works. You need the full history because sometimes innocuous looking "simple fixes" are neither innocuous nor fixes, and because obscuring the history makes auditing and merges more difficult, not less.
> It's just confusing and wastes people's time on review, and makes things like blame and bisect harder to use.
So blame and bisect are poorly written therefore you should hack around them using a convoluted process that obscures the true history and raises the risk of requiring people to redo weeks worth of work by overwriting branch histories on shared repos. Great idea.
The authors of Fossil and Sqlite did a complete breakdown of everything wrong with rebase and what a proper tool should do, so I won't belabour the point further:
Yes, that was my point. Have you read it?
> You need the full history because sometimes innocuous looking "simple fixes" are neither innocuous nor fixes, and because obscuring the history makes auditing and merges more difficult, not less.
Obscuring logically split and curated commits makes auditing and merges more difficult. Obscuring pointless edit history of the developer makes them easier.
> So blame and bisect are poorly written therefore you should hack around them using a convoluted process that obscures the true history and raises the risk of requiring people to redo weeks worth of work
blame and bisect are powerful tools that can work sensibly in various repository topologies. However, if you're either intentionally putting garbage into your topology, or not utilizing it well because of ill-defined idea of "clean history", you're simply doing yourself a disservice and induce unnecessary mental load.
> by overwriting branch histories on shared repos. Great idea.
Who said anything about overwriting branch histories on shared repos? It has its uses too (it can be useful in some cases when you're a downstream working on a project maintained upstream, for example), but it's not what most projects will ever want to do. That's not what rebase is there for.
You should not be able to use `--amend` during a rebase.
For me editing all my changes onto the commit I'm working on with `git commit -a --amend` (or as I've aliased it, `gcaa`) is automatic; I do it 500 times a day, just to save my work. But I can't count how many times I've been in the middle of squashing commits and accidentally typed `gcaa` and amended someone else's commit after fixing a merge conflict, and it's super annoying to unwind (if you realize after typing `rebase --continue`) so usually I end up just giving up and starting over. I really wish amending to a commit that wasn't one of the ones you're rebasing was just totally disabled.
I guess there are some other small complaints, like the annoying reversing of `--ours` and `--theirs` from what makes sense (yes, it makes sense if you have the internal model of rebase instead of the intuitive one, but that's stupid), rebase's tendency to pick the wrong parent commit if you've accidentally amended someone else's commit (and therefore lag a while and then produce a rebase log of 1000 commits or something), and the utter tedium of editing the rebase log to replace every instance of "pick" with "s" for squash except the first, since almost 100% of the time what I want to do is squash everything (and use the last commit message, not the first, and definitely not all of them munged together which is the default).
I would love a separate command or a flag, like "git rebase --tip" that does all of this automatically for my otherwise extremely elegant workflow (and I'm gonna be really bummed if it turns out it exists and I didn't know about it for the last 5 years...).
Probably easiest as a little shell function like
gcca() {
local GIT_DIR
if ! GIT_DIR=$(git rev-parse --git-dir); then
return 1
elif test -f "$GIT_DIR/REBASE_HEAD"; then
printf 'Rebase in progress: commit --amend is disabled\n' >&2
return 1
fi
git commit -a --amend "$@"
}
rather than an alias?[Edit] I forgot about rev-parse --verify, which simplifies this further:
gcca() {
if git rev-parse --verify REBASE_HEAD >/dev/null 2>&1; then
printf 'Rebase in progress: commit --amend is disabled\n' >&2
return 1
fi
git commit -a --amend "$@"
}
This also leaves you still able to use commit --amend long-hand if (for example) you want to edit one of your own commits during rebase -i.You could try reverting the first commit on the HEAD once you finish the rebase. This is of course assuming your branch and the last commit don't touch the same files.
Then on the final rebase the commits are automatically ordered with s and f as appropriate.
Although I do a fair bit of amending, too.
There's a false dichotomy nobody addresses here, which is the notion that there needs to be such a thing as "the" history for you to get the benefits of a clean history.
If all you really want is a linear history, then just do merges, and make sure the "first parent" is the main branch (which you can enforce with tooling). Now you can just traverse solely the (linear!) sequence of first parents, which is exactly the same view squashing would have given you, except without the information loss.
If for some reason you can't stand the idea of something branching off your main branch at all, then set up a separate job that automatically squashes everything onto a branch that only it can write to (or branch from). Now you have a truly linear history with nothing branching off it, exactly as you would've had with squashing. And you can always reproduce it on demand.
That way you avoid the information loss, and can always do archeology on the full evolution graph if needed.
Don’t merge the base branch into a feature branch. Rebase to “update”.
Do use rerere and the curse of fixing the same conflict over and over is (almost) gone.
Don’t rebase (or force push for other reasons) a shared branch. Rule of thumb here is you can probably rewrite history if you work with _one_ coworker in a branch but any more than that and you’re more likely than not to upset someone.
Do rebase -I HEAD~N to reorder/reword/squash into easily reviewable sequences of commits.
Don’t force push after review, until the review is complete. This keeps the history of the review process but you can later merge the fixups with the commits they logically belong in right before merging.
Do use Merge, Squash and “Rebase+FF” as appropriate for merging PR. There is no best solution for every scenario so prescribing “always merge” or “never merge” or similar isn’t helpful. A good rule of thumb though is that IF a branch has merged from the parent branch to update (which I suggested was a “don’t”) then avoid merging it back. A branch that was updated that way is better to e.g squash when merging back.
I've had to talk to soooo many developers about this. I want to see what changed since my last review, not restart my review.
This is usually caused by merging an upstream branch (e.g. develop) into your feature branch and then later trying rebase it.
Effectively the commits you've merged in from develop undo the changes you've made in your feature branch. You fix them but the foreign commits undo the changes again.
The solution is actually pretty easy. Use git rebase --interactive to remove any commits from the rebase that aren't directly part of the feature work.
You may still have an odd merge conflict to fix but you'll only have to do it the once and everything should go smoothly.
I would also recommend never using the same commit message twice. When you have a list of 10 commits all called "Wip" it's hard to tell which are obviously duplicates that can be deleted.
git rerere is the even easier solution.
> fixing the same conflict repeatedly is annoyingI use it all the time and I really like how I can make garbage commits (wip, test) and then squash them into atomic commits which are easy to review and later on easy to bisect when inevitably mistakes happen. Sure I've fucked up too when I was learning on how to use it and those were some painful mistakes but only through using it and making those mistakes have I learned to use the tool to great advantage (clean history).
The whole purpose of source control is to reliably track code changes so you don't lose anything and can revert to any point or recover from bad merges. Since rebase permits you to violate this core purpose and literally lose the entire history of code changes, then yes, it is stupid.
Getting away from dependence on the ephemeral is why git exists.
When I hear people griping about rebase, I assume that nobody took the time to teach them how to use reflog first. Once I had an understanding of reflog, I could mess up all I wanted (without pushing) and recover. In that environment, rebase can become a very useful tool. Without being able to recover, rebase becomes a tool of confusing irreversible destruction.
No, it's stupid because it's really common for people to fuck it up, and because the purported benefits (clean history) are not something which matters.
You've told me that it doesn't matter to you, but that hasn't changed my mind.
It's not a matter of objective truth, it's a preference.
I would have made this part of a root-level comment but I doubt anyone would read it, but: I think what gets lost in all these git debates is what language/context are we talking about? a shop that churns out javascript and releases to prod every 8 hours is very different than a C++ shop that writes safety-critical software. Their git needs are very different, and having an "I make my bed every day at 5:30am before I go for a 5mi run and come back and drink my juice and eat avocado toast" git regimen may be appropriate for some codebases but not for others where "I woke up hungover at 10am with a partner whose name I cannot remember, in a bed that is not mine" regimen. I think countless human-brain cycles are lost to bickering between these 2 camps.
https://www.mercurial-scm.org/doc/evolution/
It's a good idea that's been attempted to be ported into a git
Better alternative I've found is squash merge - topic/feature branches are squashed as a single commit instead of bringing down each individual commit or creating a merge commit. You're history is cleaner, you're able to revert stuff easily, and it's really hard to mess up since it's just an atomic last step you do in your workflow.
https://www.git-scm.com/book/en/v2/Git-Tools-Rerere
It's often easier to resolve a conflict during a merge than during a rebase because it presents you with left, right, and the common ancestor. You're also only looking at the tips of each branch. With rebasing, you're replaying each commit one on top of the next so you lose the common ancestor information and you may also have conflicts that won't exist at the end.
Another tip: if the other branch has changed a lot since you last rebased, even a single merge may have a lot more conflicts than you want to deal with all at once. In this case, consider a series of intermediate merges since you're going to throw them all away anyway.
...which is exactly what is suggested in the section linked in the article, https://github.com/kimgr/git-rewrite-guide#split-a-commit.
* Regarding "weird interactions with merge commits" - `git rebase --rebase-merges` tends to help most of the time, since, during a rebase, merge commits are skipped by default (even if they contain changes).
On any code base I've worked on that's larger than a small FOSS project, I've found that this simply isn't avoidable. Yes, there's merge commits but, for reasons I won't go into, I think those are worse than the alternative of rebasing and making code reviews difficult.
> One way to avoid this is to push new commits addressing the review comments, and then after the PR is approved do a rebase to reorganize everything.
Not realistic when working on a code base where PRs are being squash-merged every hour and the code review lasts for days.
The best middle-ground is to avoid rebasing until the current wave of feedback has been resolved, even if no one has actually approved yet.
But if they collide, you have to resolve all the merge conflicts anyway, and then with rerere you should have relatively little additional work on the final rebase.
Why squash merges? I have a number of team-mates who make local commits on feature branches that make the history look like a series of less-than-useful commit messages. E.g. wip, wip, wip, wip, make it work, wip, wip. All of the context for the change is actually on the PR, so it's really only helpful to see the PR message and have a link to the PR for the discussion on the change.
It is particularly useful when doing difficult merges regularly. Invariably I'll find a mistake in the merge and start over (before pushing, obviously); the second "git merge" remembers the previous resolutions so I don't have to solve all the same conflicts again.
Similar for difficult rebases that may need multiple attempts.
Git remembers resolutions across branches and commits, so in the rare case where (say) a conflict was solved during a cherry-pick, rerere will automatically apply the same resolution for a merge with the same conflict.
I think the reason it's not on by default is that the UI is confusing: when rerere solves for you, git still says there is a conflict in the file and you have to "git add" them manually. There is no way of seeing the resolutions, or even the original conflicts, and no hint that rerere fixed it for you.
You just get a bunch of files with purported conflicts, yet no ==== markers. Have fun with that one if you forget that rerere was enabled.
been using rerere for years, never seen this behavior.
I'm still searching for a way to manage long-lived Postgres submissions, the most challenging git scenario I've encountered. Julia's post finally got me to brain-dump my current process, something I've meant to write down for a while now:
https://illuminatedcomputing.com/posts/2023/11/git-for-postg...
This link could almost be an "Ask HN": if any of you have suggestions to improve my workflow, I'm all ears. (I asked around a bit last May at PGCon, but didn't get any concrete advice there. Maybe it's too complicated for a hallway off-the-cuff discussion.)
One thing that annoys people when rebasing is, if you need to rebase a couple of times and also you change the commit history of current branch, you might end up solving the same "conflicts". To avoid this, you can use git rerere, this basically saves your conflict resolution, and if the same conflict is encountered, it resolves it automatically: https://mirrors.edge.kernel.org/pub/software/scm/git/docs/gi...
If that was during an 'edit' phase and not a conflict, then you get a blank state, but you need to commit any change there anyway before applying the other patches so commit s also the right thing to do... If you had intended to amend instead of a new commit then you can just squash the commit again later, or reset (soft!) HEAD^ to undo the commit and add again/amend.
When I was at amazon their internal PR tool handled them just fine so I would do them in that case.
There's more nuance to the OPs comment along the lines of
> Don't rebase on branches others are working on
On a pr branch I usually at most expect others to pull and do a build to run it locally so I'm not very worried about wiping out changes.
I do a lot of rebases, but they are trivial rebases in a feature branch, like changing the order of new features and fixups, and then squashing the fixups to make a nice PR. Don't try to do smart weird rebases!!!
From time to time I have to make a smart weird rebase because my small commits order is just a mess, or I have to split a commit or something unusual. The important first step is to make a new brach as a backup to hold the version before the rebase. If I mess the rebase process I can just go back to my version before the rebase and start again.
If I need to do a non-trivial rebase I always start by creating a backup branch so I can delete my fuck up and start over trivially, which was a hard learned lesson.
A repository shouldn't be a dump of unmeaningful commit messages, but a curation of best contributions at the time of commitment.
- if they have done any cherry-picking - if they have done any backporting - if they have reverted any commit that made it to main - if they have ever used bi-sect
I think rebase is harder to learn than merge, but once you get used to it, it lets you have history that is easier to debug, and use without extra effort
You want to "pretend" you all took turns at making the code better, like Alice goes first then everyone downs tools while Alice makes her chnages and then Bob picks up Alices work and does his changes, then Charlie starts
The difference is that Bob can start while Alice is working and all he needs to do is right before checking his stuff in, he grabs her fixes from master, and then applies his fixes on top of hers as if he had started after she had finished. Sometimes they both worked on the same code and he needs to figure out what's safe but hey that happens any which way.
As long as you ensure fast-forward merging onto master only it's kind of simple.
Annoying. but simples.
Use real merge commits and `--first-parent` as your default view.
It's unfortunate that to make `--first-parent` default you have to either edit your git config or grow a few new habits, and I still think there should be at least few more UIs that are focused on `--first-parent` with optional "drill down" instead of raw subway map diagrams. Subway diagrams look cool in screenshots, but so much of the complaints about "clutter" and "mess" in git seem to be just that people don't actually want to read the subway diagrams.
1. Open a new branch and do development there. You can rebase and force-push all you want.
2. When ready for review, open a PR of the branch, and never rebase again.
3. On approval, Squash-Merge your PRs and include a merge comment linking to the PR.
This way you can "clean up" your development history in your branch, maintain history of changes requested in a PR, the entire development history is available at the PR link later, and you can revert an entire PR by reverting a single commit. git --force (acts like git --force-with-lease does now)
git --force-anyway (acts like --force does now, "anyway" is just an example and should be harder/longer than -f/--force)
I understand that force-with-lease didn't exist first but this needs to be rectified.Get some updates, then merge master branch. Seems to be pretty straightforward to me.
git fetch -p --all
git pull
git merge master
(Handle merge conflicts in WE of choice.) git push
Boom. I'll let the CI platform squash commits post-merge.Option for gitlab if it's a small change and you don't want to run a full suite of tests.
git push -o ci.skipNow I make a single commit per PR that I commit —amend —no-edit until it’s merged. I sometimes have to rebase it onto main for conflicts but that’s easy.
Fresh clean branch, no commit history, create pull request.
I'm not convinced there's any value to incremental commit messages. This simple, clean, and undoable as long as I keep my initial branch
We've seen this in npm. Npm was supposed to be simpler than maven - then it slowly rediscovered the reasons we need package signing, support for circular dependencies, and all the other messy things that go into package management.
While I don't have numbers, it seems like a huge percentage of the "this problem should be simple, let's build a new app" projects either fail or recreate the same gnarly problems that led the existing projects to be complicated.
For example, Darcs (older) and Pijul (newer) are based on patch theory, so Git's rebase issues are moot there.
*) Of course, "better" is subjective.
FWIW, Mercurial's been there for almost 20 years.
git branch
git checkout
git clone
git diff
git pushHowever, life is much better with rebase than without it.
Shared branches is a bit of a faff just due to coordinating.
I branch, merge main in regularly, make a PR, squash back to main and delete the branch.
Linear history, single commits for a single piece of work.
Easy, clear, no fuss.
The PR itself contains full history of subwork. GH can recover branch from PR. No work is lost.
Dude! Learn the tilde notation
I'd rate the benefits of version control, in rough order of importance, as:
* It allows multiple programmers to work on one body of code. This has always been it's main use. Rebase screws with this because multiple authors updating a rebased branch creates a cluster. The simple fix is rebases to a branch are only allowed when you are the sole person working on it. (Atlassian is wrong. The branch can be public. It's multiple writers that creates the problem, other people reading and reviewing a public branch isn't.)
* A backup. I've lost more than enough work to know the importance of backups yet if I have to do it manually I still don't do it regularly enough. git==backup. Wonderful.
* Assist in reviews. Actually, I'm not sure how you would do reviews without it, as version control system both highlights the differences and serves as a communication medium. Rebases help here, as they let the author parcel up a body of work to make the reviewers job a lot easier.
* Make open source contributions auditable. To be fair I've never used this personally, but I use software that depends on it - like the kernel, so I rank it pretty highly. If becomes very difficult to anonymously introduce malicious changes when the version control system is tracking who made every modification. In an amazing coincidence, a branch is effectively a block chain which makes it hard to change. Rebases could muck this up of course if the "only personal branches may be rebased" rule isn't enforced - but it normally is.
* Bug hunting. Blame and bisect are the main tools. In my experience compared to the previous points this gets used very rarely, but bisect in particular can save a lot of time on those rare occasions. Before version control we did bisects using by restoring backups. Notably git blame still works perfectly if you follow the "only rebase branches you own" rule.
* Going by the discussion here, some people spend time on archaeological digs through source code repositories. It seems some of them prefer their digs to be dirty (aka rebase free) and others like it clean (rebased).
Interestingly, most of the noise here comes from people arguing about the last point. That strikes me as about important as the colour of the bike shed. The only other place rebase effects is reviews. Reviews are a problem everywhere I've worked. Everybody prefers to be doing something else, so making them as friction free as possible is a worthwhile goal. Rebasing does that (and so does unit tests).
me: Do you even git bro?
If I understand what rebase does correctly, it just adds all my commits in my remote feature branch to the HEAD of the branch that's being rebased onto. That's why after doing it and finishing it, one needs to do a push force because the head of the remote feature branch would diverge. But... Why? How is this better or different than just merging master into the branch? Gitlab for example and Intellij show all the branch changes and commits with their hashes so it all can be cherry picked or reverted if needed quite easily...
Has anybody who used to "merge master into" and now uses rebase that has a much different view on it being better?
* Merge Branch 'B' into 'main'
|\
* \ Merge Branch 'A' into 'main'
|\ \
| | |
| | * B Commit 2
| | * Merge remote-tracking branch 'upstream/main'
| * | A commit 2
| * | Merge remote-tracking branch 'upstream/main'
| | * B commit 1
| |/
| * A commit 1
|/
*
Compared to a commit graph where feature branches are rebased to replay their commits against the tip of main before merging: * Merge branch 'B' into 'main'
|\
| * B commit 2
| * B commit 1
|/
* Merge branch 'A' into 'main'
|\
| * A commit 2
| * A commit 1
|/
*