* use filtering commands like "git log -S"
* press the "annotate" button in my IDE and can see which commit introduced each line
* run "git bisect"
* use "tig" to drill down through the history of a file (shortcut "," is "move to commit preceding current line's blame commit")
...every step of the way, I get a meaningful description of why a change was made and what other diffs were necessary to achieve that change. And not just "fix", "bug", "PR commments".
In PyCharm, I can see which commit introduced each line, regardless of branching. Same with drilling down through a files history. Is this an IDE limitation you're seeing?
> every step of the way, I get a meaningful description of why
Isn't this more about commit messages, than anything else?
> Is this an IDE limitation you're seeing?
I'm using Jetbrains too.
* `git blame --first-parent`
* `git bisect --first-parent`
* At least one "tig-like" with a --first-parent first UI: https://github.com/kalkin/git-log-viewer
Makes reviewing a set of changes prior to a merge much easier. It's nice if there's a 1:1 correlation between a commit message and the actual patch contents.
Im sure you've dealt with the case of reviewing a colleague's changes with a commit message like "Enable logging in foobar module" and the patch is actually enabling foobar logging and a bunch of other stuff.
This makes bisecting your git history to identify and fix bugs much more difficult.
If the git history is clean, you can just read the commit messages and implicitly trust the developer if clean git hygiene is in place (as opposed to actually needing to read the whole diff on a per-commit basis to find out what _actually_ happen at commit XYZ, despite it's message).
I'm a fan of rebase myself, but understand the point made above. For me, the biggest pro of a clean history is when doing `git blame`. If the history is clean and the commits are good, it might solve my issue. On the other hand, if the commit in question is a huge mess of unrelated things it doesn't help me at all. I also find it way easier to review a PR with a clean, well-described history.
Every time you do a git rebase, you are literally asking your source control system to lie about history. If you mess up, and you eventually will, you're then forced to manually figure out what the history really was despite being lied to. If you mess it up, well, good luck.
I used to work at a company where someone (we never figured out who) in another group would rebase every few weeks. We didn't find out about it until their stuff was pushed then released. The result was that features which we'd written, QAed, and released to production would simply disappear a few weeks later. With no history suggesting that it ever existed.
Have you ever been pulled off of a project to go fix a project from a month ago which has disappeared from source control? You don't know what happened, you no longer have context, you've just got complaints because your stuff no longer works.
Is your desire for a "clean history" worth potentially creating THAT disaster for other developers on your team???
You get clean history by not merging branches with 50 intermediary "fiddling with X" commits in them.
Wouldn't this be trivially solvable by git bisecting your deploy branch?
I mean that's the equivalent of reversing your JCB through a house on a building site because "the house was not there a week ago when I last moved the JCB".
or am I missing something?
We had 4 teams, each released on a schedule. Each team had a branch. When a team released, it was merged from master, then each team pulled. The person who did the pull would change, and it was often a rebase.
So my team released, some other team rebased badly. There would be no sign of problems for us until after they released. But since 2 teams generally released at once, and people didn't remember who actually did the merge a few weeks earlier, it was hard to figure out who was actually messing up. (I had suspicions, but no proof.)
I've seen rebase used appropriately since. But that disaster left scar tissue.
Yeah that would leave a mark.
What others here, including myself, are advocating is rewriting your own history before you share it, which makes a very different set of trade-offs.
How do you determine which combination of changes are yours, and not accidentally undo changes that other developers made in the last month?
Could git reflog have helped? Maybe sometimes. If we had the foresight to have saved the right commits from the past, sure. I think we did start saving old branches just in case it happened again, but then we had to sort through a month of changes to figure out what to keep, what to change, and what conflicts there might be.
Remember, the person screwing up didn't know he screwed up. And the person trying to fix it is doing so a long time after the fact. It was a disaster.
The true history is not recorded in your normal commits either. Every time you modify your source buffer, that is the true sequence of events. This truth is lost already as you undo/rework things before you commit. You're ALWAYS manipulating and telling a false story of history whether you realize it or not.
Commits are a tool that give stronger backup/undo protections over simple file saves and in-memory editor undo lists. Just because you happened to save your work in a commit doesn't mean it should be instantly be regarded as holy history. Not anymore so than if you simply saved the file.
I think the bar for "holy" history should be whether it is published to a shared branch.
Every single point in the commit history represents an actual state of a repository at a specific point of time, along with the information of which point or points were next before it. This is all part of the true history.
This is not a full history - you don't have every keystroke, abandoned commit, switch between branches and so on. But nothing that you're being told is wrong.
As soon as you do a rebase, you're rewriting history. You're claiming that there were specific points of time with specific states that never actually existed. You're losing information about points of time and specific states that actually existed, which someone once considered important enough to do a git commit over.
The difference becomes important if that someone, which at a previous job was me far more often than I would like, tries to go back to the historical commit. And finds that it is gone without a trace.
Agreed. But the rewrite occurs in your private branch. It's history is just as private as the undo list in your editor. No one cares about what's going on in your editors undo list. And by the same logic they shouldn't care about commits in a private branch.
> You're losing information about points of time and specific states that actually existed
If you avoid rebase, then you end up "rebasing" without rebasing. You "squash" intermediate states by never recording them to begin with.
Failing to record history is not superior to squashing it.
> And finds that it is gone without a trace.
I don't have the details, but it sounds like someone rebased a public branch. Yes that is bad. But it's sort of like saying we shouldn't drive cars because someone chose to drive the wrong direction down a 1 way road.
The rebase becomes part of the public branch eventually, inflicting your lies on everyone else.
> If you avoid rebase, then you end up "rebasing" without rebasing. You "squash" intermediate states by never recording them to begin with.
If only Git had a third alternative, a way to... entangle two diverging branches of history without destroying or rewriting either. You could say it would be a bit like a car merging into a highway.
> Failing to record history is not superior to squashing it.
"We don't know" is at least an honest statement. Claiming that you do but then making up some nonsense is something that the LLMs do enough of already.
What do you think of the editors undo log? It's a very real historical log. Should it be treated as "holy" history too? If not, what makes the undo log less true/important than a private git commit log?
Squashing is often dumb and unhelpful, because you're now re-summarizing the points in time that you already considered worth highlighting when they happened (when you had the most context to judge them!).
Rebasing is lying about the order and/or context that those changes happened in.
Your undo log is comparable to squashing, but not at all to rebasing.
And then again, the first-order vs second-order summarizing distinction matters, and you already capture the second-order summary in your merge commit. Squashing is just destroying information for zero practical benefit.
> private
You keep using that word, but branches are often a lot less private than you think. Push it to get a colleagues' input on something? Congratulations, it's now public. Created a pull request that you want to revise? Already public.
Fun note though, I argued this directly with Dr Hipp (principal author of SQLite and of Fossil, inventor of the ‘git rebase is a lie’ argument) and during that discussion, he agreed to soften the language on the Fossil pages. They are still hyperbolic, using the word ‘dishonest’, and continue to distort the reasons and usage behind rebase, but he did remove some instances of the word ‘lie’ and ‘lying’, which is progress.
It’s a bit of a shame that they haven’t found the strength to frame Fossil in a positive light without trash-talking the competition. There is a good-faith argument for Fossil vs git, but they’re choosing not to use it.
BTW rebase produces a new commit ordering, but does not modify the old one.
> You’re claiming that there were specific points of time with specific states that never actually existed.
No. You are asserting intent on the part of git users and git that has never existed, you have misunderstood what git history is. The git history is not a claim that the state at that point existed during development, you are projecting your own goals that are not shared by git or git users.
> You're losing information about points of time and specific states that actually existed, which someone once considered important enough to do a git commit over.
Hehe this is so full of assumption. You write it like I’m rebasing someone else’s work, but you already know I’m only rebasing my own commits, and I’m the one who decides what’s important enough to do a commit over.
I like commit early, commit often. I want to make small incremental commits that don’t display to others that way and I expect to put small commits and fix ups together later into a single useful commit with only one commit message.
Committers always have a choice of which of the changes present in their working tree they stage and then commit. The commit history is always a flat approximation of the real evolution of the files in the repo.
Alternative DVCSes which support this workflow include: fossil
But I have never seen a commit disappear while rebasing, ever. That workflow is busted somehow. They were doing it wrong.
This argument reminds me of a scene from Yes Prime Minister [0]:
> Humphrey: The minutes do not record everything that was said at a meeting do they?
> Bernard: Well of course not.
> Humphrey: And people change their minds during a meeting don't they?
> Bernard: Well, yes.
> Humphrey: The actual meeting is a mass of ingredients for you to choose from.
> Bernard: Oh, like cooking?
> Humphrey: No, not like cooking. Better not to use that word in connection with books or minutes. You choose, from a jumble of ill-digested ideas, a version which represents the Prime Minister's views, as he would, on reflection, have liked them to emerge.
> Bernard: But if it's not a true record...
> Humphrey: The purpose of minutes is not to record events it is to protect people. You do not take notes if the Prime Minister say something he did not mean to say, particularly if it contradicts something he has said publicly. You try to improve on what has been said, to put it in a better order. You are tactful.
> Bernard: But how do I justify that?
> Humphrey: You are his servant
> Bernard: Oh, yes.
> Humphrey: A minute is a note for the records and a statement of action if any that was agreed upon.
I think the analogy is pretty clear. A pull request does not record every single little change you made when writing it. You choose, from a jumble of ill-digested ideas, a version which better reflects your intent as you would, on reflection, have liked it to emerge. It doesn't matter that it's not a true record, since its purpose is not to record events but to communicate ideas. You try to improve on the commits as they have been written, to put it in a better order. You are tactful.
But because GitHub and other tool’s version of rendering history just flatten merge commits into spaghetti we’re stuck with squash merge. Thanks GitHub.
I recall when I first switched to git at work and the team was insisting on a "linear history", I was bemused: Could these developers really not handle a merge graph? It was bizarre how something straightforward in other VCS is suddenly "messy" amongst git folks.
It is like moving to one's favorite part of town.
As someone who has had to put up with rebase for several years now, my life is definitely not better. And things that were trivial in mercurial are now complicated (seeing the actual chronology - both in the log and in the graph). That graph with multiple branches that git developers find messy can actually be really useful.
Also sometimes you decide you want to backport some change to other releases, and if commits are in a good state, it is much easier to do this.
Whether this is feasible at all depends largely on the care developers put in structuring their commits.
A free form textual interface to document everything about why you made the changes you just made? Why not maximize the value of this resource!
I’m not necessarily on Team Rebase, but isn't this just as likely with merging gone wrong?
A badly done merge can indeed ruin code. But you'll always have the versions that went into the merge, and the merge itself. Your history has all of the information to recreate exactly what happened, find what changed, and then figure out how to fix it.
A badly done rebase not only ruins your work, it also removes from the branch any record of your work having been done. Unless you can find the right stray old commit which is not yet cleaned up, there is no choice but to start doing it again from scratch.
I find it easier to run git binary search with it like this too.
This being a principal reason for VCS, I very much understand the motivation.
If your goal is to be able to revert the codebase to a previous version, then you want your history to a series of well prepared, atomic changes where at each point the software is actually functional.
At that point I have a much better idea of the scope of my changes, and I can revise them into a few coherent commits, rather than a mess of "WIP" commits that are not a useful history to keep.
Why bother with this step at all? It's literally pointless and serves only to stroke an (IMO) silly aesthetic preference.
> rather than a mess of "WIP" commits that are not a useful history to keep.
I also disagree that WIP commits are not useful history. You might have explored 2 or 3 different abstractions to solve a problem, and picked one but it turned out to be the wrong one, and one of the others would have been a better choice, but now you've lost the history where you explored these options. Are you suggesting it's no loss to erase those other commits and the context around which you thought it wasn't a good choice at the time?
What exactly the word "version" means depends on context. If you put your essay under git, you probably want to track how it was changing with time. When you hack on some codebase and throw things at wall to see what sticks, you want to be able to go back to the previous attempt should your next one turn out to be useless. You may even want to commit things just because it's a convenient way to send things to be built by CI - so you basically produce "versions" to test.
But when you collaborate with others to develop a project, nobody cares about whether you made typos during your hacking session and had to come back to fix them. It's not a useful information and it never becomes a "version" of a shared project, because why would it? It's just confusing and wastes people's time on review, and makes things like blame and bisect harder to use.
Or when you put a tutorial under a git repo, with each commit representing the next step to achieve a certain outcome. You may have tweaked each step gazillion of times to perfect it, but that "history" is completely irrelevant for the resulting repo. It's meant to store "versions", not "history". Those may correlate, but don't have to.
A git repo is a data structure that you operate on. Treat it as such and use to achieve your goals.
What's funny is that "squash on merge" strategy gets you worst of both worlds. You don't get nicely curated versions in your project because if someone actually cared to fine-tune their MR you just throw that information away, and the rest doesn't care in the first place anyway. Rebasing and squashing is an incredibly useful tool for everyday use by developers in their work, but it often gets used as a band-aid for lazy developers instead.
In other words, the history.
> It's not a useful information and it never becomes a "version" of a shared project, because why would it?
Because the real world doesn't match your ideal of how software development works. You need the full history because sometimes innocuous looking "simple fixes" are neither innocuous nor fixes, and because obscuring the history makes auditing and merges more difficult, not less.
> It's just confusing and wastes people's time on review, and makes things like blame and bisect harder to use.
So blame and bisect are poorly written therefore you should hack around them using a convoluted process that obscures the true history and raises the risk of requiring people to redo weeks worth of work by overwriting branch histories on shared repos. Great idea.
The authors of Fossil and Sqlite did a complete breakdown of everything wrong with rebase and what a proper tool should do, so I won't belabour the point further:
Yes, that was my point. Have you read it?
> You need the full history because sometimes innocuous looking "simple fixes" are neither innocuous nor fixes, and because obscuring the history makes auditing and merges more difficult, not less.
Obscuring logically split and curated commits makes auditing and merges more difficult. Obscuring pointless edit history of the developer makes them easier.
> So blame and bisect are poorly written therefore you should hack around them using a convoluted process that obscures the true history and raises the risk of requiring people to redo weeks worth of work
blame and bisect are powerful tools that can work sensibly in various repository topologies. However, if you're either intentionally putting garbage into your topology, or not utilizing it well because of ill-defined idea of "clean history", you're simply doing yourself a disservice and induce unnecessary mental load.
> by overwriting branch histories on shared repos. Great idea.
Who said anything about overwriting branch histories on shared repos? It has its uses too (it can be useful in some cases when you're a downstream working on a project maintained upstream, for example), but it's not what most projects will ever want to do. That's not what rebase is there for.