Git rebase in depth
git-rebase.io
git-rebase.io
I find that terminology terribly misleading and when I was learning git and rebase it confused the heck out of me.
No commits are harmed in the operation of `git rebase`. All the commits you had in the repo before the rebase are still in the repo. Git rebase creates a new sequence of commits and after doing its work relocates the branch name to the tip of the new sequence but you can easily access the previous commits if need be:
$ git co feature-branch
$ git rebase develop
$ git co -b before-rebase-feature-branch feature-branch@{1}The same changes are still in the repo (edit: I should have said branch here), but not the same commits, because the parents and children change and therefore the hash of the commits.
It is very important to be aware that the history is changed, because the previous history can not be without issues merged with the new one, which is the main pain point and the most problems that arise from a rebase.
Each commit has a link to its parent, and represents the tip of a linked list. .git/objects is a heap of all commits (and other objects), and .git/refs contains a list commit IDs that define each head (e.g. master). git rebase will often introduce new versions of a commit to the heap and update the heads to reference new histories, but the old commits stick around and can be accessed through the reflog - with their full original history intact.
A good UI should allow you to recover from mistakes. Like the trashcan vs rm example everyone is using.
It's good that git doesn't permanently delete stuff, it's bad that you need to be a relative expert to know that. If git branch showed rebased branches and told you they would disappear in x days then beginners might feel less fear and embrace the power of git faster.
git rebase has an "abort" feature when you need to redo it. git has tag & branch & stash features if you want to save what you're doing before you rebase. The problem with keeping and showing rebased branches are that 1- you don't need them after the rebase is successful. You only need them when the rebase is going badly, and 2- you'd have a lot of unnecessary noise pile up. I often rebase multiple times before every push. I don't want to see them all, you probably don't either.
That said, I fully agree that git's UI could be better and help beginners feel less fear!
They kinda are, but that's like saying that deleted files are not deleted, but stick around for a while.
While technically true, for most practical purposes _rm <file>_ deletes the file. The fact that each and every "git 101" manual has to explain how to recover deleted commits, means something is wrong.
It's like saying: "Here's the key, and in case it doesn't work there's a pry bar in the garage". This is usually a pretty good indicator that the lock is broken.
Git has a very good safety net when you know how to use it. The problem is knowing how to use it, not that it’s not there.
It’s a legitimate point that git’s UI sucks, that’s what’s wrong, and everyone agrees. But learn how to use the reflog and you will see the light!
Someone who has never done it before won’t know how to retrieve them, but that’s hardly a surprise.
I don’t understand what you mean. Why would you rebase a branch and then merge the rebased version with the old version of that branch? I don’t think that’s a common workflow, that is a specific case you should avoid.
Typical use of rebase to clean up the unpushed portion of a personal repo before pushing doesn’t normally cause any problems. The main thing that pops up for me is trying to re-order my own dependent commits when I forget that one commit touched the same code as another. Squashing them during the rebase is the easy fix.
They are still in the repo, but if no treeish item (eg. a branch) points to them, then they'll eventually get garbage collected.
Still, glad to see people are trying to elucidate git rebase. A small subset of its functionality is fundamental part of my workflow and I wouldn't know how I'd use Git without rebasing.
>git gc tries very hard not to delete objects that are referenced anywhere in your repository. In particular, it will keep not only objects referenced by your current set of branches and tags, but also objects referenced by the index, remote-tracking branches, refs saved by git filter-branch in refs/original/, or reflogs (which may reference commits in branches that were later amended or rewound). If you are expecting some objects to be deleted and they aren’t, check all of those locations and decide whether it makes sense in your case to remove those references.
Getting folks to use it would be very difficult due to network effects however.
After having rewritten history locally (with Mercurial this is easy, powerful, safe and with a lot of tooling available) everything you need to do is doing an hg push --force.
For one, hg histedit with its curses interface (default since 4.9) is sweet.
Linus wouldn't claim to be a brilliant designer of user interfaces. It's totally conceivable that somebody could come along and develop a new way of talking about Git's functionality, implemented by a new "porcelain". Changing the vocabulary, creating a more comprehensible map of the internal logic of the thing...
In reality, git is actually far from ideal tool, with its own weak points and use-case scenarios which it simply not or badly supports. So it's never "just you" don't feel bad that it's "hard." On top of this, even the scenarios that it supposedly "good" supports demand sometimes totally "illogical" combination of the names and parameters.
git is used in spite of its flaws for different reasons. Some actions are really very fast, faster than by the competition. Sometimes that is a reason enough. Another is -- we have to use what our colleagues use. Even another: once you become familiar with it, even if you were aware of the weirdness, it can stop annoying you. Still to be able to realistically compare it with something else, you have to at least try it. And that something else too.
Anyway, once you get rid of the actual complexity in git, it's an easy step to work on the added complexity to the UI.
Unfortunately, the reflog is confusing and hard to use correctly in the case of an interactive rebase with multiple steps. It is hard to figure out exactly how far back you need to go in the reflog to get to moment before the rebase started if you want to start over. It also just so happens that its when an interactive rebase goes awry that I really want to reach for the reflog to fix the damage.
When you turn on Recyclable Commits, every commit in the reflog shows up in the history tree just like any other commit. You can see exactly where they diverge from your other branches and can work with them as you normally work with any commit.
Same thing for stashes: check one and it just shows up as part of the commit tree as if it were a normal commit.
I've used SmartGit for years and highly recommend it over the Git command line for the way it gives you so much more insight into the state of your repo.
When you run git reflog after rebasing, you will see lines like the following:
29d82ac HEAD@{6}: rebase -i (finish): returning to refs/heads/your-branch
29d82ac HEAD@{7}: rebase -i (fixup): Commit message 2
4f8e996 HEAD@{8}: rebase -i (pick): Commit message 2
f3a954e HEAD@{9}: rebase -i (pick): Commit message 1
f74b8a5 HEAD@{10}: rebase -i (start): checkout origin/master
The line listed after the one that has rebase -i (start) is the commit you were on before you started the rebase. If I screw up a rebase, then I will stash any uncommitted changes and run a git reset --hard to the commit listed below the rebase -i (start) commit I see in the reflog and start the rebase again.I see the stash as kind of like a private remote, in that I can freely put whatever messy or experimental or half-baked WIP I like, gaining the benefits of a commit without inflicting it on anyone else.
It is rather unfortunate that there is no convenient documented shorthand for "show me the reflog of the current branch" (`git log -g` gives you HEAD's reflog). That said, after a bunch of experimentation, it seems like `git log -g '@{0}'` will give you the reflog for the current branch. Apparently this works because e.g. `git log -g '@{2}'` gives you the reflog for the current branch skipping the first 2 elements.
"Pretending the average user will know how to get things back to how they are is silly."
The reflog is faaar more complicated to use that any of the day-to-day git commands.
https://www.git-scm.com/docs/gitrevisions#Documentation/gitr...
git branch local/foo
If you mess up too hard, check it out again.Why else are you going to go through the trouble of rebasing master if this isn't the goal you're shooting for? I'm a big proponent of commit history hygiene but even I can't defend rebasing master except for egregious things.
I think the only time I rebased master except for this was to fix a poorly executed mass file rename that broke git annotate.
Like you, I felt much more comfortable using git after learning this.
Yeah, completely unreachable commits have 30 days, even when you run git gc. The default reflog for a branch is even longer: 90 days!
I often try to reassure people new to git that "if you'll just commit often, there's basically no way you can lose work so that I can't help you get it back. Apart from deleting the whole repo folder. Don't delete the repo folder.".
I also have the habit of preventing non-fast-forward pushes on origin/master, which also helps when I can tell my team that they can't trash the origin even if they try.
# Merge commit
git merge --no-ff -m <message> <hash>
# Squash and merge
git merge --no-commit --squash <hash>
git commit -m <message>
# Rebase and merge
git rebase --force-rebase <hash>
From https://stackoverflow.com/a/52301456I believe in clean, linear history, and strongly prefer rebase-based workflows to merges. That's actually one of the reasons I chose Phabricator for my current place, as it is also very opinionated towards the same way of working.
Edit: oh, and to answer your actual question, the third one.
Years ago when I was just reading about git instead of using it, I saw sentiments along the lines of "always use feature branches and merge them so your thoughts and process can still be looked at later". In the last ~5 years or so I've worked professionally, I've not once wished I could reference intermediate commits in my own code or someone else's. I've found that ambiguities and clarifications can and are caught during the code review process.
I'd say another couple of benefits:
- It's relatively easy to teach git workflows when the log/graph is linear, and similarly it's _way_ easier to reason about your workspace when you have an actual production codebase.
- Merge commits can make certain operations like reverts and patches harder to reason about
One command I often use is git blame which allows me to find the commit that's associated with a particular line of code. Then I can look at the commit message and the diff against its parent. Perhaps what I'm changing may undo a bugfix and I wouldn't have realized it without reading the associated commit message.
When I look back at the history, I'm only interested in seeing the changes that actually made it through, not every intermediate alley and dead end in between. If those dead ends are significant discoveries/results, I document them elsewhere.
But either way works fine. It just gives you a different history. My team likes merging because they don't understand exactly what happens when rebasing. In that environment `git log --topo-order` is practically a necessity, though.
If you are merging master into a feature branch that has ongoing work and continues without merging back to master, that's the problem.
If your feature branch is short-lived, it can be easily rebased.
If your other branch is more like a release branch, with lots of work that can't be rebased easily, some times you can't really avoid a merge from master without communicating it first. If your team is large or distributed it might not be practical to say "release has moved to (rebased ref), please catch up"
In that case you should treat merges to release the same as merges to master (they should be finished bits of work that are considered published) and any unmerged features for the release, are kept on feature branches that are based on the release. They can be rebased after the point where master is merged back into release to avoid the nasty merge conflicts.
git merge —-squash <tree-ish>
When merging into the baseline git merge —-no-ff <tree-ish>
The reason for the latter is subtle. Yes, a perfectly linear history is nice in the aesthetic sense. However, the merge commits are artifacts of reviews that are useful in process audits, which QA and QC like to preserve. $ git co feature-branch
$ git branch before-rebase-feature-branch
$ git rebase develop> HEAD names the commit on which you based the changes in the working tree. FETCH_HEAD records the branch which you fetched from a remote repository with your last git fetch invocation. ORIG_HEAD is created by commands that move your HEAD in a drastic way, to record the position of the HEAD before their operation, so that you can easily change the tip of the branch back to the state before you ran them.
https://www.git-scm.com/docs/gitrevisions#Documentation/gitr...
It creates new commits and moves the branch to the tip of the new sequence of commits. No existing commits are changed or deleted.
- Summarizing history: squashing "implemented subfeature A.A" and "implemented subfeature A.B" into "implemented feature A"
- Rewriting history: moving commits around, changing the base commit, and so on
In my opinion summarizing history is acceptable, you're making a creative decision that certain information will not be useful in the review/when trying to understand the code in the future.
Rewriting, on the other hand, is essentially lying. You're creating repository states that never existed, and which you have never tested. In the worst case, consider the following history:
* F: (master) Merge branch 'component2'
|\
| * E: (component2) Fixed component 2's integration with 1
| * D: (component2) Merge branch 'master' into component2
| |\
| |/
|/|
* | C: (master) Refactored component 1's API
| * B: (component2) Implemented component 2 that depends on 1
|/
* A: (master) Base
Yes, it could probably be completely linearized, but that would be a horrible idea. Commit B will leave the repository in a completely nonsensical state. Sure, you could squash in E to mitigate it (since, luckily, nothing else happened in component 2 in the meantime), but then you're still stuck explaining what will likely look like a bunch of really weird design decisions compared to if you had designed against component 1's new API immediately.Historical context matters. If in doubt, don't rebase. Never `git pull --rebase` blindly.
No one is advocating that you use rebase on your public branches or basically any branch that has been "published". We are talking about feature branches or spikes or branches that exist just on one developer's machine.
Whether you've published the true history earlier is irrelevant to that discussion.
If I make three commits and then realize that I should have included something in the first commit, I use rebase to create a new sequence of three commits that has the corrected version of the first commit. I haven't shared those commits with anyone, this is just work that I've done locally.
Are you seriously advocating that creating a pull request with: (A, B, C, A-fixup) is better than using rebase and then creating a pull request with: (better-A, B, C)?
You think that second case is "lying" because I didn't show the intermediate step that included the mistake?
You can mitigate most of the damage if it is convincing enough (for example, go through B' and C' and make sure everything still makes sense at each point), but realistically nobody is going to do that, because it's pretty inefficient way to spend your time. And even then, you're still removing context (unless you're just fixing a typo).
> I'm advocating that you use rebase to improve the quality of your changes that will be reviewed before merging or even before being reviewed at all.
That was clear from the start. But the fact that X breaks Y doesn't imply that Y is a good idea when X doesn't apply.
Us mere humans make mistakes all the time. Typos, omissions, false starts, and so on. What is the value of throwing that raw set of events at a reviewer or complicating the understanding of the changes when viewed in retrospect from the future? What is the reason you call curating the work into a more polished form "lying"? Why do you think the time spent being intentional about changes isn't valuable when compared to the time spent by a reviewer (or your future self) to sort through the flotsam and jetsam of your intermediate work?
You can do it in a way that isn't harmful (as I mentioned earlier), but good luck getting a team to actually stick to that. It also doesn't help that pretty much no tooling encourages doing it properly.
I made a typo when writing this reply, and pressed backspace to correct it. Is use of the backspace key lying?
I think you're placing a value on "history" that doesn't map onto all users of "rebase", or E-Mail client drafts. A lot of advanced users use it as the equivalent of "save" in an editor, sharing all those intermediate states is more noise than value v.s. crafting a sensible patch once you figure out what you want/what change to make.
The true problems begin once you start creating commits that represent repository trees that you never tested or reviewed, for example by editing past commits (invalidating any testing you've done of commits after that point), deleting past commits (aside from squashing an unbroken sequence of commits, or deleting them if the squash would result in a no-op), reordering commits, or rebasing commits.
But yeah, the history you push to a canonical branch should generally be made up of commits that have all been tested in isolation. The rebase command doesn't make this worse, but better, e.g. with "rebase -i --exec='make test'".
I also prune out history of some false steps taken. Have you never written a program and done something like "I'll use a hash here <save><compile><test>, no actually a list makes more sense <save><compile><test> ...". Those intermediate steps are commits for a lot of advanced git users.
Sharing all your mistakes-as-you-go-along with the world doesn't help anyone, I'd typically be sending you a 100 patch merge request for some rather trivial change instead of 1-3 sensible commits.
I rather review the final patch series with changes in logical order and not necessarily in the order the code was written or with intermediate work that was later reverted or changed again. I do look at commits, because also every commit message counts and is supposed to explain the individual change.
The messy reality is valuable, when talking with your teammates about how the work was done and what kind of bumps were along the way. It's not about looking over their shoulders, it's about using data to develop together as a team, eliminating hinderances, etc - if you have the mutual trust to do that. And of course you yourself can go back and look for patterns of mistakes or problematic areas in code based on your history.
Like I said, to judge the change its, the whole PR diff is usually the most useful unit of inspection when you just want to see what happens. And if it's a big pr, you can of course always merge child PR's or branches against the big PR/branch, and look at the merge diffs.
The science principle of publishing your experiments, including failed ones, has the same benefits in sw engineering: others can build on your failed attempts, or save time by not replicating them.
Of course you do! And it's not am inefficient use of your time, because it helps reviewers now, and yourself when you're bisecting later.
> And even then, you're still removing context (unless you're just fixing a typo).
You place that context in the commit message.
And how do you detect that you forgot / was too busy to do it, when you go back 6 months later? It's fragile, "fail-open".
https://github.com/tianocore/tianocore.github.io/wiki/Laszlo...
Yes. A failed attempt is still a useful signal that people shouldn't try to simplify back to that way in the future (and why not). It's also a useful starting point in case the reasons it failed no longer apply.
Commit histories littered with commits that get back-and-forth reverted are frickin unreadable though. Extremely annoying to bisect, painful to comb through when looking for changes, noisy in git blame, etc. There's a ton of downsides for what in practice is very rarely even an upside.
Say upstream is at A.
I clone it in my local work-space, and make a few commits over the course of a few days. So my local is A B C
During this time other changes have been merged into upstream, so upstream looks like A D E
I now have two options. I can try to merge from upstream or rebase off of upstream. Merging introduces a messy commit history that quickly becomes difficult to follow. Rebasing removes my local commits, applies the changes in upstream, and then re-applies my local commits.
So after rebasing my local is A D E B C. There are no messy merge commits. And ideally, I can squash my local changes into a single feature commit, so upstream ends up incredibly tidy.
At no place in this process is there any dishonesty or lying. I haven't changed the history upstream, which is the source of truth. What's the issue here?
No, your local is now A D E B' C'. Commits aren't just a diff between two tree snapshots, they are tree snapshots.
Hopefully you test and sanity check C' before submitting for review, but it's very unlikely that you're going to give B' the same treatment, making it more difficult for people to understand the history in the future (as well as breaking `git bisect`).
And even if you do, are your coworkers going to? Consistently? No CI tool that I'm aware of will enforce this for you.
> There are no messy merge commits.
No, but the underlying messy workflow is still there. You've just swept it under the rug for the sake of aesthetics, at the cost of future comprehension.
> At no place in this process is there any dishonesty or lying. I haven't changed the history upstream, which is the source of truth. What's the issue here?
Those are completely orthogonal concerns. You're presenting a false version of the repository state.
The common mantra of "don't rewrite public history" is about not creating a mess of duplicate commits, it doesn't imply that rewriting history is fine as long as it's not public.
The exception to this is when a feature becomes more involved and has several logical steps, or any kind of history worth providing. This should be rare and when it happens, use merge commits to preserve history.
Not polluting the history with N trivial branches for every 1 branch that needs historical context, is a benefit of this.
Your commit history is just a somewhat arbitrary recording of your code at certain points in time that you choose. Rebasing simply makes that less arbitrary, allowing you to document the way your code is built up in a structured way. Rather than having to decide on the spot whenever a certain combination of code is a good candidate for a single, atomic commit, you can make that judgment with the benefit of hindsight.
Nobody is suggesting that rebase be used to change the history of a released or published branch (master, develop etc.). If that is your concern and the reason for you not trusting a "rebaser" then you are simply mistaken, you are arguing against an imaginary workflow for which no one is advocating.
Rebase should be used only to curate the commits on a feature branch and to keep the feature branch synchronized with the upstream branch.
* F: (master) Merge branch 'component2'
|\
| * E: (component2) Implemented component 2 (now with updated component1)
|/
* C: (master) Refactored component 1's API
|
* A: (master) Base
So we've lost B and D from your example, but who cares about those commits?That's on you if you don't test every commit. I don't care if you had failing tests (or even build) when you were writing your feature. I care that every one of your patches (and thus commit) does one logical thing, and that tests passes (making bisect useful).
You want your PR commit history to tell a coherent story. No one care if a writer had 15 bad draft of their story before publishing, and the same apply here.
And the cherry on top is that you can't easily tell later if you are looking at rewritten history. So if the above kind of rewriting might have happened in your project, you will essentially not be able to trust git history anymore as a record of engineering decisions.
This has been a common misunderstanding of git in the past, but thankfully is fading now. I was hoping it wouldn’t come back to haunt this thread. I don’t know where the extreme and hyperbolic idea of using git the way it was designed is “lying” and creating “false history” first came from, you aren’t the first person to suggest it, but it’s neither correct nor helpful to use that kind of language. This theoretical philosophical ideal that there’s only one true history is trading away things git was specifically created to do, as well as the practicalities of real world software development, in favor of a strange unrealistic and abstract notion that once git commit has been used the commit should never be touched again.
Everyone knows and agrees that rearranging already published commits is a bad idea. Not because it’s “lying”, but because it causes problems, costs other people time, and can even inflict irreconcilable merge conflicts on their work.
Cleaning up your own commits before you push using interactive rebase is not just a good idea, it’s the way git was designed, it’s what Linus does, and it’s kind to your team. This includes reordering commits and pulling with rebase.
> Historical context matters. If in doubt, don’t rebase. Never `git pull -- rebase` blindly.
Maybe you could back up your assertion with some examples of why it always matters, and why that justifies using words like ‘never’?
Your rhetoric is ignoring the real-world fact that on a large team, the majority of commits at any given time are orthogonal to each other, and that the parent commit you end up with is completely arbitrary.
Not only do I use pull -- rebase, I always git config --global pull.rebase true, and I frequently recommend others do the same.
Having merge commits in master every single time someone checks in is incredibly noisy and it inflicts friction on the entire team to force everyone to read the noisy log. I’ve always worked on teams that decided to take the more practical approach of one-off commits should not have a merge, regardless of when they happen, to keep history cleaner, and feature branches with more than a couple of commits or by more than one person should have a merge commit, to keep the master branch from having broken commits or unfinished features and so it’s always bisectable.
In your local repo yes, for next 30 days. Then git will garbage collect them. But when you force push new pointers to GitHub, the GitHub repo will lose access to the unreferenced commits right away.
But of course those 30 days will give you plenty of time to go back, if you changed your mind about the rebase.
> If you've made a mistake and in so doing lost commits which you needed, then git reflog is here to save the day.
Seeing that sentence in the article immediately makes me think the author has poor git practices or even understanding. reflog is useful at times, yes. But it is not what saves beginners from losing work while practicing their rebasing skills. What saves them is that they didn't rebase an important branch. They made a temp branch, pointing at the HEAD of their important branch perhaps, and they screwed up the temp branch.
If people would just make this one thing clearer to beginners, it would help a lot of people learn git more easily and with less fear.
The full version of the command as
$ git rebase --onto new-base start end
takes the commit range (start,end] and re-commits them on top of the new-base commit. The commit range doesn't have to be a full branch and you don't even need to be on the branch to run the command this way. It's very intuitive and I nearly always use the full version now.I've also gotten into the habit of "pinning" my branch before I rebase so that I have it in its original form. If the branch name is my-branch, then the command
$ git branch my-branch{-hold,}
which is a handy (bash-specific?) shortcut of $ git branch my-branch-hold my-branch
leaves you on my-branch and creates a new branch label called my-branch-hold that points to the same place.EDIT: clarification of pre-rebase branching
git branch my-branch-mulliganForce pushing branches is what you do when you have pushed a branch that you expect to modify. Why would you do that? Because that's how Github and Bitbucket have taught people to conduct PR's.
If your immediate reaction is that "rebase is bad UX", ask yourself whether or not pull requests are good UX. I honestly think rebase is great, but that pull requests are extremely bad UX, and the UX blame is misplaced on rebase when where it really belongs is on pull requests.
I thought that Github and Bitbucket encouraged people to push up additional commits to fix issues in their PR. So, a typical PR will end up with a commit history like:
Implement a feature method
Add calls to new feature method
Update to version 1.2.3
fixing missing semi-colon
addressed comments
one more thing
now its working
People who force-push are the ones who are trying to keep a clean commit history (meaning you don't have those extra 4 commits). So, your point about a PR being bad UX versus a rebase is correct, but not for the reason you state.Ideally they'd use git commit --fixup=<sha> -p, and then git rebase --autosquash when the maintainer approves the merge, but few people care.
(And what I'm saying isn't really accurate any more since GitHub does show force-pushes in the pull request UI these days and one can run git range-diff on that. But this wasn't possible last year.)
The idea that git is good because it is difficult to use is just "git snobbery", as is the idea that it must be difficult because it's a DVCS.
There's nothing to stop git having two levels of the API, one exposed for tools to build off of with the full complexity and another for every day use.
And for anyone who doesn't know the primitives, I posted a brief summary on Mastodon the other day:
https://cmpwn.com/@sir/102038690003388821
Also recommend Pro Git's chapter on git internals.
I don't need to know how postgres does paging, indexing, tree diffs, etc to be able to write good SQL.
I don't need to know how typescript compiles to javascript to use typescipt.
I don't need to know how my engine works to drive my car.
I use all of those every day.
But the suggestion here is that as every day users of git should learn git internals to be able to use the tool better.
That's a tooling failure. It's not a failure of the git design or git fundamentals, it's a failure of the git cli.
GIT COMMANDS
We divide Git into high level ("porcelain") commands and low level
("plumbing") commands.I'm not sure if I do, but that's the conclusion the man page + your comment would seem to support.
Things could be better. There's an existence proof. It just lost the mindshare war and so now we're stuck with Git, which I still have to look up basic syntax for because its command set is contradictory and makes no sense. (Is it git <x>? git <y> --x? git <z> <a-b>? Something else entirely? Who knows!)
(I've also had to work with various IBM CVS, and they are universally garbage. When I get frustrated at git, all I have to do is think back to those.)
So yes, Mercurial is better, but is it worth the effort? Not in my experience.
That being said, I use git daily and find that I'm able to do everything I want and more, so I'm not looking to make a switch. Unfortunately, most people don't care to learn how to use git beyond "checkout, commit, push, call for help".
I forget all the reasons why RTC isn't great, but the main one: if the server goes down, you're screwed. This happened several times, and we basically went to the pub instead of working. Slow to check out. Streams sucked compared to branches (especially when the server admin restricted creation of streams, meaning you simply could not branch at all if I'm remembering), and the capability to stash changes/switch branches to work on different work items if one was blocked was also more cumbersome. Code review was terrible.
A centralised paradigm does simplify things a lot mentally, but the workflow suffers IMO.
I guess you can pull a branch or email a patch and diff it with the diff tool of your choice.
There's a few options for gir review UIs, e.g. Gerrit or GitLab. Kind of unix-y, just have the VCS be a good VCS.
> a constructive proof is a method of proof that demonstrates the existence of a mathematical object by creating [...] the object.
> This is in contrast to [an existence proof], which proves the existence of a particular kind of object without providing an example.
I don’t think the argument is that the difficulty is a virtue, nobody is trying to be snobby, so try to avoid jumping to that conclusion.
Git just has some inherent complexity. Git does have a steep learning curve that is the root of a UX problem. But it’s not clear what better abstractions there are or how to simplify git. Lots of people have tried to make a higher level porcelain, and the issue isn’t going away. You are welcome to suggest & create a git wrapper that makes it simpler and less dangerous.
Perforce is easier to learn, so you might try using that instead. I use both and I’m becoming more and more frustrated with Perforce because git is so much more flexible and safer and easier to use once you learn how to use git.
For the former you want flexible history, distributed repos and freedom to do whatever you want.
For the latter the history should be sacrosanct and the repository is better be more or less centralized.
git tries to sit on both chairs and therefore has to adapt a quite awkward position.
I also don’t understand your larger point about git and what the problem is. For almost everyone using git, the pushed history is sacrosanct, and the main repo is centralized. The main workflow for rebase is to clean up before making commits public.
"Version control" is a log. I should be able to return to any point in history at any time. I should not be able to destroy this history no matter what I do. The history should be backed up in a remote location.
With git every once in a while these two come into conflict: I'm using my branch as a change management system and then someone else pulls that branch and makes a couple of changes on it. Then I force push and the mess begins.
Regarding centralized repository: say, I have two working copies. In goode olde subversion there was the "master" version in trunk and two local versions, one per working copy. Three versions altogether. Pretty clear which is which.
Now in git I have:
- master in origin repo
- origin/master in wc1
- master in wc1
- actual file in wc1
- all the same in wc2
Seven potentially different versions of the same exact file. That's even without mentioning a stash.
> Then I force push and the mess begins.
That is your problem right there. Force pushing over published history should always be avoided. Don't do that, it is very unfriendly to others you work with. Just always pull, resolve any conflicts, then push your changes without forcing them. If you need to force in order to fix a serious mistake, then notify everyone first and have people hold their changes, then pull everything, resolve the problem, force push, and notify everyone again to "force pull" by first fetching, then reset their branch to what's in the origin verison; using reset --hard will let them avoid having any conflicts after you force pushed. Consider carefully whether the mistake even warrants a force push, or if you can make do with new commits on top that fix the bad ones.
> Regarding centralized repository [...] Seven potentially different versions of the same exact file.
You're conflating multiple different topics. The existence of multiple copies of a file isn't related to which repo is the central one, nor is it some kind of problem.
The origin repo is your central repo. Your downstream repo has to copy from the upstream/central repo if you even want to work on the code. What you called "copies" in origin/master and master are branches, not copies of the file. The only single copy on your machine from your point of view is your local workspace, which expanded from your "master". stash is something that happens behind the scenes to your git database, it's not making more working copies. wc2 is another repo or computer, it's not even relevant. None of these copies you're talking about are visible to a user except the one working copy.
That is my problem right here.
There is no way to tell the difference between "I push to make something public" and "I push to back up my data in some safe location".
That said, making backup data in a safe location isn't what git was really made for. If you really don't want history to be there, you can use cp or rsync, or you can git clone -depth 0 from the backup machine, or just use backup software. There's "bup" which is a backup tool based on git...
Git is good because it is extremely elegantly designed. There are blobs, trees, commits, and refs. When you understand them, you understand pretty much everything about git.
And yes, the interface is a mess.
But once outside that world, the user if faced with the problem that "rebase" is just a special case of "merge" and shared all the complexities and edge cases. And that's hard for fundamental reasons. Git has tools for this too, but their interface shares the complexity of the problem domain.
My last few workplaces were either not git, or relied on git pull w/o rebase. My current workplace has rebase as part of their flow, and I find that, like, 60% of my PRs require force push, which bothers me greatly. Everyone else just shrugs and considers it part of business, but I know what I WANT to do should be nicely aligned and not encounter that problem.
Unfortunately, every explanation I'm like "yes, yes, branching trees, I get it..." and then I'm suddenly in the "...and it says things are different and I don't know why". And because this always happens when I'm trying to get some fix in, I never have the time to study it to figure out what is really happening. It's just "--force and promise myself the next time will be different".
Not really sure what that means as rebase can be used in multiple ways, but you might want to try using:
`--force-with-lease` instead of `--force`
> This option allows one to force push without the risk of unintentionally overwriting someone else’s work
To the degree I understand it, sure:
We have master branch A
I create a feature branch B
Both get updates. Someone will do a rebase of B to the most-recent A and push that. (In my previous workplaces they would have just pulled the most recent A)
Here's where the confusion comes in: If I get the updated B but A has updated again, I cannot pull A nor rebase to A and successfully push the result without forcing. IIRC, on push it complains that my local branch is not up to date, but if I pull it will tell me I am up to date.
At least, that's what I think is the timing - since this involves multiple people I'm uncertain of what exactly occurs and the order, nor why problems are inconsistent. We don't have that many feature branches that have multiple people contributing AND requiring updates from the master branch, but it happens often enough.
FWIW: the line noise you typed means "Make a list of the changes since three commits ago, let the user edit the list, then apply them". Other than -i step and the length of the list, this is basically a noop -- you're rebasing on an ancestor of HEAD! I mean, yeah, git lets you do that, but I don't know why you expect the syntax for irrelevant nonsense to be simple.
But let's humor you and try to use that syntax for something real. If you typed "my_version_tag~3" it might make sense -- you want to back up to the commit before whatever automation might have added for a release and put your current work on top of the earlier tree as if they had been developed as part of the release. And you have some junk in your current tree you don't want to expose to the customer to whom you are going to hand this test tree, so you want to remove it interactively.
That... sounds like a useful trick. But it's complicated. And the syntax is complicated. So what's a good syntax for the previous paragraph's action?
The "line noise" I typed is an extremely common command that is used to clean up commit history, e.g. squash all your "WIP" commits into a few nice ones. Yes the idea of rebasing onto the same branch is just weird, that's kinda my point, since it's the only way to edit history that git supports (as far as I, or anyone I've ever seen answer a question about how to do this knows).
Git doesn't have "reorder patches" feature. Maybe it should. But the fact that its rebase tool can be abused to do this doesn't say anything about the interface value of "git rebase".
Honestly, it seems like the root cause here is that you don't actually do branch rebasing very often, don't see the value of having a rebase tool in the tree, and are just complaining that the trick you do need isn't well supported by the rebase tool you don't use or understand.
My complaint is that git's UI is terrible to teach people, to understand it you have to understand way too many internal details of git.
Yes, making a wrapper for `git rebase -i HEAD~x` (and `git rebase -i <other thing that refers to a commit above head>`) would satisfy this UI nit. It wouldn't satisfy all UI nits, this is just a relevant example.
As for why not submit it myself, I'm sure I'm not the first person to complain about this, drive-by UI changes to a project are the absolute best way to get a ridiculous inconsistent UI. I don't submit it myself because I'm not willing to commit the time to become a core maintainer of git, and without being a core maintainer I don't feel right trying to push UI changes in.
Huh? That is exactly what git rebase -i is. This isn't some kind of "abuse" or "trick", this is precisely what interactive rebase was designed for. What is making you think that rebase is only to be used when merging two different branches?
Also, emacs magit makes this even easier, if you don't mind a more GUI-like interface.
The only difference is that one's interactive and the other isn't. If you use `-i` and simply exit out of the editor, the effect is completely the same as not having used `-i`, isn't it? I think moving `rebase -i` to a completely new command could potentially make things more confusing. That new command would be an extension of `rebase` and so could be used interchangeably.
A hypothetical new command would be a simpler version of "rebase" that comes with the restriction described above, that it's not actually changing base.
`rebase -i` doesn't restrict that, does it? There may be people whose workflow includes things like `git rebase -i --onto foo bar baz`. That you don't use it is another matter.
> A hypothetical new command would be a simpler version of "rebase" that comes with the restriction described above, that it's not actually changing base.
You want to remove features from git? Why?
If you didn't mean to say that instead of having `rebase -i` we should only have this hypothetical command, then you can do:
git config --global alias.edit-history 'rebase -i'
Though you may want to add to that to make it impossible to use edit-history to rebase, too. I mean, you did say you wanted the restriction, right?Your right when you say I don't "just" want to introduce an alias because I want to restrict the arguments.
The other issue with that solution is I don't just want a solution for me. I know how this works now, I've already memorized the magic incantation to edit history and later spent the time to understand why the command does what it does (the same goes for basically all the other common git commands). What I want is a solution that works for everyone, out of the box, so we can stop wasting time teaching git internals.
I generally prefer transparent systems to those that try to figure out what I really want to do.
I'd never use it on a public or shared branch, but `rebase --interactive` and `--exec` are some of my most-mused git aliases.
also: 'git push --force-with-lease' (worth repeating)
https://stackoverflow.com/questions/30542491/push-force-with...
Apparently it will think you know about the changes if you've fetched them, even if you've not merged them. And some systems auto fetch in the background.
Also, unless the diagram is confusing with the alignment, the "rebase to rebase" example seems to be implicitly assuming --onto, because the last common ancestor of 'master' and 'feature-2' includes 2 commits on that branch: the first of 'feature-1' and then the one that is only on 'feature-2'.
One thing I do in our team is provide git novices with some useful aliases, one of them being `fpush = push --force-with-lease`. It's saved a few people losing work along the years
Put those in your ~/.gitconfig:
[alias]
fu = "!f() { local msg=\"fixup! $(git log --oneline -n1 | cut -d ' ' -f2-)\"; git commit -am \"${msg}\" && git rebase -i --autosquash HEAD~2; }; f"
fuc = "!f() { local msg=\"$(git log --oneline -n1 | cut -d ' ' -f2-)\"; if [[ \"${msg}\" != "fixup!"* ]]; then msg=\"fixup! ${msg}\"; fi; git commit -am \"${msg}\"; }; f"
xx = "!f() { git reset --hard && git clean -f -d; }; f"
From now on:git fu = (git fix up) combines last commit and all staged changes into one commit with the last commit message
git fuc = (git fix up comment) commits all staged changes as a new commit with the last commit message, but prefixed with "fixup!". Except if the last commit message is already prefixed with "fixup!". Now you can work on the same thing but commit incremental steps. In the end you just do: git rebase -i master (or similar) and they will be all in one commit.
git xx = get rid off all unstaged/uncommited changes in current directory. Very destructive, very useful.
Last but not least. If you have a lot of "fixup!" or "squash!" commits and need to interactively rebase without autosquashing them do:
git rebase --no-autosquash -i master
fixup! Exact title of commit
We actually use that workflow during PR reviews on Github. That is, someone comments on a PR and the person who opened the PR will make a fixup commit with the appropriate title and reply to the comment stating that it was addressed and link the fixup commit by its abbreviated sha1 value.At the end of the PR review, the person who will merge the PR will run:
git fetch origin
git rebase -i --autosquash origin/master
git diff @{u}..
This basically updates the remote tracking branches (including origin/master), rebases the feature branch on top of the remote tracking branch for master, and then checks if any changes were introduced in the rebase process by running a diff against the remote tracking branch.If there is no diff or the diff only shows changes made on the upstream master branch since the feature branch was created, then they can run:
git push -f origin feature-branch-name
and then merge the PR.With this in mind, acceptable curation of the history for me is squashing or separating commits and neither of these require rebase. But I wouldn't object if a colleague chose to use rebase to accomplish this.
I would, however, object to reordering commits or rebasing since this alters the chronicle.
Which I guess leaves me with my opinion on rebase:
You can, but you don't need to. If you are going to, be sure you understand what you are doing and don't alter the chronicle.
A clean git history is not a vanity project. It can be used as a tool in further code building.
As mentioned by another poster, the fact that a whole website is needed to explain the concept illustrates the design and UI failure.
This is an uphill battle that can't be won until a next-generation interface becomes usable by mortals. If that can't be done due to complexity, it's a lost cause for average developers paid for delivering business value.
Your comparison is not apt in many ways. For one thing, open-source governance is extremely different than managing private codebases maintained by a single company. Nobody in open-source governance is reading git histories to figure out whether employees are struggling, who is performing, who is not performing, who is overworking, where the project is, whether or not people are duplicating their efforts, how to report progress to clients, etc. And besides, the Linux kernel uses an email-based, merge-oriented workflow anyway. Kernel patches are submitted via email. That's not at all representative of any company that I have ever worked for or any company that I know of. The Linux kernel history also has a network graph that is literally unviewable on Github because it's so complex. Again, not representative of 95% of the work for 95% of users on 95% of projects.
I'm not saying you shouldn't care, just that I don't believe this strategy will become mainstream unless it gets much easier to understand and use.
But maybe you use your history for something else that I haven't considered.
(--fixup type commits aside).
Practical benefits from not squashing history:
- Can bisect to find bug introduction.
- Can annotate/praise/blame to find who/when some change was made.
- Adam Tornhill's "Code as a Crime Scene" argues that it'd be beneficial to consume VCS history to provide health metrics on the codebase. (e.g. use VCS to check which sources have many contributors (thus potentially high defects), or check for "lost knowledge" from developers who have left).
- Can build/run an older version of the software.
But is there really a big advantage from putting time into maintaining a sequence of commits? EDIT: Ah, I see another comment point out that "maintaining a nice history" tends to mean fixing very borked commits. That makes sense. :-)
The fact that I had a bunch of stupid typos and broken tests that I didn't realize were broken before I committed doesn't need to be in the final history. What I really want for the preserved history is the conceptual chunks of changes I made along the way.
If you have a safe work atmosphere, and yor teammates reviewing the work can discover pitfalls in your project's workflow, you as a team can have discussions about it can improve your test stuff. And you can maybe go back through the history and see how many times this kind of normal human mistake with other branches and developers.
You can still diff through the PR as a unit before merging, without getting bogged down in low level commits.
Code and by extension history should be easy for humans to read. For that reason, the signal is very much worth keeping and polishing, but the noise is not. Documentation of false starts, appealing but ultimately problematic design choices, and “why” information belong in comments, commit messages, or design documents — explicit rather than implicitly littered around the history.
If you submit a PR to my project on GitHub and it consists of 20+ broken, non-atomic commits leading up to the final one, I'm going to ask you to clean them up and squash into one.
EDIT: I meant to include this link, which is a pretty brief yet thorough explanation of when/how/why to rebase: https://www.atlassian.com/git/tutorials/merging-vs-rebasing
Also, I've this gist on how to think about git: https://gist.github.com/nicowilliams/a6e5c9131767364ce2f4b39...
• In a number of code blocks, angle brackets are not escaped, and so text doesn’t appear in the end result.
• `<span style="color: maroon">#include</span> <stdio.h>` lacks the semicolon
• The paragraph immediately after the #conflicts heading lacks its <p>.
I also think it would be worthwhile mentioning in the first footnote a hazard of empty commits: that they will be dropped by default when rebasing.
Am I missing a config setting you explain somewhere?
You are probably thinking of 'git log', which uses the reverse order with the newest commit at the top.
Not sure why I was convinced otherwise, must be log's behaviour as you say. Had to test to believe it!
Still, rebasing is an important workflow to master even if you work on GitHub or on shared branches with others. If you mess up and force push to the wrong branch, you can always go back to the reflog[1] and check out the old version, then force push it.
[1] https://git-rebase.io/#reflog
And in the immortal words of Doug Gwyn(?): "UNIX was not designed to stop its users from doing stupid things, as that would also stop them from doing clever things."
In a rebase-oriented workflow, every commit has exactly one parent, and the commit history is entirely linearized. Merge commits have more than one parent, which means going backwards through the history means that you have to essentially navigate the branching structure to navigate the history. That makes it pretty hard to do!
A linear history is comparatively easier to reason about than a history with all of its branches. It's also, in a sense, less "true"! So it depends what you care about. Do you care about telling the exact story of everything that happened to everyone on the project, or do you care about the history of changes to the one shared copy?
Linear histories also make it easy to step backwards through the history one commit at a time. I personally just do `get checkout HEAD^` over and over, walking through the history in reverse, since most breakages are noticed within a few commits of their occurrence. I really like being able to do that. A lot of people think that's useless!
If you have one copy of the project that is considered authoritative, and all developers are synchronizing via that single authoritative copy, creating a clean and linear history is possible.
I didn't make it past this line. git is extremely good at keeping history immutable; it's also good at creating new, alternative histories and moving between them.
You mentioned "altered" history, which to me isn't much different from "edited" or "modified", though maybe that was just a slip and not your preference?
There does need to be some term for it. If I push a reordered branch, I clearly did more than nothing, even if we don't want to call that modifying history.
I do question if any term will work perfectly, or if there is any full solution other than one learning git to an intermediate level before discussing it. The problem, as I see it, is that there are (at least) two concepts of history, there is the DAG of commits, where ancestors are older, and there are branch-heads where you can change what they point to in unrestricted ways (pretty much the reflog). I'm not sure English has a really good metaphor that's going to capture everything.
Quoting from this document below.
# Rebase topic branch on the main development branch (optional).
git checkout TOPIC-BRANCH
git rebase master
# Edit commits, e.g., last 3 commits in topic branch (optional).
git checkout TOPIC-BRANCH
git rebase -i HEAD~3
# Force push rebased/edited commits to the pull request (optional).
git push -f origin TOPIC-BRANCH"Beginners to this workflow should always remember that a Git branch is not a container of commits, but rather a lightweight moving pointer that points to a commit in the commit history.
A---B---C
↑
(master)
When a new commit is made in a branch, its branch pointer simply moves
to point to the last commit in the branch. A---B---C---D
↑
(master)
A branch is merely a pointer to the tip of a series of commits. With
this little thing in mind, seemingly complex operations like rebase and
fast-forward merges become easy to understand and use."Suddenly everything makes sense now!
This works really well but I rarely see people talk about it. It means you don't have to use something complicated like github or gitlab to keep up with rebases. It works for everyone.