Git Workflow Basics
blog.codeminer42.com
blog.codeminer42.com
Rebasing can cause the loss of history and developers should be as careful with it as system admins are with `sudo`. I can't recommend any workflow that includes it without treating is as a terrifying and scary thing. How easy is it to accidentally remove a line during interactive rebase and lose all work associated with it?
This is why my team and I moved to squash merging. Sure it has it's own drawbacks, but they're far less worrisome than rebasing. If you screw up a rebase, the history is re-written or force-pushed by accident. If you screw up a squash merge, you can still check out the intermediate commits if you know the hash.
We won a Ruby award for our work on Git Reflow. There are big improvements coming this week that can make it easy for teams to tweak the workflow to suit any special needs you might have. It works on github and bitbucket and automatically creates pull requests (and makes sure they're reviewed.) Gitlab support coming soon (maybe this month).
We need more articles like this! Thank you for your work!
Until the original branch is deleted and the refs are garbage collected, anyways.
It seems strange to me to advocate for squash merging on a premise of not losing history. A squash merge is a rebase.
> A squash merge is a rebase.
Well, it's not a rebase in the literal sense: the base commit isn't changing. It _is_ modifying history, though.But my point was just that a merge squash is just a specific incantation of the git-rebase tool. And it is one of the most history destroying incantations, rather than the least.
C -> D -> E
/
A -> B
then you "squash merge" S------
/ \
A -> B -> F
where S is C + D + E.Whereas a rebase + (now fast-forward) merge would be
A -> B -> C' -> D' -> E'
and, if squashed during rebase -i or by some other means A -> B -> S
It seems you're saying a "squash merge" is the last one, I thought it was the second one.Also it's entirely possible I'm wrong, I'm a fan of merge bubbles (sometimes rebased for clarity), and avoid large squashes in general, so maybe I just don't understand how people do it. I don't know why you'd bother to keep both a merge and squash commit around, though.
I don't know why you'd want to do it that way either! But I also don't understand why you'd want to "never rebase", so when it comes to git, I assume that someone has reasons for doing all kinds of things I don't understand.
(One thing I really love about git is that it's a tool I use extremely often (it's something like, 20% of the commands I run in my terminal), yet I learn something new and useful all the time.)
I wouldn't recommend it to others as I don't trust them to read the man page and understand what git-rebase does. Those of us who use git-rebase also know how to recover the refs since before they are GCed, though I've never had to do that.
It's dangerous to pronounce certain powerful features of a system as off limits for day to day private use. That's a different pronouncement than a decision not to share "why everyone should use git-rebase often".
An exception can be made for topic branches, especially in a pull-request workflow. These branches could be rebased / amended to update the final result, even after they have been pushed already.
That's authorization.
> anything accessible without any authorization.
Even if it does require authorization, it's considered public in regard to this discussion.
Why is this a conversation?
It does not matter if you pushed it to a private repository where some other people still have access to; they might have fetched that history and committed on, which causes unexpected results when you rewrite such a branch.
I tend to use published history / commits for this reason to make it less confusing.
And I also mention in the post that we should always use what's best for our teams, so if the squash merge works best for you, go for it! :)
This is not only not easy, it's actually very difficult. If you drop something in an interactive rebase, you can reset your HEAD to the HEAD commit from before your rebase. It's a bit arcane, and has its own dangers, but it's also important to be clear that rebases are not destructive unless a git gc runs between your rebase and realizing you made a mistake. It's also the equivalent fix to checking out intermediate commits for a squash merge, as far as I know.
Don't get me wrong, I don't think rebase should be the first tool you reach for and I don't particularly like “rebase everything to master” workflows. But it's not as dangerous as you're making it sound, IMO.
This is what Mercurial Evolve tries to solve. There's nothing wrong with rewriting draft commits. The only potential problem is rewriting public commits. Mercurial uses phases to distinguish drafts from published commits and you may optionally designate certain repositories as non-publishing, so that they can be used for collaboratively editing draft commits.
A similar de facto convention on git is to only rewrite certain branches (e.g. feature branches) but never rewrite others (e.g. master). Commits that are local-only can be rewritten at will. Mercurial just codifies this convention via phases.
git diff origin/master branch
This way you can see if the diff looks like what you expect it to be. If it doesn't and you believe you messed something up while rebasing just
git reset --hard origin/branch
and redo the rebase
Our workflow is:
- locally, commit to local master or a local branch
- occasionally checkout local master (if necessary) and pull using the 'rebase after fetch' option.
- if we had local master commits, fixup any conflicts in our code
- if we were working on a local branch, checkout that branch and rebase it on the new master, fixing any conflicts that arise
- If our work is complete, possibly do a final rebase to reorder and squash local commits, and fast-forward merge to master if we were on a branch. Finally, push the local master to share our work.
Note that we never rebase anything that's been pushed. Also, if we're worried that a rebase is potentially complex and error prone, we create a new branch at the existing HEAD so that the old commits don't get lost in the reflog. Once the rebase is done, we can delete that branch so the old commits can be garbage collected.
Getting rid of merge commits is literally the only benefit of your workflow over the standard branching model.
This may sound counter-intuitive, because you're thinking it's the same conflict either way. Most of the time it is; if you and someone else changed the same bit of code, the conflict will be shown to you the same way whether you merge or rebase. Those are easy to fix. What's harder is when someone reorganizes some code without making significant changes to it. In a merge, you'll see changes all over the place, but in a rebase git can usually figure out the new line numbers, and may not indicate a conflict at all. The other area where my team has had difficulty with merges is in Visual Studio sln and csproj files. When you add new projects to a solution or new references to a project, git can present a very confusing diff during a merge conflict. But for rebase only your additions are highlighted, and most of the time you can solve the conflict with "use theirs before mine".
The fear of rebase seems to always come from its ability to "delete your work", but that fear is almost always unfounded and based on a lack of knowledge on git's internal structures. Of course, one of git's biggest and well recognized faults is that its UI makes no attempt to alleviate those fears.
When I'm dealing with my own work branches -- I absolutely rebase my commits before submitting a pull request to my team. I'll either squash irrelevant commits or I'll reword poorly written git commit messages. I also almost always default to `get pull --rebase` as well, as I hate the noise of merge commits littering up my commit log.
https://git-scm.com/docs/git-reflog
> …or force-pushed by accident
Well, you could nuke git folder "by accident" as well :) Jokes aside, don't force-push to "other's" branches).
git push --force-with-lease origin/other_branch
so you at least won't trash any commits your collaborator pushed after you last fetched their branch. git will see that your origin/other_branch doesn't match origin's other_branch, and will bail out so you can pull the changes and decide how to incorporate them.if you're going to store commit hashes externally instead of using the reflog then it makes no difference whether you use rebase, merge, or even cherry-pick or reset.
also, sudo itself poses no risk; it's much more important to evaluate what you're running, instead of a blanket restriction on what is just another tool. if I run "cd /; rm -rf *" on my desktop, it doesn't really matter whether I'm running as root or as my user, I'm going to have a bad day. "curl | sh" is equally as dangerous as "curl | sudo sh".
Git rebase is absolutely a vital part of the developer toolkit. If you're not using it, you're missing out on a big timesaver and git feature.
I think you must be talking about rebasing public commits/history. In the article he is specifically talking about private history and has a nice big warning against rebasing public history. Linus has explained the distinction pretty well before: http://www.mail-archive.com/dri-devel@lists.sourceforge.net/... .
Wouldn't your tests catch when lines get elided by manual merge conflict resolution? Have you see mjd's Git Habits[0], which lays out a foolproof way of rebasing without losing lines?
I didn't mention that in the post because it's intended for beginners and adding that info there could maybe be a little too much :/
But what if you need to do ship a full pipeline of a branch for multi-repository project? I clarify the issue better in my post.
The model described in this post is basically the model used with svn, and it brings with it its own vast set of problems that active uses of topic branches are intended to resolve. We've been down this road before and it's not roses and sunshine.
* Want to see the history for a specific feature? Impossible in your proposal, native in a feature/topic branching model.
* Want to do a code review on a specific feature? Again, impossible in your proposal, trivial in a feature/topic branching model.
* Want multiple developers to work under the same codebase with minimal conflict resolution and clear separation of tasks? Very hard under your proposal, easy in a feature/topic branching model.
You also say that using topic branching means not doing CI. I've used the idea of proper feature branching for years, and have never _not_ had a CI process. CI tools are most definitely ready for (and are quite welcoming of) workflows like git-flow. I'd be happy to speak at more length about how we implement it if you'd like, but I assure you it is all but complicated.
about the history for a feature: most branching based workflows prefer squashing commits when merging, so most probably the history is lost whatsoever.
you're right I didn't explain how I do code reviews - and, effectively, I usually prefer pair programming when pushing, and after-the-commit code reviews - but before the feature is toggled on by default.
There's a reason for this, I usually say that a review should review the status after a merge, not just a change; many a times I've seen reviews for a PR that miss the whole point, along the lines of "the change is good, even though the resulting merged code is complete mess"; on the contrary, if you review a certain commit before toggling a feature on, you're basically declaring that the code, at the point, is basically good. Yes, it may be hard for a large codebase, and often reviews are done by looking at what changed, and not at everything.
But I've seen many, many, many stupid errors done or overlooked because people just looked at the change and not at the whole code after such change.
I can't really explain the depth of how pleasant, and how surprising, this pleasant surprise was. Thanks!
But if you're rebasing your commits, haven't you lost that? The concerns about a "clean commit graph" seem more aesthetic than functional.
> more aesthetic than functional
It is easier to understand a cleaner history than a messy one. > haven't you lost that?
You've lost the ability to look through one kind of history, but not other ones. bisect still works if you've rebased.You cannot get things perfect on the first try; this is part of the whole principle of code review. When my patch is perfect, except for that one little typo, what should be done? Is a history with two commits, one amazing, one saying "fix typo" with a one-character diff, or one commit that's perfect, an easier to understand history? What is actually lost by "throw[ing] away history that you can never get back"?
If it had been right in the first time, that history would have never even existed in the first place. So you end up with the exact same thing.
If I were bisecting that repo, it's a lot easier and more useful to be able to point the finger at the one commit that actually changed the line, rather than having to parse the one monster squashed commit to find the one line that introduced the bug.
Furthermore, if this is a PR that's open, then the "bug" would have never even landed. So looking through history to "find what caused the bug" would have not even been a thing.
When writing a feature, I use git to save my progress, and once I'm done, I would like to present the feature as a clear and complete changeset. A commit is me saying, "These are the changes I'd like to make to the codebase", and I feel that argument is easier to make when I present one thought, rather than the dozens of thoughts I had on the way.
If I were perfect, then I would commit in a way that was one cohesive thought, but I'm not. My commits are often, "Did the thing", then "redid the thing, with better testability", and "Re-redid the thing, fixing some fundamental bug in how I did the thing at first", etc.
When we write bigger features we always need at least two dev. One writing the front end (HTML templates) and one the backend (whatever populates the templates, makes sql query). We both need immediate feedback. I design the models around the templates so i need the templates at least partly to work. He needs the models to properly do his work either (if he wrote the templates before I write the models, the development of the whole feature would slow down and it is not nice to only work with a lorem impsum all the time. We also get very detached from the actual feature that way.
How do you manage those situations? Just do one branch per feature and if that feature requires more work, then just let two people work on that feature?
What is a good workflow around git?
At work, we use deploy branches for a few repos, and integration branches for others. Personally, I use rebase for my own projects.
https://www.atlassian.com/git/tutorials/comparing-workflows/