Take Advantage of Git Rebase
about.gitlab.com
about.gitlab.com
One reason to regularly push commits can be to trigger CI/CD, security scanning and dev/staging deployments or review apps.
GitLab team member here - before joining GitLab in 2020, I was a Git/GitLab trainer. https://www.netways.de/en/blog/2018/05/24/releasing-our-git-...
Thanks for the pointer into learning something new :)
And I am currently waiting for my GitLab backpack for contributing something (very small, but apparently it counted!) so thank you to the team for that!
And also thanks for contributing, every contribution counts and helps :-) Maybe see you at a future hackathon :) https://about.gitlab.com/community/contribute/
--fixup if you have a modification to an earlier commits
rebase --autosquash amend the fixup(s) to the commit(s) they belong too.
(Of course you need to know what you doing. Otherwise you create conficts tedious to resolve in the rebase)
honestly the commit graph as a first-class value yields a lot of ideas in my mind, we should be investigating compute at the subtree layer more often.
Same. I use it to reorganise branches before submitting change requests or merging:
- Combing or splitting commits so that each commit contains a small unit of a logical change
- Reordering commits to related changes are near each other
- Reordering commits so changes that are depended on by later commits occur earlier.
- Rewording commit descriptions to have a lengthy explanation, if needed
- Removing small commits that fix minor typos
It's a great way to turn the commit history of a project into one of logical changes, rather than a simple temporal log of every change.
(I had a colleague who was so used to CVS and couldn't get his head around not using version control as a time tracking device.)
This seems a bit like an unhealthy obsession. I would be curious to know if the benefits of having a perfect git history outweigh the costs. A temporal log of changes is good enough for me.
Also in many cases I would think that the fix ends up being even smaller (or bigger) than a single commit (the commit isn't actually atomic).
Maybe the codebases I've worked on weren't active enough for this to be a common problem. But also, having too many cooks in the kitchen may be a part of the problem here too.
But don't underestimate the value of a good commit history, especially if you are using git blame to understand a change from several years ago. Wading through small typo fixes gets in the way.
Git encourages you to use branches, and rebasing allows you to squash simple typo fixes and minor changes, so you can review a change with a few commits that are easier to understand instead of a dozen.
Used with cherry picking, you can also move commits to other branches. I often find that bigger features changes turn into multiple branches and PRs, which is easier to review.
The cost isn't very much. Maybe a few minutes at most reorganizing a branch.
It also forces me to review the branch before submitting a PR. Sometimes I catch issues or use that as an opportunity to split into multiple PRs.
I use git blame to find who touched the code, but not necessarily why. And during reviews I spent most of the time looking at the totality of the code changes (along with tests and PR notes) rather than deducing what's going on by looking at one commit at a time.
Overall I see the history as a log, which is only an artifact and not a product on its own.
Reviews are very rarely happening on the commit level, so having lots of small "meaningless" commits (lint fixes, fix spec, add thing) hasn't been an issue but having squashed merge commits (effectively a rebase that squashes everything) has made release management so much easier for our team.
It also has the benefit of just less cognitive overhead of micro-managing my feature branch commits that are WIP or currently in-review.
My rule of thumb, always rebase unless you hit a series of merge conflict while rebasing. Then either merge, or squash your branch then re-run the rebase.
I'm the same. I commit frequently. A lot of those end up being `--amend`, but a lot of time I will rebase and squash/reorder/etc. In fact I'd say the majority of time I push changes I've rewritten my local history at least once.
According to any git history you'd find from me on the server, you'd think I never accidentally commit syntax errors, break unit tests, and that when I do big structure changes or refactors I get it the way I want it on my first attempt.
> including configuring git pull to rebase rather than merge
It annoys me a bit this isn't the default -- at least for teams that work from a single centralized git repo (eg: almost all corporate dev, and the majority of open source). I've done it for years and have never had any issue with it.
It doesn't rewrite anything but your local (non-pushed) history, which is totally fine. (Unless you are working with multiple remotes, in which case it causes chaos, which I think is why it isn't the default).
Are there other drawbacks?
Why? What's the difference? You can still diff the previous version of the PR with the current version and end up with the same thing that an add-on commit would give you, but ready to merge as-is.
I can't imagine being able to easily enforce that without asking people to edit the correct part of their commit. It's maybe more difficult with gitlab/github interfaces where changing the middle of a sequence of commits will not render very well, but in email based workflows it works fine.
On the other hand, being able to bisect a project without having to worry about whether an unrelated issue is causing you to traverse the wrong branch of the bisect is an enormous advantage compared to the minimal effort required of keeping track of a modified (rebased) commit in the middle of a set of commits under review.
https://git-scm.com/docs/git-bisect#Documentation/git-bisect...
I imagine something something githooks. However it might be enforced, it seems like a miserable way to develop.
This is if course mostly valuable if you don't squash commits on merge. Otherwise, the extra rebase work isn't that valuable.
Very useful when you've created a branch(A) based on another branch(B), which in turn was based on master, but in the meantime master had a few commits added, so while it's trivial to rebase B with master, rebasing A with an updated B won't work.
With a merge commit I am fixing the resulting work on both branches, which is easier than merging the in progress state in the current branch.
Anytime resolving conflicts becomes pretty intense, where you're feeling like there's a significant risk of introducing a bug in the resolution process, I think it's better to use merge commits because then it's preserved in the merge commit how the merge conflict was resolved. If things were done incorrectly, you can trace it back to the merge commit and analyze it and you don't potentially lose important intermediate state that you might need to recover.
On the other hand, If you're touching the same code in 3 separate commits, then perhaps they shouldn't be separate commits. In these instances, I often do 2 rebases: 1 to squash the commits without reparenting them, and then another to reparent the squashed commit onto the upstream branch.
Also, if you use merge commits for merging PRs, it's easy to cut through the noise of the commit graph by using git log's --first-parent option. Once you are using this, it's a lot less important to keep an ultra-tidy commit graph.
If I see a ton of commits like "fixing formatting" or "oops", I want those squashed away into proper commits.
If the changes you make are atomic in their own right and I can check out that commit, compile it, then run tests, and it passes. That's perfect. It works good for git-bisect, but the trick is to get everyone on board in the project to do that.
For public libraries I maintain, it's squash merges all the way. I like a clean history and the ability to check out that commit, compile, test, and run cleanly is perfect.
Every team I'm in I strongly advise against such commits honestly. They don't help anyone, I emphasize being able to find things you changed if you need to find them. You can even enforce a format for commits.
Funnily enough, I saw a wave of "fixes bug" "now really fixes bug" that got auto rejected by some commitcop utility at a former job, the guy was freaking out cause it wouldnt take all his changes. I guess they wanted to force him to stop doing such awful commit messages.
Commit messages should be useful and historically descriptive.
Same with me, but people still ignore it.
As opposed to a series of small commits, that would make git bisect actually useful.
Before this there was some work to bring the codebase as close to 3.x paradigms as possible.
As long as the two points before and after the work are copacetic we were satisfied. In that sense it was discreet.
From the Git CLI, without any reference to Git* platforms, it is not so obvious when searching for a commit that introduced a bug, e.g. using "git bisect" for binary search. Reading a 10,000 lines git diff can be harder than a smaller commit that also explains the reasoning in the commit message. Speaking from own experience and programming mistakes in a small team, focussing on clean commits and a good history tremendously helped in stressful debug situations. Until you hit a compiler regression bug, but that's a different story then ;)
I'm personally still very fast on the Git CLI, but I also know that there are a variety of CLI and UI tools out there that can help with analysing large Git commits. Potentially in the future also AI assisted that tell us which change a diff caused a performance regression in a release 5 months later. Or we don't need it at all because Observability driven development enabled to see these problems before merging and code reviews, e.g. the memory leak but only when DNS fails. True story from ~2016, more in my KubeCon EU talk at https://www.youtube.com/watch?v=BkREMg8adaI and project at https://gitlab.com/everyonecancontribute/observability/cpp-d...
This is impossible to enforce or guarantee at scale. Squashing PRs, though, is practically fool-proof: PRs already represent a single, atomic unit of work that passes all CI checks and is safe to merge. No such thing is true (or should be!) of individual commits within that PR. Whether we like it or not, a branch commit really only represents a "save point" for a developer.
I'm not sure 100% of the commits compile & pass all tests - there may be some mistakes - but generally we're in a pretty good state, and the clean git log is being successfully used for bisecting.
If you want even larger scale - if I understand correctly, the Linux kernel practices a similar thing, which is where we got this practice from (ScyllaDB founders came from kernel development). And since Git was originally created to help developing Linux - that's where you want to look for good practices.
Sure, you can git bisect to the exact commit that introduced a bug, but that commit was part of a larger PR, and you probably can't revert just that commit alone. So what was gained?
git revert MERGECOMMITHASH -m 1 # revert to first parent of merge (revert changes of second parent/other branch)
You can include good descriptions in your PR merge commits including discussion comments and everything. Some PR tools automate that, some do not. I don't know any technical reason why any PR tool automation would create better squash merge commit messages than normal merge commit messages.(ETA: Also you may want to reconsider your workflow if you rely on reverts that often. I had to look up the command argument, even though I knew it existed, because I haven't needed to do it in a while and I'm very thankful for that.)
One thing squash and merge does is make reverting trivial.
A PR doesn't necessarily have to be a single commit to be a good PR. Sometimes, it makes sense for a PR to have several atomic commits.
Those tiny messy commits should be fixed-up into the atomic commit that they are fixing.
If you are doing anything that involves rewriting the history you are doing it wrong.
It's not about "pretty"; commits are a form of _communication_. Do we send emails without editing before hitting send? It's a means to optimize for easier reviews through better comprehension of the changes, which also leads to faster reviews. Our colleagues don't want to read a bunch of intermediate commits.
> If you are doing anything that involves rewriting the history you are doing it wrong.
Care to elaborate? What's your general strategy?
Writing clear, atomic commits is a good idea regardless of whether you use rebase or not.
In your email analogy, rebasing would be altering previous messages in the chain. That doesn't make rebasing look good!
> > If you are doing anything that involves rewriting the history you are doing it wrong.
> Care to elaborate? What's your general strategy?
Not the parent, but to me the most important part of VCS history is accuracy. Rebasing commits (or cherry-picking them) changes their context; that context can be important for understanding why things were done in a certain way. For example, imagine we're digging through the following history, to understand how some feature 'bar' works:
* Add workaround for Error(foo) in feature bar
|
* Implement feature bar
|
* Bump dependency baz to eliminate Error(foo)
Why was a workaround for Error(foo) added, if that error had already been eliminated? Did that dependency change not work? Is there some more permanent way to eliminate Error(foo)? Is the workaround still needed?Compare that to the following, more accurate history:
* Merge
|\
| * Add workaround for Error(foo) in feature bar
| |
| * Implement feature bar
* | Bump dependency baz to eliminate Error(foo)
| /
|/
Here it's much clearer what's going on: the dependency change was not in place when that workaround was added. Hence the workaround shouldn't be needed anymore. Rebasing the 'feature bar' changes on to the 'dependency baz' changes throws away that information.[pull]
rebase = true
edited for line spacing
like thisCreated a MR for the Developer Evangelism Hacker News handbook to add this formatting tip, and some more https://gitlab.com/gitlab-com/www-gitlab-com/-/merge_request...
When I pull and it runs into a merge conflict, I want to see that error/log on the CLI. Reason: Sometimes the automated pull-rebase takes a very long time to resolve conflicts after each step. I prefer to first run
git fetch
git diff branchname origin/branchname
and then decide my strategy :)Bit off-topic but since we share .gitconfig tips - I upgraded to a recent Git version and enabled the "git push" option to setup tracking automatically. No more "git push -u origin branchname" actions. https://gitlab.com/dnsmichi/dotfiles/-/blob/main/.gitconfig#...
[push]
autoSetupRemote = trueMy practice is to rebase all my pending PRs each morning to make sure that the prior day's activity is coalesced.
If you wait weeks and weeks to do the first rebase for a big change set, you can wind up visiting the same files and conflicts way more times than is logical. In my experience, this results in a much greater chance of screwing something up along the way, further reinforcing for some developers that the rebase is bad.
Timely rebase is the answer.
You know what works almost every single time? With tiny conflicts if any? Merge + squash on gitlab.
git rebase --abort
to return to where you were beforehand.I find that git is actually very good at explaining what is going on and what you should do next by just running git status in the middle of some process like rebasing and reading what it says.
Even if it all goes really wrong the git reflog has always saved me.
Rebasing is just a way to make extra work for yourself
2) Before you rebase take a note of your current branch's HEAD commit.
3) If things go horribly wrong, you can force your branch to point at the original commit and it will be as if nothing happened. Well as long as you do it immediately. At some point git gc will collect up unreachable commits and delete them.
But to answer your actual question, there are various ways in which what others on a team do affects you:
- CI building 'merge branch master' all the time on master, because of merge-pulling upstream changes after committing
- spaghetti merge feature branches (that you might be reviewing, working jointly on, or as below)
- git log, blame, etc.
LWN has a great article:
https://lwn.net/Articles/328436/
Linus Torvalds explained it best:
Beats the hell out of:
* deadbeef (master) Merge branch 'master' of upstream
|\
| * feedbeef A proper commit message from upstream
* | deafbeef A proper commit message of the actual change
|/
which not only hides the real commit message in CI, but also flips around the history so master's snaking around between upstream's path and the person not pull.rebase-ing's path.1. 1 PR, 1 commit.
2. 1 PR cannot have more than 50 lines of product code added. Any number of lines can be removed. You can have up to 100 lines of test code.
3. Every PR should include a set of tests for the added/changed functionality. They must pass.
4. Git merge is forbidden. Everyone must rebase.
5. Every PR goes into master. Nobody can push to master. You can create as many feature branches you like, but definition of done is that your code is available on master.
6. Identify relevant existing test cases and make sure they are passing.
7. Master must always be in a state that it can be deployed instantly.
Hard disagree.
Landing a single class/module to implement new functionality may take more than 50 line of code and 100 lines of test, even though conceptually the additional functionality is simple and it amounts to mostly boilerplate and looks like every other class that implements the pattern.
Some large plumbing refactorings or CI refactorings are difficult to do in small chunks. I'm perfectly fine with a PR that replaces one internal API with another one and then has hundreds of small copypasta fixes littered around the codebase to swap out the APIs. Anyone should be able to read that without cognitive overload.
Tests are also commonly over 100 lines of code since they tend to be repetitive by nature, and I'd like to not see those artificially broken up into smaller PRs since the most important question I've usually got is if all the cases are covered or not. Breaking 300 lines of tests up into 3 different PRs to satisfy some kind of line-length metric means now the most important question I have needs looking at all 3 simultaneously and that is deeply counterproductive and totally useless. And small code changes with lots of comprehensive tests which fill out cross products of different API usages are great and should be encouraged.
And while I don't personally like the dozens-of-tiny-atomic-commits approach, I deeply don't care if other people do it or not. I've never found it useful to read their PRs, or even after the fact, but I won't stand in the way of them doing that if it is what gives them enjoyment (OTOH, I never do that).
If you find yourself unable to subdivide the task into small units, maybe your system isn't well architected. In that case this won't work for you.
You can use `git rebase` to perform the so-called "squash" action, that squashes multiple commits at once.
For example, the latest commits in the branch look like this:
123 WIP1
456 WIP2
789 WIP3
012 FeatureA
WIP1,2,3 as commits should be merged into a single Git commit, and be put on top of the FeatureA commit - so to speak, use the FeatureA commit as a new base.
Git uses the rebase command to do exactly that - you'll perform an interactive rebase onto FeatureA commit. The interactive rebase allows you to define specific actions.
"pick" keeps the commit. In order to perform a squash of the WIP1,2,3 commits, you'll keep the first WIP1 commit, and tell Git to continuously squash the next commits WIP2 and WIP3. Each operation is done at once.
pick 123 WIP1
squash 456 WIP2
squash 789 WIP3
pick 012 FeatureA
This results in a new history and commit - squashing changes the commit sha checksum, with new content in the Git commit, as well as a new base commit.
567 Squashed
012 FeatureA
"squash" in an interactive rebase keeps the individual commit messages for each commit, and at the end, the editor for commits allows you to edit the final messages. This can be helpful to review the squash action, and for example amend or abort it.
If you plan to not keep the individual commit messages, "fixup" throws them away and can be used as alternative action.
"git rebase -i" has more options - you can also stop at a specific commit, and amend it, e.g. when a "git add" was missing a file earlier.
Note: Any change to a commit within the Git history forces all later commits to change too, as the linked commit base changes too, thus regenerating the sha checksum. This entirely rewrites Git history trees, and can be very invasive - if you intend to keep specific commit IDs (for release tags for example), ensure that the workflows to not allow rebasing on certain branches. One possible workflow is to keep the main default branch protected, disallowing to rebase and push a changed history, add git tags there, and only rebase in a Merge Requests branch prior to review/approve/merge cycles.
Moving from Git commands to GitLab - GitLab also offers a Merge Request option to squash commits automatically when the MR is accepted, so that all commits in the MR branch are squashed, and you don't need to do it manually. https://docs.gitlab.com/ee/user/project/merge_requests/squas...
Last but not least - the quick action /rebase in a MR comment allows to trigger a rebase via the UI too, thus not requiring client side clone/fetch/rebase. More tips in https://about.gitlab.com/blog/2021/02/18/improve-your-gitlab...
I like using explain git with D3 to show how things work - https://onlywei.github.io/explain-git-with-d3/#rebase
Compare that with merge - https://onlywei.github.io/explain-git-with-d3/#merge
At the end, the HEAD has the same content, but the structure of the graph is different.
Note that rebase falls in the set of "rewriting history" operations in git and so pay attention to the caveat in the explanation:
> For this reason, you never want to rebase commits that have already been shared with the team you are working with.
Rebase and reset are two of the commands that need to be done with caution and full awareness if working with commits that have been shared with other people.
There are probably tools built around merging too.
> You've never lived.
Some people like having the context from the complete the history of little commits without the risks of rebase breaking something.
Some people like using rebase to group all changes needed for a feature into isolated(ish) commits.
Some people just like the aesthetics of a straight master branch without the clutter of little "corrected typo" or "fixed bug for real this time" commits.
git revert $(git rev-list COMMIT43^..COMMIT123 -- path/to/thing)
I've been on teams where everyone _hates_ what a stickler I am about good VCS hygiene until they realize something that looks like it's going to be a big pain in the ass at first glance is doable with a one liner.And then the Overlords shut down all the VCSs because "this isn't a software company". Then everyone sneaks around using weird homebrew portable tracking widgets, or, more often, just gives up.
Is rebase handy? Oh yeah.
git rebase --exec 'make test' main
The --exec <command> flag allows you to run any shell command after each rebased commit, stopping if the shell command fails (which is signaled by a non zero exit code).Just discovered the --rebase-merges option, worth exploring if you want to edit a commit under a merge commit, but don't want to mess up the merge commits.
Currently I create a text file, add this, then move this up to where I need to insert it and edit this commit.
Not exactly sure of your order of events wrt inserting a commit and editing it (I.e. edit commit message? Or changes?) and modifying it (I.e. see previous) but am happy to try and help if you want to flesh out your example a little more.
E.g. Commit 1 Commit 2 Commit 3
Using an empty fixup commit, e.g.
git commit --allow-empty --fixup=commit2hash
You would end up with
Commit 1 Commit 2 Commit 3 fixup! Commit 2
Which will adjust automatically (when starting interactive rebase with autosquash) to
Commit 1 Commit 2 fixup! Commit 2 Commit 3
You could then simply change the rebase 'fixup' instruction to 'reword' to edit/remove the 'fixup!' part of the message, or 'edit' to continue the rebase uo until that commit and then stop.
That said, again, not clear on your use case but maybe that might help you or another reader.