Also the fact that we didn't rebase and used merges everywhere was a major contributor to no one ever breaking their git repos, something that git seems notorious for elsewhere.
Also the fact that we didn't rebase and used merges everywhere was a major contributor to no one ever breaking their git repos, something that git seems notorious for elsewhere.
The "history should be exactly what you did" argument - which many people make - is really funny to me because a pull/merge-only strategy only preserves the _wrong_ history. As a tech lead, for example, I absolutely do not care one bit about the date of a commit, or when the developer started working on it, or what was the commit they started working on top of. That may be "what really happened", but it's worth nothing in the grand scheme of things. When a commit _has made it into the product_ is the only " what really happened" there is, and that is what I care about. And a linear history makes this much easier to analyze and understand, reducing cognitive load considerably.
Also, it's strange that you see merges as a contributor to keeping repos from breaking, as my experience is the opposite.
I advocate for a rebase-based strategy wherever I go as it helps developers push better code, it actually curbs hysteria-driven "merge it as fast as possible no matter how shit it is" cases, and I see how it turns the Git log into an actually useful source of information for developers and other personas. People start reading the logs!
The log should track the product's evolution, not the developers' activities.
I personally prefer a more rebase-heavy approach, but what we had worked very well for us.
Git is a development tool, not a product release tool. If you want to see the product evolution you could filter to just merge commits, or just merge commits in a specific format.
If you want to keep track of releases specifically, then use tags, that's what they're for. I suppose you could make a separate branch/repo where every commit = a release, but that opens you up to merge conflicts without any benefit over tags
When I say "the evolution of the product" I really mean "the "evolution of the code". When a small feature branch with 5 commits - four of which say "wip" and the last one says "added color support" - gets merged as is, and all relevant information is held hostage by whatever Git platform the company is using this week and not inside the repository itself, the log is not useful to me regardless of any strategy.
But in a different setting I would not necessarily insist in the same way.
Yes, it can be annoying if your developers are committing nonsense, but then just tell them to not do that, or to rebase locally before pushing.
If you find yourself troubleshooting a bunch of nonsense commits, you can just do a diff to the merge commit, and it will show you all the changes. But you also have the option of figuring out exactly which commit caused the problem, and seeing it in context. If I see an error in the middle of a bunch of commits that look like "trying x with y." Then I know that this is a tricky problem, and the developer was lucky to get it to work at all. If it is in the middle of a standard looking commit, then the developer didn't struggle with this. So maybe they didn't put enough effort into it, or maybe it is a rare corner case.
When I'm troubleshooting other peoples problems, every bit of information helps. Especially when the developer who introduced the problems is no longer with the company. Squashing commits removes some of that information, without providing anything that I can't approximate by using merge commits in logging/diffs.
I can definitely see how those intermediate commits can provide more information, but there's a tradeoff. More often than not, they do not provide me much value, and instead give me bloat, so I prefer to keep things simple.
Telling developers not to do something is like telling a kid not to push that red button. The average developer chooses what's easiest _right now_ and thinks it's someone else's job to fix the mess at PR time. And they're afraid, because that one time five years ago they ran a rebase without knowing what it does, and lost some code without knowing it's actually right there in the reflog, and since then they are deathly afraid of Git. I know how to use log and diff and all the others quite well, but most don't. So I'm trying to make things easier on everyone in the long term, not the short term.
"well, it landed on this rebased commit that's huge. I guess it was a kind of useful, just not as useful as we'd like".
It's like the first time I saw the essay calling ORM's the "vietnam of the software industry". I remember reading it and wondering who the hell would use ORM's in that manner?
Apparently a lot of people, but if you're using rebase because you don't know how to create commits that build and are functional then I submit the issue is with you.
As an independent contractor working with many different companies, unfortunately, virtually everybody.
git tag -f pre-rebaseJust an aside, but if there problem being solved is harder than it looks then this deserves a comment explaining why in either the commit message or in the code itself.
git add .; and git commit --amend --no-edit; git push origin --force
Rarely, I have to do pipeline work on repositories. You'd normally see twenty "Fix Jankins Issue" commits on the main branch because of some nonsense that only happens when you deploy UAT or whatever. Once I learned this little gem, this is also how I manage my feature branches mostly. But also my employer's fleet of laptops has been aging and I've had to do 3 swap outs this year, so I like to keep my in progress work pushed up just in case. git commit --fixup HEAD
and you can tack an -e on the end to add more notes in the commit message body. You can fixup prior commits by supplying their short hash, which is how I originally discovered this: I essentially wanted to amend a commit farther back in my branch.This makes a commit with a “fixup!” prefix that works with
git rebase --interactive --autosquash
you can also form that kind of commit directly, and there is also a “squash!” directive. Now you don’t have to force push amended commits. And it helps sometimes when you accidentally amend something you didn’t mean to–now you can just soft reset to HEAD~1 and try again.I don’t even bother locally rebasing to autosquash it all anymore since we use squash-to-merge/rebase in github PRs now.
function git-commit-fixup() {
git commit --fixup ":/$*"
}
alias gcf="git-commit-fixup"
# Looks for the most recent commit that matches its arg
# eg: we have three commits with messages:
# "fix: the thing"
# "feat: 5 percent cooler"
# "test: test coolness"
# then we do some work and git add, then do:
gcf cooler
# now we have 4 commits: "
# "fix: the thing"
# "feat: 5 percent cooler"
# "test: test coolness"
# "fixup! feat: 5 percent cooler"
# And autosquash will combine the fixup commit with the appropriate semantic commit as you say.
Sadly I haven't been using it much as someone introduced a bunch of commit lint git hooks that choke horribly on the "fixup!" part. And you can't pass --no-verify while rebasing.We have had multiple security incidents because some developer left a credential file inside the local git clone (no, not all tooling supports out-of-tree stored credentials). Blind 'git add .' is the first thing I teach my developers not to do.
My typical workflow is to git add . to see the mess I’ve made then decide how to clean it up. If I’ve mistakenly added a credentials file, the fix is to add it to the gitignore and unstage it, not JUST unstage it.
Not saying that you shouldn’t do both, but maintaining a gitignore and completely removing the potential problem for other people seems better than pretending your tool is more limited than it is.
Default layout is a bit weird IMO, here's what I'm doing instead: https://u.ale.sh/my-git-cola-screenshot.png
Sounds like you need better tooling, tbh.
git add .; git commit --amend --no-edit;
with the `-a` option git -a --amend --no-edit;I'm pretty sure rebasing locally is exactly what the person you're arguing with is arguing for. The original comment in this thread was saying you should never rebase, always just merge.
Merge+squash eliminates this and it works every time. One or two clicks on gitlab for example.
"then use tags, that's what they're for"
No they're not - that's definitely a useful way of using them, but they are just labels.
Why is it useful to know this? Well, when you know your tool better (how it operates, not the porcelain or CLI), you have better insights and are able to use it better.
You can manage Agile-style "features" with nothing but hashes and tags, no branches necessary. Branches are actually somewhat antithetical to distributed development, they're a useful concept, but that's all they are.
I do tend to agree though, re-base is superior to merge, in a product setting. If you want to track a feature set, having the set of commits which represents that feature set is better than a litany of nonsense tangled up in twelve feature "roots" (branches which have been merged together).
In a "do what I want" setting, rebase and merge are about equal, though. If I want to work on 3 features independently, I'd like to be able to easily see both features in parallel. I also would like to squash my features to single commits, and rebase them into feature branches where I can then merge/rebase/whatever those into my final "product" branch.
It's funny that HN can't agree on what Git is.
But, depending on your git workflow, when merging to main, I prefer a merge commit so I can see the tree of activities that lead to any particular release.
Also when you have to cherry-pick fixes into an older release branch, you get to fully appreciate squashed merges. If you did not squash, you need to cherry-pick all commits from the merge. If the feature branch was not rebased and the dev merged main into the feature branch multiple times, then the branch commits are inter-mingled with main commits and it is so easy to mess up the cherry-picking. Just imagine if the dev had to fix conflicts...
All that makes squash commit a time-saver.
Only reverting the entire feature is easier. Reverting a single change (for example, because it introduced a bug but is not critical to the feature itself) becomes much harder after squash.
No, it's just as easy, while needlessly throwing away other useful information.
> If you did not squash, you need to cherry-pick all commits from the merge.
...or use a single `git rebase` invocation.
> If the feature branch was not rebased
Linear workflows with merge commits usually require feature branches to be rebased on merge (just like squashes, just without the actual squashing), so that's not a problem at all.
One day he’s blaming Mark for a regression in the code. Being pretty loud about it in fact. When I look at the bug, it sure looks like the sort of mistake Mark would make, and the annotation says Mark. Only the thing is that I’m the one who reviewed this code and I know Mark so I was looking for exactly this sort of bug and was pleasantly surprised to find that he was learning and had dodged that pitfall. So I go excavating the history and sure enough, that bug wasn’t in the code I reviewed. It was in Steve’s merge resolution. Fuckin’ Steves, man. And the fact that git lets you do things like that is not a great feature either.
Of course Steve can just commit terrible code to main-line, and Mark (or yourself) are always stuck fixing their code, but that's what review/testing are supposed to be for - maybe blame the reviewers and testers in that case instead.
Yes, Steve should have kept things honest. But being the boss, you need to be careful not to pick favorites and to treat everybody in a similar fashion. People are very touchy about how they are treated within the "tribe".
Also, if you identified the issues both of them are bad at, are you all solving them? Being "terrible at merges"... how does that even work? Isn't that kind of an important skill? As their boss, their know-how is your responsability too, are you solving it?
Sorry if I have misjudged the situation, I obviously don't know it first-hand, so you will need to see for yourself if the above is true. There were just too many red flags (for me) in your comment to let it pass... And the reason I see them is that I have misjudged colleagues in the past, and wish I had known better then. Ah well.
This, so much this! And the price you pay for it is a slightly more difficult "insert". We started to enforce linear history in one of our bigger repositories (about 100 devs) about two years back; the first months were quite the ride (I had to do plenty of support-sessions to recover 'lost' changes). But the devs really started to see the benefits, and once they got the hang of it which was actually faster than I anticipated for most, it was smooth sailing. Many actually started to embrace it and advocate it for other repositories as well.
For me, it became also evident that filtering for people capable of learning git (rebase, cherry-pick, reset etc.) was very good at finding out who I'd want to work with and who not. It's really not that big of a deal, the UX of the CLI might be lackluster but the underlying datamodel is rather straight-forward. It's such a quintessential tool in our every-day-workflow that it's really worth putting a bit of time into understanding it, and if someone can't or doesn't want to, well, it might just be better if they work somewhere else than I do.
Developers shouldn't try to merge branches with wip/wip/wip/wip histories either, that's just garbage. Commit messages are documentation, fix your documentation before you publish.
Fully agreed, but that requires first of all to understand how this works and second requires you to run commands locally. If you're unfortunate enough to have to use e.g. bitbucket-server at work like I do, you'll always see the full graph, there which is A LOT easier to grok if it's linear. And since that's what's most devs look at (instead of git-log using some extra options) and also where CI-state happens to be reported (green/red build), that's worth a ton :)
I've found that the history rarely doesn't matter at all to me. Finding out who modified a specific code section (git blame) is usually good enough.
Generally, bigger picture stuff works best with cleaner histories as they mop up a bunch of unnecessary and distracting details, and neatly package things together. But doing so also means you're getting rid of, well, the details. If you need them later - and some poor bastard always will - you're just screwed.
Unfortunately all we've got are commits, so you're constantly fighting different groups and even different people who value the benefits of different approaches due to their positions, histories, or preferences.
This isn't even a half-baked idea at this point, but at first glance something like a meta-commit which just contains more commits and a message seems like it might be better. The top-level commits could just be the 'clean history' while deeper levels could record more of the as-happened details.
My favorite is how apparently rebasing causes developers to write better code. If you say so.
When people set ridiculous absolute rules, what develops is an underground of people who don’t follow the rules and in some cases get a thrill from subverting the dystopia.
This is a false premise. There is a product (and its various versions), and the team that develops it. Both have a claim to a meaningful definition of "history". What you are arguing is that the 'history of developer' is more fundamental than a less noisy 'history of the product development'.
Does it really matter (and need we record it for posterity) if developer x used n commits to post n incremental changes to a well defined software unit of the product?
It seems a comprise position of (1) no history rewrites before code review, followed by (2) post code review cleansing squash before merge would satisfy all concerns and history records have also served their purpose.
. A developer's timeline is relevant to her team lead, not the product manager.
Code styles are one thing, what I type into my terminal is another.
At my current employer, we used to use Perforce, and in that world you're totally right that committing (submitting in Perforce terminology, IIRC) did share changes. In that context, we developed a lot of bad patterns, losing code or holding up other people's work while a developer got their work ready to share. Transitioning to git has been super painful, due mainly to people treating git as if it's the same sort of thing as Perforce...
It sounds like the grandcomment had a ban against rewriting history across-the-board, which would help make git idiot proof. I love rewriting history, not because it's what I wished I had done but because it's what I am going to want to review when I have to.
Rewriting history is a great way for gitiots to shoot themselves in the foot.
(From the guy who force-pushed on a personal project yesterday to resolve a situation with multiple remote heads - I am ashamed)
Not to say that doesn't make for a good learning moment.
Commits in git are immutable. They're identified by their hash, so they have to be. What's more they have the hash of the previous commits so the whole chain back to the first commit can't be changed. You can only add new chains.
As a consequence, if your main branch points to commit abc123 and your feature branch points to commit def456 then it doesn't matter if you merge, cherry pick, rebase or dance the fandango, if you point those branches back to those commits, the branches must by necessity look identical to the way they looked before you did anything.
And you can find out where they used to point in the reflog.
To my original comment - having to force-push in order to resolve heads - is there a "correct" way to do this that doesn't feel gross?
Essentially, if you have more than one person making changes to the same piece of code, the method for resolving them is:
1) Pull - gets their changes
2) Merge - puts their changes together with yours and you can reorganize them at that point.
Note, this was done with Mercurial which doesn't have the concept of a stage, so Pull feels like it has a slightly different meaning when you look at it that way. One [suggestion][1] for Git if you want to achieve the same effect - getting a Fast-Forward at the end - is to Rebase, then Merge last.
Part of me knew this, I simply forgot and wanted to get out of this particular hole.
[suggestion]: https://www.atlassian.com/git/tutorials/rewriting-history/gi...
I'm not going to demand you explain more, but if you want to explain more then I'll try to answer the question you had about whether there's a better method.
Also Pull is generally a shortcut for "Fetch then Merge", and just getting changes is Fetch.
It's pretty freeing to realize that it's basically impossible to lose committed code.
You can get the same effect by reverting-to/checking-out a merge commit, or doing diffs between merge commits. You can also get a fairly clean history by only showing merge commits.
My rule of thumb is that any commit that has been pushed to a shared remote should never be re-written[]. If you're going to rebase, do it on your computer before pushing, or on your own repo before opening a merge/pull request.
[] exceptions would be removing accidentally committed secrets or large files that are no longer needed.
For the changes I do, not only I rebase them all in a single commit, I don't even merge branches. I cherry-pick my changeset. Clean commits, clean history.
Whatever I do before a push is my business and no one's else.
For anyone else changes, that's it, anyone who is not me: history is untouchable. No rebases, no squash, whatever it is already pushed, must stay as it is.
When I do a pull: git stash, git pull --rebase, git stash pop
I encourage everyone to clean their commits before a push. A clean history is a good history.
I like linear histories too, but if you have a production branch and a development branch you need merge commits between them.
But a long-lived branch needs to be rebased before review, because conflict resolutions need to be reviewed. In fact reviewing the version without the conflict resolutions is a waste of time, because that version probably won’t ever be deployed.
Do you have a better git workflow in mind than what is described above?
That gives you the history of exactly what git commits you made, not exactly what you did and how you went about solving the problem though, right? git describes changes to the source.
It seems like you're getting the worst of both worlds, trying to understand how the source changed is much harder when you have peoples' experiments and draft changes and false starts littering the history. And understanding the approach to problem solving and why decisions were made is pretty unsatisfying by digging through that stuff too really. That seems better kept as documentation and/or in the commit logs.