Things I wish Git had: Commit groups
blog.danieljanus.pl
blog.danieljanus.pl
This is a false dichotomy. I'm greedy, I want BOTH. Give me story mode when I'm just browsing the repo, but offer me the option to switch into commit-by-commit mode when I want more detail.
I don't like the name thought, commit groups seems odd, I prefer feature view or something else. The reason is commit groups sounds to me like groups of people who are allowed to commit.
I don't really follow your naming critique though. "commit groups" seems like a fine name, they are groups of commits. What you describe I would call "committer groups".
I'm asking totally unironically; every company's flow may be different, and for good reasons which I'm oblivious about. So I'd gladly read if you had time to explain.
The important thing is that your tools, specifically git, works for you, and not the other way around, so 'bad' commit messages like 'before', followed by 'after' are totally fine while working in the branch as long as it gets cleaned up during merge. However, they are the antithesis of useful months or years down the line while playing code historian. (Hilariously, the answer to the question "what idiot wrote this code", is sometimes the reader!)
So is it a lie that sometimes after lunch on a Friday, the act of programming is often a series of off-by-one, off-by-two, off-by-one-the-other-direction compile-test-edit-commit loops? And that we'd prefer to be thought of as a genius that wrote some fundamental-to-the-company code 5-10 years later, with beautiful commit messages that live up to some platonic ideal, rather than "that one dumbass"?
It's a lie the same way that people who wear makeup are 'lying'. It's true under a very specific, weird framing, but it doesn't really agree with reality.
There are notable high profile exceptions like publicly viewable patch series against the linux kernel, but you're deluded if you think those aren't edited before being released for human consumption.
It encourages people to Commit Whenever They Want To, secure in the knowledge that people will never see my -- I mean their -- feeble intermediate attempts at working code. Committing frequently is good. It means that reflog will often have interesting stuff in it, bisecting your feature branch might have a decent chance of finding obscure bugs discovered during development, etc. etc.
It's a godsend once you truly Get It that all that ugly intermediate nonsense can be removed before merging. I suspect that people who advocate against the "rebase+rewrite" philosophy do Not Get It and work in a way where even their branch-local commits are pretty meticulous and tested, etc. Nothing against those people, but they can still work fine in an environment with rebase+rewrite being the default.
I also have literally NEVER heard a good argument for keeping the full commit history (with WIP commits, etc.). There is nothing to be learned from it unless you're specifically investigating peoples' commit habits.
ETA: Some IDEs have a Local History thing where you can basically see all versions of a local file (snapshot every 60s or whatever). Do you want that in your repository history? I don't think so.
Maybe their company ranks employees by number of commits?
I do feel for these people though. Depending on their company they could try and fight it and change it (can be possible depending on things like company size, is this being introduced or well established, your influence level w/ the deciders etc) or simply and RUN and never look back.
... but, yes, if you find yourself in that type of situation and powerless to change it[0], move on.
[0] "Maybe give it a couple of tries. If that doesn't change anything, give up. No reason to be a damn fool about it." (Paraphrasing Churchill, I think? Anyway, not claiming originality.)
As they should, and tons of them certainly have gone through draft version with "garbage" commits on temporary git branches on the computers of their author. The thing is, that kind of draft version before you even want to communicate with other humans (or at least a high number of them) is rarely useful for long term understanding of the history, so it is extremely useful to have a somehow cleaned-up history in project. That you do some cleanup just before submitting a series by email is just a detail: and even then the rewriting is arguably even more pronounced than most other project because your first version of a complex change is likely to be rejected because the maintainers want some improvements on some aspects, so you submit a second (then maybe a 3rd, etc) series and most of the time you rewrite the history in each (and not just add a new patch on top of them)
So I don't think this is even an exception, on the contrary! Git history should not change on branches that are widely used, like the branch of Linus for the kernel, or maintainers, etc. For work in progress used by a single dev or highly synchronized two or three persons, maybe even skipping some processes (e.g. code reviews) while doing it at first, I better not see that in the "clean" history of the project because 99% of the time this has very little value and extremely high noise.
No, exactly the opposite: it's a lie to pretend that you can write perfect code the first time, and it does no-one any favours in the long run: not yourself, and certainly not those who come to learn from you in the future. And it's not normal or expected the way makeup can sometimes be; junior programmers will be genuinely deceived and this will cause real harm.
Keep the history, warts and all. Best case someone might learn something from it. Worst case you're no worse off.
Why would my messy history be useful for bisection? The places I committed, it may not have even fully compiled except for the very last commit. In that case, to separate the code in a useful way (such as 3 commits, one for each feature, each of which compiles on its own), you'd have to do a bit more work and create new commits, which again means either rendering the original commits pointless or disregarding them.
If you really do make most of your commits not compile then I can sort of sympathise with squash-merging, but if you merge then worst case it's a one-liner to bisect while only looking at "mainline" history (i.e. only the merges to master, the equivalent of what you'd get if you'd squash-merged), whereas if you squash-merge then there's no way to bisect back through the original history.
That's pretty interesting. I know with me, that is definitely not true, because 90% of all commits would just be the message `wip` which makes Git Blame incredibly hard to use.
Git release branch history should be kept immutable, because this is a way to see how things were in past, for troubleshooting, for ensuring that the code under source control is actually the code you deployed, and / or the code your downstream depends on, etc.
On your feature branch, you can do anything, as long as you can later cleanly merge with the main branch from which releases are cut.
I'd say it's a completely normal practice to rebase, split, and fixup your commits on the feature branch, in order to present a clean picture to the reviewers, to easily show that a new test catches the error in the old broken code, and the new fix actually makes that test pass, etc. Nobody cares about what happens on your feature branch but you. Its commit history is not holy, it's a tool like other tools. Somebody depends on oncoming progress of my feature branch while developing their own? Well, `pull --rebase` regularly, same as you do with the main branch.
Squash that history during the merge to the main branch. Do not delete the feature branch, uncheck that checkbox in Github repo settings. Clean history representing completed features: check. Detailed explanation of development in code: check.
On the topic of the article: to my mind, "commit groups" can be sufficiently well implemented as branches, or as tags and ranges of commits between tags. For some very complicated cases, they can be implemented as actual text tags inserted into commit messages.
Let's call it differently: dabble branch and share branch. It's the share branch where you interact with others. You are not bound to have exactly one release branch, and often you don't, when you backport stuff to older releases. But this is a branch you keep in order because you share it with others.
Your dabble branch is your playground. You can do weird things, make stupid mistakes, fix them, etc. You do not share that branch with others much, except to let them see its current state. They do not depend on it, and not expect it to be nice.
When your portion of work is done, and you (maybe several of you) want to share it with other collaborators, not involved in the process of your dabbling, but interested in the result of it, you may choose to clean it up. You can reorder commits into logical spans, and meld them. You can split a commit that does two unrelated things, and describe each separately. You get rid of all the noise (if you produced any), and form a nicer picture for your collaborators to review and understand. You do it because you care about their time and sanity.
Then you merge the result of your dabbling into the share branch, squashing commits into one. This keeps the history of the shared branch(es) observable. If anybody wants to step back, they have your original dabble branch, which you now abandon and create a new one.
Dabble branches should be short-lived, a couple of days. You can have many long-lived share branches for features that take long to develop, etc. Share branch history usually does not need cleaning up, so there's usually no point to rewrite it. It allows to merge it periodically with other share branches, if any.
As for caring about your reviewers' time and sanity, rewriting commits that they may already have seen is the opposite of that IMO. Any decent review tool will let you review a single combined diff for the whole branch, and that's what a reviewer who hasn't been following your progress will use. Meanwhile if a reviewer did happen look at your branch yesterday, taking away their ability to view just the changes since then is doing them no changes. (This is especially true when it comes to applying changes from review feedback - if I requested a couple of small fixes then I want to review a commit where you made those small fixes, I don't want to have to re-review the whole PR because you rebased)
I spend considerable time these days making my commits as clear and readable as possible, and between rebasing, squashing and amending messages, it's quite likely there there might be dozens of intermediate commit IDs for every single commit ID that enters the repo.
This way to see feature per feature I do git log --merges and to see the commits git log --no-merges
Your story mode == commit mode provided you can divide at a good enough granularity. This is generally not a problem though.
The best way I've seen by far: Prepare a fast-forward merge, then merge it with --no-ff. You end up with a linear history of commits grouped by the merge commits, can see either view in git log using --first-parent or not, and bisect can find the actual commit when needed.
This top-level comment shows what the commit history looks like doing what I said: https://news.ycombinator.com/item?id=27723435
I also didn't just say rebase because that could also mean squashing commits manually and isn't what I'm talking about.
It limits how commits are grouped a bit. But that limitation might be good to prevent feature creep.
I often do this, but then have some integration related change that’s required after.
> git show --first-parent COMMIT # perspective of target branch = branch you were on when merging
> git show -m COMMIT # perspective of both branches
Lengthy explanation and more ways to do it: https://stackoverflow.com/questions/40986518/git-show-of-a-m...
Also the following works in my experience and is more easy to memorize because the command literally says what you want to know:
> git diff COMMIT_BEFORE_MERGE..MERGE_COMMIT
This article shows what I mean: https://euroquis.nl/blabla/2019/08/09/git-alligator.html . (The only thing missing is the rebase before merge, since the example image shows a case where a merge could have hidden some changes.)
That said the rebase followed by --no-ff merge is my preferred approach as well.
git show [commit]^1..[commit]
The ^1 incantation means "first parent", so in the feature branch style I'm describing, this would show the changes on all of the commits for the feature merged at [commit]. You can use `git diff` instead of `show` in the same way, if you just want the cumulative diff.It seems to work on the repository I've been contributing to, as it shows each of the commits that were part of the feature branch. The merge commit itself is empty, as desired, avoiding the up-thread concern of unreviewed auto-merged code.
"git bisect" supports "--first-parent" too.
git log --grouped
and "commit-by-commit" mode is git log
Are you proposing something different?For a view of commits that shows when they entered the current branch, rather than when they were authored (i.e. group all commits from the same PR together in the history):
git log --topo-order
For only viewing merge commits (as if you were using a squash-and-rebase strategy, but without having to rewrite history at merge time): git log --merges
Assuming you are using branches as groups of commits — kind of the point of a branch — and you're using merges with --no-ff (which is the default on Github for the Merge button — https://docs.github.com/en/github/administering-a-repository... — and is necessary in this scheme to prevent fast-forward "merges" that would mess up viewing the merge history) rather than squashing and rebasing, `git log` shows you the "true" history, `git log --topo-order` shows you the chronological history of when commits entered the main branch, and `git log --merges` shows you the zoomed-out "clean" history of PR merges.Many people don't know about these features because Github doesn't have an option to view the log that way: only git does. TBH I wish Github had offered a --topo-order and --merges selector to their commit log view before offering a myriad of PR merging/rebasing options: git provides a wealth of solutions to viewing commit history without forcing you into destroying parts of it with squash-and-rebase.
Edit: IMO, part of this issue is because of git's bad default log view. The --topo-order view is a much more useful view for reading the history of the repo (both for knowledge building and for debugging) than the default chronological-by-commit-timestamp view; typically you don't care what date a commit was authored on, as much as you care when it entered the main branch. I can't think of a single time I've actually wanted the default log view, and yet it's the default, and thus it's what Github shows.
For my new project I've already decided some months ago to use merges with --no-ff when merging in work by other people, along with a kernel style "everyone gets their own repo" model.
A lot of the discussion around Git workflow is basically noise created by the weak and hardly ever improved GitHub UI. True. They prefer building AI to building better log views on their core product. But, even much more advanced Git clients like the one in IntelliJ don't properly expose all the different ways of looking at git logs, and git has so many flags and options that it's nearly impossible to know them all. Even if one day you spend the time to read the user guide from cover to cover, newer versions will add new features and you won't see them. If the GUIs were better, I suspect all these discussions would quickly dry up.
o-o-o o-o o-o-o
/ \ / \ / \
o-------o-----o-------o-->For instance one of the biggest annoyances with git is it's a pain in the ass to find the the merge of a commit into the mainline (aka the next child with more than one ancestor… probably), which can make it difficult to go back from a commit to a PR unless it was a single-commit pull request.
edit: also https://stackoverflow.com/questions/8475448/find-merge-commi..., and also github shows this (but only for github PRs, not other merges)
It encodes the necessary information in the commit graph, without introducing a completely new concept (commit groups). It’s true that Git doesn’t give you the tooling to get that information out of the box, though.
A bog-standard merge already does that.
"git rebase" doesn't create a new hash if there's no change as a result of the rebase. But GitHub's PR rebase button always creates a new hash even if there is no reason to. (Its PR merge button does not do this; it will merge a one-commit PR without creating a new commit).
To add to the inconsistencies, GitHub doesn't sign the new commit when using the rebase button so it doesn't show up with the green "verified" icon - even if there was no need for a new commit anyway. Yet when using the merge button it does the opposite - if it doesn't need a new commit, your signed PR commit is merged to the main branch and shows as verified, and if it does need a new commit GitHub signs it (if the PR commit was signed) so it still says "verified" (even though it's really GitHub's key that was used, not the author's).
For this reason, when I merge PRs I avoid the GitHub UI, and use "git rebase -S" locally followed by "git push". This does what the PR rebase button should do.
In addition, if people are feeling charitable, the branch will be cleaned up prior to merge with an interactive rebase to squash out "Wip" commits and hopefully leave a nice clear set of self-contained commits that provide a logically separated view of the work that went into the change.
If you're not doing that, might it not be better to squash? The point of not squashing is to preserve valuable granular commits. If you have WIP commits, those don't have much value.
I suppose it depends on how clean the feature branch is. Personally, I have horrible commits, knowing I will clean them up later.
But I believe this also suffers from the problem described in OP where you can't tell if HEAD^ (or any parent) is from master or from feature branch.
It can be done atomically by having a tool perform the rebase and merge, and never merging anything by hand.
That’s necessary to implement the “not rocket science” rule anyway.
Summary: every feature is merged with 'main' as the first parent. Then, whenever interacting with history, you tell git you only want it to consider --first-parent
And he's right. If someone does the merge while checked out on the feature branch, then commits it _as_ the master branch, then the first-parent concept breaks.
Don’t do it for the same reason that you have a convention of useful commit messages.
Which makes it nearly useless. Not totally - the people who know about it benefit - but it's extremely far from safe. The only way to really enforce this is to add this kind of thing to your "main" remote repo that everyone pushes to.
e.g. Add new feature foo to main menu
...
The problem is getting all the visualization tools on the same page.
Bazaar got this right. Their official UI displays a linear history with the ability to expand any merge commit to show the side branch.
https://commons.wikimedia.org/wiki/File:Bazaar_Explorer_-_Lo...
But I can see why people might still desire more explicit constructs to handle history organization. The expectation that the commit graph should be both logical and historical causes a sort of cognitive dissonance, which these sorts of conventions don't fully resolve.
l = log --graph --abbrev-commit --date=relative
Im not sure about collapsing history though.They're a commit with two parents, one of which has no real meaning anymore because it doesn't exist anywhere else. They break a lot of tools because of this.
What do you mean? That's the feature branch, what's the issue?
The second parent is the last commit of a series of commits starting at a commit, sometime in the past on the main branch, followed by all the new commits that represent the feature that is being merged.
Merges are horrible, if you don't have a mental model of git that aligns exactly to the above.
Say for example, I'm looking at a freshly cloned repo. There's a first commit and most-recent commit on master - I can identify them with git-log. The problem is that I cannot view the path of commits between these two if I'm only interested in the commits made when the current branch was master (which is generally the case unless I want to drill into a feature branch).
Disallowing merges makes the problem go away but that removes a lot of options in terms of work-flow.
When I'm working none of the commits are made when the current branch was master. They are made on branches, where commits are finalised, tested and signed, and then master is fast-forwarded to match the branch tip commit (or something intermediate if I'm satisfied with that). Conflicts are detected by the fast-forward and reconciled locally in the branch, re-tested and signed off again.
It's rare that the current branch is master, and no commits are created while on master.
So what would you like to see?
Which "master"? Since git is fully distributed, there's no central repository which contains the true "master" branch. It's perfectly valid to have two independent lines of development, both naming its current branch "master" (in separate clones of the repository), and later merge one of them into the other. As another example, consider the branch named "for-linus"; take a look at the Linux kernel git history, and see how many independent branches all named "for-linus" are merged on each release.
It would need to be transferred on the side and not in the commits themselves to not rewrite commits in case of merges from other remote. Similar to tags, but instead of only pointing to single commit it could be a list, let's call them labels. Git log would then add option like --labeled-by=origin/merged-on-master
git checkout feature
git rebase main
git checkout main
git merge --no-ff featureThe best part about this workflow is that history remains linear; it is very easy to track the history of changes (since there is never a "branch" with changes on both sides) while at the same time you keep the ability to visualize where the start/end points for a given feature were.
It also works with nested branches! You simply create a new branch2 from your branch1, and then merge --no-ff branch2 to branch1.
I like that this poster is at least thinking about what he wants the history to look like, but you’ll never get it good with a broad rule. It really takes the exercise of rewriting every change you make from a stream of consciousness set of WIP commits into a logical series of incremental patches with clear commit messages. And doing that every time you want to push anything (especially when it is “just” a WIP or feature branch). After a months practice your sense of taste will kick in and tell you if a --no-ff merge seems appropriate for the current chunk of work.
Some people are really great at writing commit messages, most are not. But for even for the ones who write the best commit messages, often their is a lot of discussion inside the actual MR which never makes it into git.
GitLab and GitHub write a merge commit that can be used to get back to that discussion, but a git blame or bisect doesn't take you to the merge commit, which means you have to spend a fair amount of effort to get to the merge commit.
Treating merges as a group would be amazing.
Something but addressed in the article is how you deal with groups of groups of groups of merges. IE topic, feature, dev, master branches.
It would be somewhat difficult to make the tooling graceful at ungrouping different levels of groups.
The reason why I prefer squash commits is not only the linear history but also the fact that I’m simply not interested in the sausage making process. If a PR becomes so big that it would need multiple commits than I request smaller patches. But I also see that this heavily depends on the team size and general setup. But I use squash PR ever since it was introduced in GitHub. Before I manually rebased the feature branches to have a clean merge.
It's also equivalent to a merge commit, and then only keeping the diff from the merging in branch.
So it's sort of a half way house, and sort of the worse of the both worlds.
Sure, no one really cares about they sausage making process, but tracking down regressions is so much easier with a proper history. Bisect is your friend.
I see this a lot and in my opinion squash on merge is just a very poor version of using rebase to put your commits into a proper state before merging.
A commit should be one logical set of changes. It's great if you can get a piece of work done in one change, but often larger work requires several changes.
Beginners seem to treat pull requests as places to pile commits until you get something reasonable at the end (the sausage making process). A better way is to curate your work and clean up the individual commits as changes are requested in review. Then you have a coherent history with commit messages that might tell you something useful.
This meant that log/blame would default to showing the high-level summary of each review, but the git commit graph still has the full gory history if anyone needs to look at it. More importantly, you could safely start developing on top of someone else's unreviewed changes and git would correctly be able to track who made what changes, in a way that breaks badly with a standard squash or rebase workflow.
To view the "simplified" linear history, you run "git log --first-parent --no-merges", which unfortunately doesn't have a config option to set as default.
Unfortunately, we lost this workflow when we moved over to GitHub, and now git blame is surfacing everyone's broken wip commits because we've stopped adding the high-level summary squash commits.
I feel that it's quite over the top. Nobody goes to the commit history to "learn programming" (or at least you should not) and if you did, it would be so confusing to see things added in a way and changed back 2 commits later because it didn't work.
The code is just half the story, it needs context, it doesn't represent the train of thought behind the changes, that's only in your head. If you want to document what didn't work write a document where you explain the different approaches you took, why they didn't work and what you did at the end. That really is super useful.
Just keep the history clean, if you ever need to revert a change you will be grateful you didn't add many changes to 1 commit or didn't change the same code in 5 different commits throughout the PR. Keep the code working at each commit, if you ever need to `git bisect` you will be grateful you did.
This benefit is much more important than having a visually pleasing history. And you can always get a visually pleasing history by adding a few flags to git log.
No one's talking about ever getting rid of any commit ID that got checked into `main` or `master`.
We're talking about getting rid of commit IDs that only ever existed on one machine and only during development.
From my reflog, I see this commit ID: `21b9f2e HEAD@{28}: commit: Fix and instrument nstray URL problem`
No one needs to know about my typo - ever.
In my version, the nesting can be infinite of course (I guess the author here would call these "groups of groups" -- but that might be complicated with the flat approach of a group being a range of commits).
But basically, I want the UI of an entire "set of commits", but then a disclosure triangle to be able to see the "true history" if it's interesting to me. It is simply the case that sometimes one is more useful than the other, and other times the reverse is true. There simply isn't a one-size-fits-all solution. As far as most people are concerned, you only ever want to, for example, unroll the entire set if a test fails. Or, you want to cherry-pick the entire set to another branch. They serve as one logical unit. But if you for example care about the decision-making process that lead to that final code change, you can see it by revealing the "inner history".
the network graph shows what's going on: https://github.com/fiddlerwoaroof/git-group-demo/network
The article also seems to imply that there's one "right" way to merge branches and that teams should stick to one approach. I disagree. I use all 3 approaches whenever it makes sense: merge commits for large PRs with more than 1 commit where you want to preserve the history of the changes, squash+merge when the history is messy (usually after code review) but the change itself should be atomic, and rebase when the history is clean, though I mostly reserve rebasing for smaller single-commit PRs where a merge commit would just add clutter.
I also disagree that Git is this flawless piece of software we can't improve upon. It regularly fails to do fairly trivial merges automatically, forcing me to manually fix conflicts, or use `rerere`. Ideally my code versioning tool would understand language semantics and be as maintenance-free as possible. Git is nowhere near this and requires quite a lot of familiarity and hand holding to work as the user intended. The amount of time and effort spent understanding and using it properly is difficult to quantify, and it's still a tall hurdle for new developers.
The git cli UX is terrible, cryptic, and forces one to think way too much. The number of GUI tools that wrap the git cli should tell us all something.
I Mercurial. It - manipulates the identical data structure (the DAG) - interops just fine with git (thanks hg-git plugin!) - its cli manages to use common sense verb names, by default, for all its functions - includes no unnecessary concepts/models (read: git's index)
As someone who stubbornly uses mercurial for all his code, git's index is the single greatest feature that git has over mercurial (not counting the Magit interface). The index allows me to _incrementally_ build up a commit. I can go back and fix things, add and remove things from the index, and only when I'm ready, commit. In mercurial, I have to hold so much more state in my head about what is ready to commit, what needs to be fixed, and what is code that needs to be reverted.
And `hg commit -i` is nice, but is not a replacement for the index. Git's index allows me to add hunks, then stop, go to lunch, review the status, remove some bits, add more, then commit.
I'm told that Mercurial's queues feature would support this, but I find it incredibly non-ergonomic, and I can't quite figure out how to use it as an index replacement.
This feels extremely pedantic to me. There are certainly workflows that produce this confusion, but the most common ones definitely don't.
For the most part, in especially most github-based flows, the 'left-most' (aka `git log --first-parent`) history of the main branch is precisely the history of the main branch, and the "commit groups" are the divergent "right" parents.
Can someone do something that temporarily breaks this? Sure. People `git pull`ing with divergent changes at least used to litter projects' histories with this kind of nonsense. But it's not terribly likely to make it into your main history these days if your upstream repo has a 'protected' main branch, which is so normalized at this point it ought to be considered the default state of affairs.
It seems like maybe the thing the OP really wants is just for the branch name at commit time to be stored as metadata in the commit. That would maybe help with pulling out intentions while looking at history.
Also, Mercurial had a kind of 'hard branch' feature that also might have resembled what's desired here, but as far as I can tell most users of mercurial found it more frustrating than helpful and used plugins that provided looser kinds of branching.
Furthermore services like Github add valuable valuable comments like the PR number that was merged.
However, this approach only works if all the developers on a project buy in to this philosophy.
As much as I dislike squash and merge, it's better than the alternative of trawling through a git history with dozens of "WIP" and "Fix tests" commits and janky rebases/merges.
Individual commits on a PR/feature branch are useful to check what's changed since the last review (I silently despise developers who force-push their rebased commits when fixing things, making it basically impossible to re-review).
I just wish it was easier to detect locally whether my development branch has been squash-merged so I could script the deletion of them.
I personally wouldn't want to see 15 commits to the main branch from one PR regardless of how descriptive the commit messages may be.
That way you maintain a linear history, you get clear marks of when each pr was merged and you can see what happened in between.
The obvious downside here is that you'd need to add yet another step in your process and come up with a meaningful system for this.
If your weakest git user cannot revert easily then you're in trouble. Enjoy being on call 24/7 otherwise. Reverting is more important than committing.
I genuinely don't understand why people care about git history so much. I've never needed to look at it in 8 years of using git.
If you deal with a service that has been matiained for a few years, this is also an excellent way to figure out what was done why and when. Or figure out if a well meaning rebase accidentally clobbered a critical piece of ancient logic. Or determine who the heck owns something when you realize it's time to split up a larger service.
The list goes on. Note git history plays directly into git blame too; it can be an excellent tool in the right circumstances.
I have allocated some of the valuable left-hand-only keyboard shortcuts in my IDE to searching Git history. It tells you the "why" when the "what" doesn't make sense, and the "who" when git blame shows some reformatter.
However, it is also necessary to use git "empty" commits as the first commit of each "group" to prevent confusing output from 'git log --graph' when nesting and sub-nesting groups of commits.
Here's an example of doing this -- https://github.com/maratbn/test_nested_sub_groups_of_commits...
That way even if a commit message is not clear or you added an improvement to an earlier commit, you can reduce the clutter a lot before sending out for code review, while still having individual commits for different parts of the code change
Maybe "poorly" is unfair as this would be relatively complex to implement in a way that has good simple intuitive ux. But from the perspective of anyone using vcs daily it does seem an obvious want, so I'm surprised it hasn't been done well yet.
cus,
Each feature = one branch
Each little change = one commit
It's annoying that they throw out half of that info.
I may be wrong here, but I would argue that git does forget something: it forgets where the branch reference used to point when a new commit is made on that branch. If a commit has more than one parent, that means some information is lost because it has to guess which parent the branch reference came from.
> Merge: 8 6
> So it tells you that these two parents have been merged together, but it doesn’t tell you which one used to be main. You might guess 8, because it’s the leftmost one, but you don’t know for sure.
Why don't I know for sure it's gotta be the one on the left?
Because of fast-forward merges.
It will be the one on the left if you, as expected, were on the master branch and merged the feature branch ("git checkout master; git merge feature"). But if you were on the feature branch, merged the master branch into the feature branch (usually, this is done to resolve conflicts), and then went back to the master branch and merged the resulting feature branch into it ("git checkout feature; git merge master; git checkout master; git merge feature"), it will be a fast-forward merge: no new commit will be created, and master will point to the merge commit which was originally on the feature branch, which is going on the opposite direction as you would expect.
The solution, as others have already mentioned here, is to do all merges to master as non-fast-forward ("git merge --no-ff feature"); that gives a consistent order to all merge commits on the master branch, and the end effect is the most similar to the "grouping" feature OP wants (if everything on master are these non-fast-forward merges, the difference between one merge and the preceding one is that "group").
- It's still ok to "git-pull --ff=only", right?
- Is it possible to enforce this at the CI level?
It looks like a bug which should be fixed. Git should create a merge commit when it sees that "the direction" changes.
If you want to futz with stuff and squash commits and rewrite commit messages because you didn't write good ones on a feature branch, fine, whatever makes you happy.
But just merge to master, don't rebase.
git log --pretty --oneline --graph --topo-order
That should at least group commits roughly by branch. Did that answer what you were asking?Please, devs
Or do you mean a hash of the hashes of all your heads?