Git branches are named sequences of commits
blog.plover.com
blog.plover.com
Git branches, the thing the "git branch" give you are pointers to a commit, or refs if you wish. It has several implications: you can move them at will, and after a merge, you lose track of what "sequence of commits" it originally was.
I came to git from mercurial, and while they are fundamentally very similar, branches are a major difference because in mercurial, branches are really a named sequence of commits. Each commit of the branch is tagged with the branch name, and it will stay that way, even after a merge. You can't change or delete branches, only append to them or close them.
Git-style branches also exist in mercurial and they are called bookmarks, so maybe you can call them that.
Now there is the higher concept of branching that includes cloning and is not specific to git, and we can indeed call a branch a sequence of commits, but "git branches" have a precise meaning, and it is not that. Anyways, git is terrible at naming things, especially when it comes to its command line, one of the biggest flaws of an otherwise great version control system.
Git could have called their branches bookmarks too.
If Bob insists on calling vodka "Bob water", and someone asks for some water to put out a fire and Bob gives them vodka, there is a problem even if "Bob water" has a precise meaning.
Yes, you can be "technically correct" saying branches are just refs, but it's not a useful statement for most users.
I believe the author makes a very valid point, and we could do with a bit less "technically correct" and more with language targeting the usage rather than technical implementation.
Git is confusing enough for many people, and we don't have to make it more confusing for them.
The problem is that this mental model isn't that useful in the first place, and often leads to confusion. Instead, it's usually easier to think of each commit as a complete snapshot of the entire codebase, that includes a link to the previous complete commit that it was made from (which in turn contains a link to the previous commit, and so on). In this scenario, a branch is just a pointer to a given commit - that's pretty much the easiest way to think about it - and the commit itself is a stack of history. Internally, git optimises for compression by removing redundant information between different snapshots.
In this mental model, thinking about branches as just a different type of tag is easier than thinking about it as a stack of commits, because each commit is already a stack of commits. Moreover, I think the snapshot model often ends up being clearer and easier to use overall. All you need is the basic concepts of snapshots, a linked list, and pointers, and the whole thing kind of just falls into place.
Since you can go back and forth between them at will, it seems odd to claim that one perspective is inherently superior. Like insisting that a chess board is actually a white board with 32 black squares on it.
I don't see claims of inherent superiority or correctness. It's about what's useful for education.
Nobody is arguing that you need to adopt a new mental model if what you have works for you.
You're correct that the two models are equivalent, but version control is about operations that you perform on the models, and those operations will not be the same for both models. You can reason about your git history as if its a series of patches, but git itself doesn't know how to deal with any model other than snapshots.
If you start by creating a mental model, the confusion goes away. Reminding people that "branch is just a ref" is just a way to push them towards less confusion.
git log --oneline --graph
And git lola for a gestalt of the repo's recent state: git log --oneline --graph --allExample Fossil graph: https://chiselapp.com/user/rkeene/repository/kitcreator/time...
You don't have to learn how git-the-tool represents things internally. However, you absolutely should learn how git-the-model-of-a-repository represents things, because that's what you're operating on. Git is a tool to manipulate repositories, just like LibreOffice is a tool to manipulate documents. You don't need to learn how ODF stores things in zipped XMLs (just like you don't have to learn how git stores things in its content-addressable filesystem), but you need to understand what paragraphs, words, pages or slides are as this is the model you're working on (just like you need to understand what commits, branches and refs are and how they form a graph).
Unlike LibreOffice, git doesn't make it easy to understand its model just by using it (you could even say that it actively misguides you, although it has good reasons to do so), so you usually have to read some docs to grasp it.
A commit is an immutable object. Whereas a ref is a pointer, literally a place on disk (a regular file) that holds an address (plaintext SHA) to the latest point in a logical chain of commits.
Meta remark:
This is also what makes it ok to delete a branch after it is "done", and why it is ok to merge a standard working branch (like "develop") repeatedly into a target branch (like release/master/main/trunk).
The semantic / meaning of a branch is transient. It is mutable conceptually and literally.
edit: formatting
Did you confuse "refs" (references) with "revs" (revisions)?
> A commit is an immutable object. Whereas a ref is a pointer, literally a place on disk (a regular file) that holds an address (plaintext SHA) to the latest point in a logical chain of commits.
Reflog shows you the immutable commit SHA and the HEAD@{N} ref. I've only ever used it to get back to a commit I've lost, never by ref, so to me it's a commitlog.
HEAD@{<N>} is not a ref - it's a rev in <ref>@{<N>} form that means "N positions back in ref's history" (see `man gitrevisions` for more rev forms).
> never by ref
When you look at reflog's output, you've already dereferenced these commits by the given ref and its history.
Try `git reflog <branchname>`.
Using terms with the wrong definition, and not precisely defining concepts, makes things more confusing, not less.
A head ref is a ref that names a branch. But branches can exist in git without refs. Branches are artifacts that exist in the commit DAG - they are dangling chains of commits that end without being merged in to some other commit. They exist, as pure platonic branches, even if they are un-referenced.
But then you can make a head ref and name one and now all of a sudden you have a named branch. As you make more commits that extend the branch while ‘attached’ to that head, the head ref follows the tip of the branch (that is in particular a thing a head ref does that a tag ref does not).
But you can add commits and extend a branch in a detached state of you like - no head refs following the branch tip. Yet the branch definitely exists. And then if you tag it, you name it.
So no, I don’t think “a branch is a ref” tells the whole story.
But without this understanding, being told "branches are named sequences of commits" is probably worse than being told "a branch is nothing but a ref". The second one is cryptic and will soon be forgotten, no harm done. The first one leads you into a false sense of understanding, and soon you'll see an operation that looks like deep magic: someone moves the branch ref and now suddenly the whole branch is a completely different sequence of commits.
The confusion experienced by many people is largely due to the fact that a lot of articles try to teach git in a way that gives a false sense of understanding without explaining how git really works, which is exactly what teaching "Branches are named sequences of commits" does.
I still rather think of "git branches" as the technically correct "just refs" and hold the separate-but-related human-only concept of "development branch" in my head. I don't think there's a better approximation of truth than that, nor do I think it's that more complex to understand.
But the fact these two concepts share the same name truly is confusing. One would do better to refer to the former as "bookmark" or "branch tip".
This establishes enough information for you to show how using a git repo actually works; at any given time, you're looking at one specific commit, either directly or via a branch's name. If you're using a branch, then committing will perform the "increment" discussed earlier, with the branch now pointing to that new commit. Showing how to create a new branch will naturally lead to the discussion about how you can have two branches pointing to the same commit; this lets you explain that adding a new commit without specifying a branch name can be ambiguous, which you can demonstrate by checking out the current commit directly rather than by a branch name. Once you've shown that adding a commit requires either checking out one of the branches you have that point to that commit or creating a new one, you can show that the same principle holds for any other commit in the repo as well, even ones further back in the history with no branch currently pointing to them. You can use this opportunity to introduce the concept of `HEAD` as the unique name for whichever commit you're currently looking at, and that looking at a commit directly rather than via a branch is called having a "detached `HEAD`", which means that you won't be able to make any changes without creating a branch at that point first and "reattaching `HEAD`" to that new branch.
If you're trying to teach git to someone who hasn't yet learned the equivalent of an intro to data structures class in computer science, it might be worth simplifying the concept of branches in the way you describe. If you're teaching someone who already understands what a tree is, you're doing them a disservice by trying to hide the model from them because they have more than enough to understand what a branch actually is.
> Creating a branch is the same as creating a tag
> Tags merely exist to pinpoint a specific repository revision
Working with Git for version control is as if your photo management tool required you to learn about inodes and b-tree superblocks in order to save a JPEG file. I just want to keep source code history, and allow multiple people to collaborate on the same project. I don't want to know anything about "refs" or whatever else is happening behind the scenes, yet it appears Git can't be used unless you are (at least occasionally) willing to look at the plumbing layer.
However, this model was not mapped well to the high level concepts that the typical user of a VCS operates in. This is the biggest issue of git: it's hard to make sense of it by its UI if you do not understand how it works under the hood. I struggled until I read the pro git book.
I wouldn't go as far as to compare this to knowing about filesystem data structures for saving a jpeg file. It's more like using an old school file dialog where you just see the bare file system and you need to know your way around your drive.
Why do you create a branch via the "git checkout" command? Why do you delete tags using "git tag -d" but delete stashes using "git stash drop"? If you want to blow away local uncommitted changes, you can use "git reset", "git reset --hard" or "git checkout (file)" - which (I think) all do totally different things.
Git's data model may be elegant, but its hard to appreciate it through the tangled mess of git command line options.
> Why do you create a branch via the "git checkout" command?
That's a shortcut for "git branch (name)", then "git checkout (name)". Or the newer "switch" which is more obvious.
> Why do you delete tags using "git tag -d" but delete stashes using "git stash drop"?
Because the stash is more like a stack, and tags are not, so "drop" without a parameter is a valid and very usual command. Yes, it feels inconsistent, but allowing "git stash -d" without a parameter would probably not be better.
> If you want to blow away local uncommitted changes, you can use "git reset", "git reset --hard" or "git checkout (file)" - which (I think) all do totally different things.
These do all do different things, so that's why they all exist. "git restore (file)" was introduced a few years ago (with "switch", mentioned earlier) to make the last one more obvious, since that's indeed always been an uncomfortable syntax for a core operation.
Git's a very powerful tool, originally aimed at a very complex code base run by experts, and was written very quickly as an emergency replacement for BitWarden. This rushed development and target audience does show through even today, but it's being annealed over time. Nevertheless, it's so good at what it does that it's taken out nearly every other VCS by just existing (ok, and the network effects of GitHub, but they choose it for a reason too).
Bitwarden is the password manager.
Now there is also `git switch --create` / `git switch -c` for this.
Perhaps in time, there will appear different front end dialects for git. Like the statistics programming language R has the base R language, data.table dialect, and the tidyverse dialect.
I am slowly remapping my keystroke muscle memory away from the footgun that is `git checkout` and using restore/switch. But boy is it hard to undo a decade of practice.
It feels like terrible ad hoc user design built over an otherwise extremely elegant and clean data model.
Sigh.
- a “draft commit”, where you can amend a bunch of changes into a commit before finalizing it.
- default to “git commit --all”, and allow the user to “git commit --patch” where needed
- as 'anthomtb says, `git stash push --patch` (I also find myself wishing for `git stash pop --patch` so you can shove bits in and out of a stash as needed.)
a) terminology: it’s one less concept to have to wrap your head around; and
b) a draft commit would also have a draft commit message, and (though I’m admittedly not sure about how well this part would work), draft parents (probably supporting refs rather than just commits as parents) so you can have multiple of them and shuffle them around conveniently. (This also sort of subsumes the stash as well.)
I made a preliminary stab at this a while back, though it has some awkwardness and I haven’t had a chance to revisit it recently: https://github.com/wolfgang42/git-draft/
Which is what I do with mercurial (& evolve), and I am happy I don't have a super-special extra concept to clutter up my already overflowing brain.
The way I think of it, the staging area is an incrementally buildable commit that is not called a commit because commits aren't incrementally buildable. So if you allow commits to be incrementally buildable, then you don't need the staging area. The only difference is you need to come up with a message for the commit when you first start to build it. Or not—make it empty, then amend it when it becomes something worth naming.
git checkout -b foo
is just a shortcut for git branch foo
git checkout foo
> Why do you delete tags using "git tag -d" but delete stashes using "git stash drop"?That is inconsistent. One has an interface of `git <thing> <options-to-manage-thing>` and the other `git <thing> <subcommand-actions-for-thing>`. I imagine what happened is the former was the original and was probably thought to be sufficient, but then it wasn't for `stash` and the latter was introduced for more flexibility. The inconsistency is probably from backwards compatibility.
It might be worth noting that, at least as far as I know, git was like the first to use or at least popularize subcommands. It'd be understandable if they didn't include support for sub-sub-commands from the get-go.
> If you want to blow away local uncommitted changes, you can use "git reset", "git reset --hard" or "git checkout (file)" - which (I think) all do totally different things.
git-reset is mainly about resetting the branch, index, and/or working tree to a given ref. git-checkout is mainly about checking out a ref, setting HEAD and syncing the working tree to it. They're different things with an overlap. I would say that's not really inconsistent. It would only appear so when one only learns specific patterns of commands for subsets of their function, like "blow away local uncommitted changes", which in this case fits in their overlap.
Another annoying inconsistency: git tag prints a list of tags. Git branch prints a list of branches. Git commit prints ... modified files? And git stash modifies the stash. You need git stash list to see the stash. What!?
I get it; its a complex tool. Its managing 4 different storage areas for your code (the repository, the staging area, the index and the stash). It also manages tags, branches, remotes and configuration. And it has multiple networking interfaces.
But I can't escape the conclusion that its just not a very good user interface. A good interface wouldn't be so hard to use. Redis is more complex, but I don't make so many mistakes using the redis cli. Awk is more powerful - but its much more intuitive. And cargo probably has more subcommands than git does, but I don't get lost in them. Git? Git is a mess.
In this domain it was popularized by CVS which merged the separate programs used for RCS.
https://www.gnu.org/software/trans-coord/manual/cvs/html_nod...
So it is possible to have both an elegant implementation, and a friendly UI that doesn't force the user to understand the internals to work with the tool.
Git's elegant model is not why it won out. Despite of its shortcomings, I suspect the cult of personality around Linus had a big role in that, as well as major services like GitHub.
(If you always want to commit everything you’ve changed, you can do that too— always commit with ‘git commit -a’ and only use ‘git add’ when dealing with new files that you want to add to version control.)
You can also have your git porcelain handle it. Magit for example has a great interactive overview of unstaged and staged changes. When I need to do something more picky than just commiting every change, I'll usually grab magit to stash individual chunks: I don't necessarily want to commit all changes in a file, sometimes I want individual lines.
You can do that with staging using the commands above, magit, or some other porcelain (I've heard good things about git kraken). If you really want to forget staging even exists, you could just commit straight up and amend the commit afterwards to get a comparable experience I guess. I've found staging to be helpful in keeping track of what I've achieved for my next "version" of the software to be added to the history, which is why I'm still using it.
> By cherry-picking changes from workdir into a commit don’t you basically make a blind guess?
No, you use the interactive mode (`git add -p`) to select exactly what you want.
If you overpicked, you can reset a single file, and try again. That can be a bit annoying if there are a lot of changes, so this is another reason to keep commits small and atomic.
You can imagine an inverted perspective where the stash should be the only non-staging area, and the working copy _is_ the staging area for the next commit. Stash away partial changes you want to defer, then test the current working copy, then commit the working copy.
You'd also want status/diff commands that let you more easily compare: working vs HEAD (what can be committed); stash vs HEAD (all uncommitted changes); and stash vs working (deferred changes).
Not sure what you mean by "unstash", since "git unstash" is not a command (on my machine anyway, so not unless it was added very recently). I'm pretty sure stashes are still modeled as commits/snapshots.
The git stash command is a little wonky, yes, but I don't think that's a data model thing. It's easy to mistake the disaster zone of Git's CLI for problems with the data model. It becomes more obvious where the problem is when you start thinking in terms of the data model, and trying to figure out what incantation will perform the relatively simple operation in your mind.
I meant pop or apply.
> The git stash command is a little wonky, yes, but I don't think that's a data model thing. It's easy to mistake the disaster zone of Git's CLI for problems with the data model. It becomes more obvious where the problem is when you start thinking in terms of the data model, and trying to figure out what incantation will perform the relatively simple operation in your mind.
I disagree. I think the staging area and its behaviour are inherently unreasonable; certainly all the "it's just a DAG of commits" people tend to be confidently wrong about what the staging area will do under a given sequence of operations.
fossil init repo; make a new fossil repository
fossil open repo; open a fossil repository somewhere
fossil add file; note: very different than git add. in fact I was very confused by git add, in fossil the repository knows what files are managed by it and there is no staging area. so you only use "fossil add" when adding a new file. If you move a file use "fossil mv" to let fossil know what you did. Along with "rm" when you remove a file. The staging area still feels like an unnecessary added bit of friction.
fossil commit
There are others I use often "merge", "sync", "revert" but they tend do what the command appears to do. Speaking of revert, the git equivalent is really strange , even among the rest of the strange git ui. Shit! I messed something up and want to revert back to the last committed change, easy, just run the command "git reset --hard HEAD"
A bigger question is what to do when you only want specific chunks from the diff in your commit. I tend to faff around with stash, interactive sdiff and hand editing the patch when I need to dissect a chunk, a situation I feel could be better.
Magit is probably the best chrome for Git I personally think.
git add file
git commit -am "my message"
Pretty simple really. You don't need to think about staging if you don't want to.
git init .
vi file (initial content)
git add file
git commit -am "first commit"
vi file (make changes)
git commit -am "second commit"
Basically if you use commit -am, you never have to worry about staging - which is most of the time imo. In the rare case where you want to avoid committing something that has changed use git add to stage individual files.
https://typicalprogrammer.com/linus-torvalds-goes-off-on-lin...
But when you add requirements like merging other people’s work with yours, movable “tags” to mark named versions, and the sort of cut/splice/move around operations you will always need because you accidentally did something you shouldn’t have… I think you end up rebuilding most of Git’s plumbing.
I'm a user of Git. Why do I have to learn about implementation complexities in order to fix problems that arise from normal version control operations?
I can't think of any other software that forces me to do this. Even compilers (at least those for mainstream programming languages) don't require me to understand their AST representation or other details of how they work internally.
I'm a user of mercurial. I've learned enough git to know it's an inferior tool. git's CLI complexity (and mental model) is patent overkill for 99.9% of all users.
Mercurial gets out of my way.
I'm able to clone and push git repos with it just fine (thanks hg-git!)
This is because both (git,hg) are tools that manipulate the same simple data structure: the DAG.
hg CLI's verbs match my mental model from decades of use of other VCS's. I'm able to perform my daily tasks with simplicity (including n-way merge/cherry-pick tasks that git "expert" colleagues often struggle with).
My 2¢.
I've used local (SCCS, RCS), client-server (CVS, subversion, perforce) and distributed (bazaar, git, hg)
It was built on git, but hid the complexity of branching + other git things behind “tasks”. You’d start a task, and silently push that branch out to everyone. It’d silently merge things in the background and handle a lot of the chores you have to do with git rebasing and such to keep branches mergeable.
It failed miserably with us because it perpetually created impossible to solve git issues. Someone accidentally removes a gitignore file and commits a config file with a password? Your SOL. It will keep coming back because it will exist on at least one other person’s device which will get force pushed back to the repo.
The weird plumbing exists because version control is hard, and prone to humans throwing wrenches into its nominally perfect system.
The real reason you don’t remember every magic git incantation is because you normally only need a specific one, once a year. But it has to be there!
[0] https://web.archive.org/web/20190226201600/https://conveyor....
If you are encoding video, you often need to know about chroma subsampling, and colorspaces, and fractional framerate and all the other absurd technical details. It is actually much worse than git.
You can avoid those technical details if you use high-level software, only stay on happy path, and avoid any complex operations. You'll take longer and produce worse quality output than if you had fully mastered the software -- but often this is OK.
This is true both for video encoding and for git.
How does this story relate to git? Nvidia could do this research itself and hide the complexity behind a simple switch. If a user turns on gsync, then make in-game vsync a noop or advise them to turn it off, do the shit that in-driver “vsync on” does, frame limit itself to MRR-3 and empty standby lists periodically while the game is running. Pretty sure git could do a similar thing for its users.
And yes, most “consumer grade” players experience their adaptive sync technology to maybe about 30%. It’s still an improvement compared to vsync, ofc.
I agree that Git could be a bit more internally consistent, and have a few more convenience shortcuts for very basic usage. But I can also see a strong argument that, as part of the Linux ecosystem, that's a perfectly good opportunity for someone who wants to build a wrapper around Git (and I believe many have).
It seems to me that unless you really only want to use Git in the most basic way possible (add all your changes every commit, never roll anything back, single branch, no stashes), understanding something of Git's internal model is less like understanding how the compiler's internals work, and more like understanding the fact that a C program needs to be compiled into object code and potentially linked into an executable before it can be run.
In the end data model of git is its killer feature because it gives me tools to deal with hard problems and do it efficiently.
What you’re talking about sounds like the work of a git visualization tool made for fixing problems. I do love a repo viz where I can see the tree of branches for helping me understand where something went wrong. The surface level stuff people use day to day is as simple as it needs to be.
To my understanding, in a scenario like
A--B--C--D <-master
\
E--F--G <- branch1
\
H--I <-- branch2
You can't identify the set of commits I and H (specifically, commits made under a given branch) without knowing (external to the information stored in commits normally) that branch2 branched off of branch1.My logic is that I’ve used tools that viz what you’re talking about. They actually have much prettier versions of the (admittedly very nice, thank you) drawing you made. Bitbucket has one. These viz would be impossible to make without the info.
I read the git manual a couple of years ago. (All except the plumbing chapters.) There are commands that unveil deeper and deeper into whatever you’re querying, and it was in there somewhere. Give it a look. It’s a very nice manual. I read the whole thing (minus the plumbing) in less than 4 hours.
If I had to guess, I would say git reflog is involved.
I used git actively without issues for years before I took the time to learn how it works under the hood. While understanding the details was fun and helpful, I can't say it really changed much about how I use the tool.
unrelated: I wish people would stop insisting on using full-width columns to display their blog content because it is nearly impossible to read on a 27" monitor.
Wrap that blog content with a
<div style='max-width: 65ch'>There's a noticeable portion of tech people who seem to believe that unused screen space is somehow wasted. No margins + no line breaks is the gold standard to them.
I don't get it, it's unreadable to me.
I can barely read the page it's so harsh.
source. Guy who has maintained/maintains many websites large and small.
So you can be practical, and apply the very simple workaround, or you can continue to tell the void how wrong those people are.
Personally I think that your expectations put unreasonable burdens on websites which may not be run by large businesses with big budgets. Niche forums and blogs run by regular people as a labor of love shouldn't be expected to have to devote a lot of resources to be accepted on the web.
https://github.com/dbohdan/classless-css
And before you say I should do that myself, again, if you want your work to be comfortable to read for the world, the bare minimum involves legibility.
The main complaint I see here is that it doesn't have rebase, and my point is the goal of the code is for the product to work, not to have a beautiful commit graph.
I've used CVS, SVN, Git. But nothing comes close to the usability that fossil provides.
You cannot work your jpeg photo collection but opening, editing, saving, editing, saving, editing, saving and then blame your tools when your jpegs are blocky and ugly.
First you have to know some basics about lossy compressing and destructive editing. Then, you can understand what steps you really want to take.
It's the same with version control systems. With git, first you have to understand what a commit and a branch is. Then, you can work.
You need to understand Git's data model, not its plumbing. No VCS will be useful for anything that isn't trivial if you don't understand its data model.
I agree that we end up needing to know about refs because of git's user interface, but I consider that a flaw or limitation.
I will make no major defense of git, but I am intrigued at the level of difficulty reported about it, compared to the obviously higher level of use that it gets.
Videos, OTOH, are stored rather arcanely, at least on my Sony, with separate directory structure for each format.
I think most software has caught up. But it did catch me off guard that my machine needed updating to support this. I also think flatpak had some trouble with it. Can't remember details.
Both of these have their place, I feel.
When thinking about commits to a branch, humans (I speculate) tend to imagine the branch as a sequence of commits.
But when operating on the branch ref itself, like when you delete it, you should be very clear that you're not deleting commits.
What's the point in arguing about it? It's a bit like arguing if `Line = { a: Point, b: Point }` is a line or two points.
It's a DAG with optionally labelled vertices (content is cyclic of course, commit history is acyclic).
It's all potato potato.
Isn't that true of a git tag too?
A tag is another kind of ref that.. doesn’t have those semantics.
When you fetch a remote and it disagrees about a head ref, you need to do some sort of merge.
When you fetch a remote and it disagrees about a tag ref? The remote wins.
Because one of the things git is trying to do is help you manage source code. Which means that it has to help you manage branching sequences of changes in a commit DAG.
Git doesn’t just have an arbitrary set of ref semantics chosen at random - it has head refs which behave in a very particular way to help with branching, and it has tag refs which behave in a way that is useful for versioning.
So that’s why I think ‘a branch is just a ref’ is a reductive take.
Git has a thing called refs that it uses for various purposes. Git provides tools for naming and working with named branches based on creating head refs. As far as many parts of git (that deal with arbitrary refs, whether they be heads or tags or remotes or whatever) go, sure: they can work with (the heads of) branches because branches are named using head refs.
But as far as humans wanting to do things with branches are concerned, ‘branches are just refs’ isn’t helpful they don’t want to do a thing to a ref, they want to do a thing to a branch of the commit DAG, so they need to know ‘how do I get git to do this thing with the commits that are part of this branch?’, and saying ‘a branch is just a ref’ doesn’t answer that question.
You're welcome to create your own configuration of refs and define how remotes handle refs. You don't need to use branches or tags, you can treat every ref as a tag if you like.
Branches and tags really are just refs. Any ref can be a branch or a tag. The refs don't behave in any way. They're just refs. The semantics are not frozen.
Take a look at git notes, or gerrit review refs, or GitHub pull request refs, or... any number of tools which build on the ref system.
git is surprisingly flexible, but ships with sane defaults. They're not gospel.
Yes, you can do anything you can do to a branch tag to any other tag (or directly to any commit hash).
Which means you can do things like ‘merge’ any of those things. Not just refs. So things that are not refs can act like branches.
And you can’t just treat any ref as a branch. Is a remote a branch? No - so things that are refs can also not act like branches.
X is Y can’t be true if there are instances of X that are not Y and instances of Y that are not X.
I actually find git way easier to understand if I don't think about branches in the way the author suggests. My mental model for git really is commits and refs, and it helps me use it fluently. When I wave my hands around and say "a branch is nothing but a ref" what I actually mean is "...and so you should understand how it actually works, so what I'm about to do doesn't look like magic".
I think most people come to git with their own mental model of what a branch is, what a merge is, etc.
Learning git is often mostly undoing their preconceptions (ie by saying 'a branch is just a ref').
But ultimately, we humans tend to think of a branch as a branch, not a ref. For instance, the previous sentence probably made perfect sense to you.
Probably people confuse about this has experiments with some other version control system. So they already has some specific meaning attached to the name "branch" ?
While that's true from a mental perspective, the on-disk format of most standard commits is indeed a diff.
eg what changed from the previous commit, as output by the diff command
[1] https://git-scm.com/book/en/v2/Git-Internals-Packfiles
[2] https://github.com/git/git/blob/master/Documentation/technic...
Maybe the only time / scenario I'd agree with you though, is when I'm just creating a commit to capture a temporary development state (eg WIP) on the way to some development objective. That's not very often though.
However, you can fix things up when ready for review by just rebasing things into a coherent sequence of commits. I really wish there was a good autocommit/push feature in Git that would help back things up but continuously rebase and compact old autocommits.
Personally I try to avoid merges at all costs. Instead, I always try to keep feature branches cleanly rebased off master. This sucks with github, so I have a tendency to destroy and recreate feature branches to avoid getting merge commits mixed up in the remote. GitHub is dumb like that. I don't really know how to fix that except to suggest that some branches on remote should be auto-rebased if it can be done cleanly. But it's still a pain.
But in terms of what a branch is, if you don't want to call it a "ref", then I'd just say it's a commit symlink?
What I meant is something else though. A linked list is a linked sequence of nodes (or a chain of nodes, if you wish to call it that). A branch is a linked DAG of nodes, if you wish to call it that. Calling a branch a sequence of commits feels like calling your extended family (or I guess your genealogy) "a sequence of people" - it misses some crucial aspects of the structure and just sounds very confused.
In each of these cases, there's a simpler underlying data structure that's directly overlaid with a set of conventions. It's the conventions on the use of the simpler structure that gives the illusion of some higher order data type.
(The distinction I'm making between these and abstraction in general is that here the abstraction is almost intentionally somewhat leaky - presenting both strengths and weaknesses.)
To the contrary - all of these can have "branch" changed to "ref" and they still make perfect sense. "ref" is just as much a sequence of commits as "branch" is. Every commit is.
It gets interesting when you compare two bramble bushes with the same provenance, see how they differ, and then how they can be reconciled.
A commit, specifically its ref, is already encoding the 'sequence of commits' that led to it - viz. it knows its parent(s).
A branch is a named pointer to one of those then, and 'nothing more', but that is already a sequence a commits.
For pretty much every other developer who's just using git for source control, "a branch is a named sequence of commits" is a much more useful way to understand them. I don't want to think about the implementation details of how git internally represents a branch any more than I want to think about whatever's going on inside a word .doc when I add a table into my document. The important thing is my branch (or table) works the way I and everybody else expects it to work.
(It's good to have at least one person on your team who understands both those concepts deeply, as the obligatory xkcd explained...)
To be clear I didn't say a branch isn't a named sequence of commits! I said it's a named ref to what is already a sequence of commits.
I think people should understand their tools - not the internals/implementation detail (necessarily), just how to use them effectively.
Today I had some post-incident reverting to do, and it was complicated by hairy merges, poor commit messages ('what' not why/the context), poor commit stucture & history, etc. And I'm sick of CI pipelines called 'Merge branch master of <remote url>' on the master branch (from doing merge-pulls of origin/master having committed to master locally - `git config pull.rebase true` if you're going to do that) - tells you absolutely nothing about the change that's actually building (because it's parent 1, not the commit itself or the merged parent 2) and causes a snaking history on master (flipping parents 1 & 2 every time someone does it).
It's like a builder/'handiman' using a combi-drill day in day out in drill mode, using it for masonry & screws too, but not caring to realise the hammer & driver modes exist.
Also it "follows along" when you commit to the branch. Named pointers that stay put are tags.
I suppose I can recover by saying 'committing creates a new commit, with HEAD as parent, and updates the checked out pointer to point at it'. Where no branch, a 'detached HEAD', is perhaps more thr special case - it's a sort of nameless transparent pointer that updates but you're only aware of the commit ref itself. Although again tag is also a special case in that it doesn't update, as you say. Really they're just all different ways of referring to a commit, doesn't really make sense to call any one of them the true way and the others special cases.
I stand by the commit being the 'sequence of commits' though. So tldr, branches, tags, detached heads are just different mechanisms for referring to such a sequence: named and updates, named and static, unnamed/transparent. But in each case, they point to a commit, and inherently the sequence behind it.
They just want you (and others) to learn. This is an important piece of puzzle that many people miss.
Would you sacrifice all those (other people's) learning opportunities just so that you don't get annoyed?
This mental model fits those models in the code. It also fits Git-Flow. Those are all authoritarian.
So is everything Linus produces.
But, the secret genius of Linus is that he creates things with branchier futures than most creations and then he lets them wander around, becoming even more aggressively future-branching. No, actually . . . the secret is what allows him to do that. It is at least two derivatives of "creates branchy things . . . go!"
Thus git. Thus Linux.
Building named sequences of source changes via commits which chain together, with the only thing that defines the identity of the chain being a ref to the end is such an elegant abstraction, which is not bad to expose; It allows for easy reasoning about source changes. Saying git branches are a seperate sequence of commits and keeping it at that would be not a good abstraction, even if you hid the implementation really well.
Ergo, a git branch is a named sequence of commits [when the implied merge base is obvious].
A couple of questions which might help clear things up:
Do you consider `main` a branch?
If you create a new branch called `feature1`, do you consider the commits from before the branch from `main` to be "part of" that branch or not?
What if you delete `main`? Are the commits from before the branch part of the `feature1` branch then? What if there are multiple surviving feature branches that share varying amounts of common history?
Unless given more information, I would consider the commits between `git merge-base main feature1` (exclusive) and feature1 (inclusive) as part of the feature1 branch.
Now, if I `git checkout -b feature1-A feature1`, what commits are part of branch feature1-A? It depends. With respect to which merge base?
Is author saying the same thing or is he saying that each commit also has some hard-coded reference to the branch in it too?
Are you teaching someone about development? Or teaching about git? Because a git branch IS a ref. And a git branch is useful because it points at a commit in a series of commits. Branches and commits are slightly different concepts depending on your VCS and it can be worth understanding the details.
If that struck as odd, consider reading this, which I wrote sometime ago when a number of pennies dropped: https://peter-whittaker.com/
(And now don't tell me that also tags could change, and branches not.. then we are almost back to it is all refs - btw why so mad about that? It is another valid viewpoint imo).
So, not all branches are sequences.
When I understood the reflog and how nothing I do is really gone (just gotta find it!), that was when I realized how much I like git.
The reflog won't help you recover things in the staging area that where accidentally `reset --hard` though … (you can still get the ones you added to the index with `git add` but not committed[1], but changes that weren't added are lost for good)
I love git, but the UX is still terrible …
[1]: https://stackoverflow.com/questions/7374069/undo-git-reset-h...
To someone who has a right mental model of git, all of these should be just as obvious. If you didn't stage your things, they were never in your repo.
Those aren't useful analogies as git cannot remove them in the first place. It's normal for a user to expect a tool to have an “undo” mechanism for its commands (with a prompt “this action is irreversible, do you want to proceed” for the rare actions where the action have to be destructive, like when running the git garbage collector manually)
I know exactly why git behaves like it does, but that doesn't make its footgun less of a nuisance. And it's all about the UX, there's zero technical reason that would prevent git from saving your work as a temporary commit before deleting it, in a way that would make it recoverable, just a lack of user empathy.
It's absolutely not the same thing. If you just want to discard the modifications you have, then stashes are fine. But if you want to move a branch to another location, then you have to use reset --hard, and then when you have an unsuspected `git reset --hard` in your shell history, the shell auto-completion can screw you pretty quick.
That's the difference between a prompt and a cli option, the first one doesn't appear in your shell history, so you're never going to have it pre-filled by mistake.
> The problem is people teaching git,
When you have a recurring problem with people teaching “something”, then the said “something” has a bad UX.
It looks like you've described a tool with 1) a terrible UI, and 2) trains users to use commands that will eventually cause them to lose unsaved work.
2. once you know it enough, git is a powerful tool that I really appreciate.
1. and 2. aren't contradictory. And I would love git even more if the UX wasn't the dumpster fire it is, but I happen to know enough of its bad UX to be able to do what I want with it.
Also, unlike the average git expert on HN, I still recognize that the UX is shit, and that you should need to spent as much time as I did in order to be able to use it at all. I'm really annoyed when people argue that a bad UX is in fact good because of some elitist reasoning.
I don't think I've ever seen someone say "a branch is a ref" with the intent to diminish the concept of a branch or the utility of it, or the special language around it's concept.
On the other hand, appreciating what git commits are, what tags are, what branches are, and understanding the "just refs" part of it, was vital to my understanding of git. This article feels like it argues against something that (1) doesn't exist, and (2) in the form in which it does exist, serves a vital role of exposing 'just enough' of git to help some folks understand the model enough to be effective with it.