Confusing Git Terminology
jvns.ca
jvns.ca
The single thing that made everything "click" together is that most things are just pointers to commits: branch names, HEAD, tags, all of them are pointers.
HEAD is pointing to the commit you're currently looking at
The name of each branch (e.g. `my-feature` points to the latest commit of that branch)
When you're on main and you `git checkout -b my-feature` then you have at least 3 pointers to the latest commit on main: `main`, `my-feature` and `HEAD`.
Every time that you make a commit on `my-branch`, then both the `HEAD` and `my-branch` move to point to the new commit.
"detached HEAD" means that the the `HEAD` (the commit you're looking at) is not pointed at by a branch.
The difference between tags and branches is that the tags point to a specific commit and do not move.
----
The other thing caught me out multiple times is that most commands seem inconsistent because git assumes default arguments:
`git checkout file.txt` is the same as `git checkout HEAD -- file.txt`
When you're on `my-branch`, `git rebase main` is the same as `git rebase main my-branch`
The difference is that the latter you can run it from other branches too.
----
Last but not least, when everything goes wrong, the single command that can take you out of any weird situation is `git reflog` which shows you all the commits that HEAD has pointed to.
Having said all that, I'm glad that git is acknowledging all this confusion and are implementing commands with fewer surprises and simpler interface as that will make it easier for newcomers to pick it up.
For beginners it is a fun exercise to understand why empty folders can't be added to git.
(Fun fact, all such empty trees would be deduplicated to a single object).
There are also annotated tags that can contain a message, have a timestamp, and sha, etc. These are proper git objects that behave a lot like commit objects, except they're still typically only referring to another git object (commit).
But now what's a "proper git object"? Is there an improper git object? Is there a proper git non-object?
When you `git fetch`, git is asking the remote to walk a tree of objects — starting at the commit object that the ref points to — and deliver them to you, to unpack into your own object store.
Git could in theory do a lot with just objects — with the whole "data state" of the repo (config, reflog, etc) just being objects, and then one toplevel journal file to track the hash of the newest versions of these state objects. (Sort of like how many DBMSes keep much of the config inside the database.)
But git mostly isn't designed to do this. Instead, git's higher SCM layers manage their state directly, outside of the object store, as files in well-known locations under .git/. This means that this higher-level state isn't part of the object-store synchronization step, and there must instead be a domain-specific synchronization step for each kind of SCM state metadata where applicable.
Tags are an interesting exception, though, in that while the default "lightweight" tags are "high-level SCM metadata" of the kind that isn't held in the object store; "annotated" tags become objects held in the object store.
(To be honest, I'm not sure what the benefit is of having "lightweight" tags that live outside the object store. To me, it looks like tags could just always be objects, and "lightweight" vs "annotated" should just determine the required fields of the data in the object. Maybe it's a legacy thing? Maybe third-party tooling parses lightweight tags out of the .git/ directory directly, and can't "see" annotated tags?)
I think no one uses lightweight tags anymore, except if you push by mistake a commit to refs/tags/something.
Can you elaborate on why that's helpful? I rarely get into a weird state with git but when I do it's almost always faster/easier to just delete the repository, re-clone, and re-apply my changes manually.
With reflog I was able to go back before my rebase and redo it again, more carefully this time. (Unfortunately I couldn't just cherry-pick the thing I wanted)
In general it saves you from having to do what you described with delete/re-clone/re-apply things.
However, in my experience the most difficult "weird state" to get out of is when you do something that removes/rewrites history. For example: deleting branches, rebasing, squashing, or accidentally getting rid of a reference while attempting to solve some other problem. The root issue is that you want to find a commit that seems like it no longer exists. If you rebase, the branch now points to a new commit that has ansestors you don't want, but the old commit is gone. The "secret" is that all commits that ever existed still exist, you just can't find them in `git log` because the pointers to them are gone. `git reflog` helps solve this by giving you a list of all commits that HEAD has ever pointed to.
git checkout cool-branch # HEAD points to commit abc123
git rebase main # HEAD and cool-branch point to commit def456, but you realize you don't want that rebase
git reflog # reflog tells you that the commit you were just on is abc123
git reset --hard abc123 # HEAD and cool-branch now point to abc123. You've Ctrl+Z'd the rebaseLOL, that hasn’t typically been my experience, but I appreciate the positive attitude.
You can get to detached HEAD even if the commit is pointed to by a branch
Or to use the name git gives to that concept, "refs." Thus reflog :)
Also, one thing that I've found hasn't occurred to most people using git, is that all the branch/tag/etc refs of your fetched remotes, are also refs, able to be referenced anywhere you can name a ref.
For example, if you ever want to say "I don't care what's on this branch, don't fast-forward or merge or rebase, just overwrite my local branch with what's on the remote!" then that'd be:
git checkout foo
git reset --hard origin/fooAs for the refs, yes, good point ;)
A command name that I read as re-flog for the longest time :D. I really wondered about the strange, strange name for quite a while before I bothered to look up what it does and found out that I should read it as ref-log and that it is, indeed, a very useful thing.
Also being a de-refr probably helps.
The joke that isn't a joke is that you really need a CS degree to use git. It's not wrong.
The git default arguments makes for a very inconsistent experience, agreed.
And then `git checkout .` vs `git reset --hard/soft` vs `git cleanup` are all super similar but very different.
Still love it more than perforce/svn (I did like mercurial a little better back in the day but I doubt that might still be the case).
----- `git reflog` is king and shows you the truth.
Once you internalize that git unlocks.
And then accidentally wiping out all of your checkpoints and generated images :')
HEAD pointer is pointing to the branch pointer (e.g. my-branch) which is pointing to the commit. (Except in a detached HEAD state.)
> Every time that you make a commit on `my-branch`, then both the `HEAD` and `my-branch` move to point to the new commit.
HEAD pointer keeps pointing at the my-branch pointer, and only the my-branch pointer moves to point to the new commit. But of course, when you now follow HEAD to my-branch to the commit, now you end up to the new commit.
> "detached HEAD" means that the the `HEAD` (the commit you're looking at) is not pointed at by a branch.
"Detached HEAD" means that HEAD is pointing directly to a commit, instead of pointing to a branch pointer.
You can have a detached head state, where both HEAD and the branch pointer point to the latest commit. If you use `git log --decorate`, for the latest commit it will show (HEAD, my-branch) instead of the normal (HEAD -> my-branch).
Ah, thanks for pointing it out, always good to learn the finer distinctions
Yes! Here[1] is a nice picture about from Git Reset Demystified article. You can also see is in the following `git branch -a` output,
* main
remotes/origin/HEAD -> origin/main
remotes/origin/main
[1] https://git-scm.com/book/en/v2/Git-Tools-Reset-DemystifiedAnd if you're feeling extra-paranoid like me, a `git rev-parse HEAD` and copying that string down somewhere safe before embarking on a tricky process with lots of merge-conflicts or other shuffling.
It's nice to have confidence that I've accurately identified which state is the last-known-good one--not one a few steps too far into the chaos-zone--and that I can usually get back to if everything goes to hell. (Barring unwise use of stuff like `git gc` or `git filter-branch`.)
Quite a lot of git is conceptually simple, but the UI has quite a few confusing names & options.
I don't see git as a tool to save text, its for coordinating changes on the same 'text' by multiple people, and doing so quite precisely and reliably, without blocking anyone's path forward. Try that with word.
It could be better, yes. But I'm always surprised by the hate git gets. Its an amazing tool and miles ahead of the tools we created for non-developers (word, google docs, etc).
Maybe its because I'm old enough to remember when subversion was king.
Is is entirely intuitive to everyone? No.
Did developers always have trouble with version control and merging? Yes.
In my experience most devs don't grok version control, period. It does not matter if it's SVN, CVS, RCS, Visual Source Safe, Clearcase, Perforce, you name it. Or now git.
I don't think that Git sucks, it's usable. But there are better alternatives.
FWIW, yes I have used others, maybe not the ones you have. I have used some of the ones mentioned in my list of examples, though not all but also others not mentioned.
[1] https://www.fossil-scm.org/home/doc/trunk/www/fossil-v-git.w...
[2] https://www.mercurial-scm.org/
[3] https://stackoverflow.blog/2023/05/23/for-those-who-just-don...
Mercurial: I was not impressed. Yes it sort of works but if you know git, you just miss the fact that everything (tags and branches) in git is just a label. At least that's what I missed in the two years I had to use mercurial.
Fossil: Haven't used it but the main advantages touted in the doc you linked to aren't advantages to me. I.e. the only thing I'd want from it are the actual file versioning parts. And then I read things like
You can say "fossil all sync" on a laptop prior to taking it off the network hosting those repos, as before going on a trip. It doesn't matter if those repos are private and restricted to your company network or public Internet-hosted repos, you get synced up with everything you need while off-network.
Err, yeah, I `git pull` my repo(s) before going on a trip and it doesn't matter if they're private and restricted to my company network and I am synced up with everything I need while off-network. So?Pijul: Sounds interesting from a first read but I'd take some time to actually read about it and try it out for real.
Veracity: `(C) 2010-2014 SourceGear`. I don't want to be that guy, but: really?
Personally I subscribe more to the UNIX philosophy here, where I want my tool to solve one thing but be great at being combined with other tools in ways that the original authors probably never imagined.
E.g. I've used `git` in many scenarios.
I've used it with just one other developer where we'd push and pull from each other's laptops via ssh and no central server whatsoever. In that same place we used git for doing backups. We used Bugzilla for bugtracking with that.
I've used `git` w/ `git svn` for about two years in a company that used SVN and nobody knew I was even using and reaping the benefits of `git`, while they were having loads of issues with branching, tagging and conflicts. I showed some people. They wouldn't listen. IIRC we used Rational Clearquest or something like that in that place. It's been a while and it wasn't fun at all.
I've used `git` w/ a central server via SSH. No, not github w/ ssh protocol, just our own central server. Special user whose "login shell" was `git receiving data`. This is the Bugzilla place actually!
I've used `git` w/ a central server via HTTP. Funnily enough this was also the same place as with SSH but ya know, can't expose SSH externally, right? So HTTP it was via an Apache Proxy. Yes that long ago. So yeah Bugzilla place again. But also used this setup at other places that used Jira. Jira 3 mind you. Did we use GreenHopper there? I'm not sure any more. It's been a loooong time.
I've used `git` with bitbucket as the central server. Definitely with GreenHopper in Jira!
I've used `git` with `github` as the central server. Jira Cloud. GreenHopper has now for a long time been "Jira Agile" or "Jira Software" or whatever the nom du jour might be since they bought it.
Do you see the one constant in here? `git`! Because it does one thing and it does one thing well and you can combine it with all these other tools and environments and integrate with them. You don't have to convince anyone that "My tool's integrated issue tracking is the best, throw away your 20 years in tool X and migrate everyone and everything over". You "just" need to convince them, that of course `git` is better than Clearcase, which is uphill battle enough for one year.
Nevermind the wikis in use in those places. Don't even remember which where necessarily. All the way from the none, through Wikimedia, MoinMoin, Confluence and Jive (Jive leaves a particularly bad taste as it was in the Clearcase place and used as a company internal social media platform everyone was supposed to use for 'everything' :shudder:)
- I've used Fossil within a team of three and we were able to sync from each other's laptop without central server. We used it also for bug tracking, with no additional stuff to install.
- I've used in the same company with larger teams on different projects on central server over https or ssh.
- We've worked with github and other external central servers of Git repositories.
There's also one constant here: `fossil` :) It's a nice discovery for me, after Mercury (which I still use for my personal projects) and they are doing the job with ease (clean and less error prone interfaces.) I'm not selling you something and not trying to convince you they're better (I don't care as I can use them and interact with your git repos and will be transparent) ; alternatives were asked and i point some out, that's all. :)
It's a hyperbole, a figure of speech used to make a point. I'm pretty sure everyone on HN knows what git is used for.
> Maybe its because I'm old enough to remember when subversion was king.
I used SVN for more than a year. Why would you compare git with SVN rather than other DVCS like Mercurial?
I know. The hyperbole makes it sound as if gits job is actually quite simple and we could easily have a better system. But I disagree with that, I don't believe it is simple. Most other tools have and are failing still at this.
> Why would you compare it with SVN rather than other DVCS like Mercurial?
Because SVN ruled the world back then and that is what I and many other devs used. Many of us went from SVN to Git without even knowing what Mercurial was. That explains why I was (eventually) so in awe of git, had I gone from Mercurial to Git I might have lamented the loss of a more friendly system. I do actually remember doing the odd thing with mercurial and it being much smoother to work with, but then git was already becoming the dominant player.
Yeah I know, that's the problem. In other words, many of us decided git is the best because many of us didn't know any better.
But it's been more than a decade, surely that's enough time to reevaluate one's position.
And the reason that it's better is that the support is universal (now). Even back when the battle was being actively waged between git and hg, the popularity of git made it a better choice for pretty much everyone.
Heh. I'm old enough to remember when CVS was king. Subversion was a huge improvement. Git was too. I have every reason to believe further huge improvements are possible. They might have to be gargantuan though to overcome inertia by now.
Large file handling and submodule/subtree.
I wish we still had that, because git monoculture also means that anything that replaces git first has to reimplement git. This means that just like ASCII or scroll lock buttons, we're stuck with git mostly forever.
git won for good reasons, it's clearly better than what came before it. It may be popular to shit on it now (similarly for jquery), but when it arrived on the scene it was clearly an improvement.
In a parallel universe, had there been a HgHub (with same strong initial iteration as GitHub), I have no doubt that mercurial would have won.
In fact, it worked so well, I stopped using "svn merge" (which took 5+ minutes on our repository for every merge), and started using "git merge" with git-svn instead (which reduced the merge time to <3s, and even the extra git<->svn sync overhead cost only 30s or so). As a bonus, git also reported fewer "merge conflicts" (svn at the time had issues repeatedly merging from a branch to trunk).
So when I ended up picking a DVCS for another project, git was the natural choice since I already knew it. I imagine there are a lot of developers who started out on SVN and took a similar route to learning git, so having a high-quality Subversion bridge turned out to be one of the critical features on the road to adoption. This advantage in adoption then snowballed via forges like GitHub.
https://www.youtube.com/watch?v=4XpnKHJAok8
He talks about what SVN got wrong, specifically, that it made branching easy and it's the merging that's the important piece. Git won because it made merging easy.
You guys keep covering your eyes and ears, pretending that Git and SVN were the only players on the VCS market. That was not the case.
Sure, SVN had problems and people wanted something better, but Git was far from the being the best alternative for the average software development team.
> It may be popular to shit on it now
I was shitting on it 10 years ago, along with a small minority, and for good reason. Unfortunately, the hype was too strong, and we are where we are.
I think svn also originally didn't have merging, or at the very least it was such a bad experience svnmerge.py was created and really common to use. Even once it got good it was mostly the equivalent of using git cherry-pick to pull commits across branches, though it did get a special "reintegrate" mode for diffing a branch and applying that back to trunk (and I remember there being something about it being possible to accidentally undo commits if you weren't fully up to date when running it...?)
Edit: I remember now, all changes on trunk had to be merged into your branch first, or reintegrate would interpret it as if your branch undid those commits and remove the changes from trunk. It basically made trunk look exactly the same as the branch did at that moment.
Making an svn branch and merging it, now that was a huge issue. "svn merge" sucked compared to "git merge".
For the project I was managing back in 2007 I started using git-svn, because importing Subversion commits into git, using "git merge", and then exporting commits back to the subversion server, was faster and worked better than "svn merge" did (back in 2007).
Weekly, I witness co-workers confused about git, I see posts online asking for help, I see articles like the one here once again trying in vain to explain something that should be simple.
In all my time coding, I don't remember anyone wasting hours trying to undo some mess they'd made in SVN, TFVC, Perforce, etc.
Tools exist to make our lives easier. If they can't do that, they don't deserve our time.
It doesn't take a huge amount of time to really learn the basic concepts and then maybe the 10 most used commands (I suggest using git cat-file -p to navigate a repo and history manually). There is a wealth of functionality and an inconsistent terminology, but you can quickly look that up when you actually need it.
We'd kind of standardized on trunk-based development because everyone was afraid of even attempting merges.
The only real advantage git has is bandwidth and disk space —- not that it uses bandwidth and disk space more efficiently —- only that we all have more of both, so it seems more useful than subversion (or whatever else you were using before) when you had limited disk space so you didn’t pass around infinite copies of everything.
Hah, I've recently wrote a post about similar issue - why we may be locked with git.
>So what's the issue here? I'm worried that just because GitHub is so good, then unless they decouple from git as letters management engine and allow any/other, then we will be locked with git.
Leaky terminology.
Imagine if you had to know OS/IDE/compiler internals for basic usage
Back when I used to walk backwards uphill for 6 days to get to school, that's precisely how it was.
Except that "IDE" was a misspelling of "I'd" and nobody ever did that.
Further it does nothing - it tells you nothing about the remote. It could be 1 second or 100 years and you would still need to fetch to determine if anything is different. Then what is the information for?
ls -l .git/refs/remotes/origin/master
origin/master is just a file on your system, you can see when it has been changed. It doesn't magically get updated. Do `git fetch origin` and if there are any changes, you'll see the timestamp change, and the contents:
cat .git/refs/remotes/origin/master
The basics of git are so simple, you can implement the core data structures and some operations in a day. It is really worth it to get to know these.
Somehow git has managed to create a very complex user interface on top of quite a simple core.
The confusion lies in that origin refers to different things depending on if it's `origin master` or `origin/master`. Eg `git pull origin master` does the thing we expect
Git is not well designed for human beings.
Is it though?
Consider a bank telling a store I am buying things at "This person has has €450 in their bank account", when at that moment I have €310. The store would be rightfully pissed at the bank for effectively lying when it is made clear later on that the transaction could not be completed and the bank answers "well, the person had €450 a few days prior to you asking us".
Without an explicit temporal information it is explicitly now.
Without an explicit status on sync status it is implicitly saying sync is up-to-date.
It's a lie of omission.
but it's not?
if you have an old bank receipt that says you have $450 in your account, but you actually have $310, you need to "get" a new receipt that has the newest value.
you do that by issuing git fetch origin. then you can git merge origin/master to make everything up-to-date.
what you have is a "paper receipt" (your checked out version) from your bank. something that, if you need an up-to-date version (from another remote), you need to request a new one (by issuing git fetch).
git is, by default, distributed, so whenever you need to see the world outside, you need to be explicit. linus made it this way because back in the day (not sure right now tbf) tons of kernel developers do work without any internet connection, and would only connect to pull/send patches.
this talk[0] by linus from 2007 (i remember watching it on google videos lol) explains really well where the git mentality came from. i really recommend it to you, since it feels like you are not really getting how git works.
You can't go to the hotel desk clerk and ask if you have any messages. Then for the next four hours keep telling people "the front desk has no messages for me" despite you not asking them in the last 4 hours. Things could change!
But you haven't talked to your bank, used an ATM, or been on the app in weeks! Your balance could be totally different - bills have came out, you got paid, interest, etc.
You are making my point! You know to ignore it because its old, outdated information.
Then why are you telling me this? How is it useful to me?
Of course youre up to date with what you last fetched - that is _always_ the case.
Why mention being up to date even? Just tell the user when they last synced with their remote(s).
But that is not what this message is about. It's confusingly worded, as many people agree, but what it says is that your local ref "main" points to the same commit as your local ref "origin/main." It says nothing about "main" on the other computer/server.
And it is not the case (i.e. you are not up to date with origin/main), for example, when you have committed to main but haven't pushed. It is also not the case when you have fetched but not merged.
This might be where the misunderstanding is. You are not always up to date with what you last fetched. Say you have develop checked out, and you run a git pull. As part of that process, git checks the status of all upstream branches, and updates your local reference copy of them (that’s what origin/develop, origin/production, origin/feature-branch-1 are: your local reference copies of upstream). Then you check out production, which you last touched two weeks ago. Git will let you know that your two-week-old local copy is behind origin/production, which is your local reference copy of what it just saw when it fetched from upstream.
A problem it has is there is now a generation of developers who don't know why we don't use centralised VC any more.
For the vast majority of companies signing up for Github/Gitlab/whatever licenses, the remote/decentralized part of git is pointless.
So the decentralized aspects of git just add a layer of complexity/indirection for a lot of use cases. Many extra "git pull"s in my workday.
git has a better experience than cvs or svn if you're far away from the VCS server, but that was solvable by having dev machines near the VCS server. I've gotten used to the git workflow, but it still doesn't strike me as uniformly better, other than if you're using git, you don't have to deal with everybody always ask why aren't you using git.
Yes, yes it is.
origin/master is not saying that the remote has/hasn't changed. It's comparing your local copy of origin/master, not giving you the status about if remote has/hasn't changed. You need to explicitly ask if remote/origin/master has changed or not if you want to know.
Which in your analogy would be like if the store forgot to actually ask the bank if the customer had the money or not, and instead relying on whatever information they have "cached" in the store. Instead, the store has to first ask the bank (remote) if there is any changes.
I do agree that it could be worded better to actually help the user understand, as it seems to be a common misconception.
Sidenote: I'd be driven to absolute insanity if `git status` started doing remote requests to check the remote origin/master status each time I invoked it.
Git hurts my brain.
That's pretty close to how holds / unresolved transactions work.
Unless it's a packed ref (https://git-scm.com/docs/git-pack-refs), in which case it's just a line in the packed-refs file.
It's not a biggy, it's just one of the little toe-stubs and paper cuts you get over pretty quickly.
I hope you "git pull -r" every time. Merges are almost always unwanted.
(very large sigh)
Can't wait to retire in three years and wash my hands of this broken trash.
I wondered how many commands have a more machine-readable output, and that led me to the git-scm page "Git Internals - Plumbing and Porcelain". In summary, Git was originally written as a toolkit for dealing with version control, rather than a polished version control system. Many of us who used Git from the early days learned to do VCS work with these lower-level commands, and we've passed those workflows on to many other people as well. This is the "plumbing" layer. Git later developed a more polished layer, referred to as the "porcelain".
I'm not entirely clear yet on which commands are part of which layer, but this helped me make sense of the newer workflows I've seen recommended in recent years. It also gives me a better way of reasoning about possible changes to my own workflows.
https://git-scm.com/book/en/v2/Git-Internals-Plumbing-and-Po...
Looking back, I suspect that I was extremely confused when starting out, but was pretending I wasn't to try and seem cool.
However, now that bitbucket has dropped Mercurial support, I'm not entirely sure where I can easily push a mercurial repo for backup. For better or worse, I am extremely dependent on Gitlab to backup my code so I'm not risking my work on a potentially failing hard disk/ssd.
I don't know. For large binary files I still use Google Drive as backup (I know 3D models are not necessarily "large" by today's standard)
One can use git LFS, but there isn't an easy way to free up the storage them occupy from the history. And GitHub LFS is about 5 times more expensive than Google Drive per GB.
I figure that the moment Gitlab sends me a nastygram about it, I'll move to S3 or Google Storage or something.
What does Fossil buy you over Git?
https://fossil-scm.org/home/doc/trunk/www/rebaseharm.md
In terms of Fossil as a technology, it's an SCM with built-in project management tools (wiki, forums, bug tracker, etc) so it does much more than git does.
The front page has more info:
Fossil was built for cathedral style development, where you’re a small team of trusted contributors. You get offline first issue tracker, wiki, forum, chat etc. out of the box and integrated into an easy to backup solution. Allegedly it doesn’t scale as well though.
With hosted services like Github, I feel like fossil doesn’t buy much more than offline first capability and simpler to use. However if you self host, fossil is dead simple run (single binary) and to backup due to it generating a single SQLite file and it’s easy to stream changes elsewhere. Setting up git with gitolite, email list, issue tracker etc is much harder (though I’ve heard gitea is easy to use).
- The git staging area, the term literally everyone agrees with
- https://news.ycombinator.com/item?id=28143078
> “Your branch is up to date with ‘origin/main’”
> ... But it’s actually a little misleading. You might think that this means that your main branch is up to date. It doesn’t.
No need to well-actually this one. Isn’t the Two Generals problem applicable here? If you are being really pedantic, it is impossible to tell whether you are “up to date” right this split second. Even if you do a fetch before `status`. So what’s the reasonable expectation? That the ref in your object database—on your own computer—is up-to-date with some other way-over-there ref in an object database last time you checked.
A simple `status` invocation can’t (1) do a network fetch (annoying) and (2) remind you about the fundamentals of the tool that you are working with. In my opinion. But amazingly there are a ton of [votes on] comments on a StackOverflow answer[1] that suggest that those two are exactly what is needed.
> I think git could theoretically give you a more accurate message like “is up to date with the origin’s main as of your last fetch 5 days ago”
But this is more reasonable since it just reminds you how long ago it was that you fetched.
[1] https://stackoverflow.com/questions/27828404/why-does-git-st...
...
Another thing: I thought that `ORIG_HEAD` was related to `FETCH_HEAD`, i.e. something to do with “head of origin”. But no. That “pseudoref” has something to do with being a save-point before you do a more involved rewrite like a rebase. Which was implemented before we got the reflog. I guess it means “original head”?
Why is that the expectation of a “distributed VCS”?
Edit: Removed accusation of "Conflating the implementation of the system with the wording of its output".
"Up to date" means caught up in time, as opposed to space. Local branch positions are more like space ("where is this branch pointing?") and remote ref state is more like time ("when did I last update the remote ref?"). I know, that's very subjective. Either can mean either.
Anyway, I think a better wording might be "Your branch matches origin/main" or "Your branch's head is the same as origin/main" or "Your branch is pointing to the same commit as origin/main" or some other tradeoff between verbosity and clarity. Maybe with the author's suggestion as a parenthetical: "Your branch ... origin/main (remote ref last updated 5 days ago)."
That’s fine. And also less partial to the implicit view that the remote ref is the thing you are supposed to be “up to date” with.
"Your branch ... origin/main, updated 5 days ago."
i don't think that's annoying. i want network operations to be explicit, not implicit. when i do git status, i wanna know the status of my repository as is in my file system.
if i want to know what's going on in another remote, i will fetch that and then compare.
This tool gets shared a lot and it feels appropriate here. It visualizes the underlying git model and the effects of various commands.
https://learngitbranching.js.org/
I learned more from 10 minutes with this tool than the preceding ~10 years of git experience. Can't recommend it enough.
> That, however, is not the only way to use the tool
And “pull request” somehow is exclusionary? No, because you can use it to talk about both inter- and intra-repository changes.
And it also handles the cross-repo case, which is a common case in the Github model of "make your own personal fork of the upstream repo and send PRs from there," which has advantages -- it allows random people to send PRs without needing to give them permission to e.g. pollute the upstream repo's branch namespace.
Though I agree the situation is somewhat muddied by the fact that you can create pull requests for branches in the same repository (even though that's not the normal workflow). GitLab's "merge request" terminology is more accurate for that use case.
If I pull your code, your code goes to me. If I request that you pull my code, my code goes to you.
...at least, that's how I always read it.
From a git perspective, there is no meaningful difference between server and client.
but that's a much less useful way to work than to denote a repository somewhere as the "canonical" one, and then all those involved pull from it and push to it.
the presence/absence of a server associated with a repo is a bit of a red herring - what matters is the workflow associated with the different repos. my own repo is just a local private one of no particular significance, as is that of my colleague. by contrast, git.ardour.org is canonical (and, as an aside, github.com/Ardour/ardour is merely a mirror).
Github took inspiration from git's request-pull command but instead reinterpreted it as merging from one github repo to another.
The underlying action of merging a "pull request" is to pull from the submitted branch to the target branch. It's no different than Linus doing it from a maintainer branch.
In short, it's not a request to a server but to another person. You're requesting that they pull your branch to take a look at the changes you want to contribute to the project.
[0] https://git-scm.com/book/en/v2/Distributed-Git-Contributing-...
It explicitly models conflicts, so you can e.g. send your conflicts around and aren't forced to deal with them immediately.
It tracks and logs all repo state including repo operations (checkout, rebase, merge, etc) and those state changes can be manipulated and reverted individually just like commits.
The working copy is automatically committed, and operations typically amend that latest commit. This alone like halves the number of unintuitive concepts new users need to learn. Combined with robust history rewriting tools this makes for a much better workflow.
"HEAD^ and HEAD~ are the same thing (1 commit ago)"
Followed by:
"But I guess they also wanted a way to refer to “3 commits ago”, so HEAD^3 is the third parent of the current commit, and HEAD~3 is the parent’s parent’s parent."
The author's language implies a contradiction they immediately prior said doesn't exist. If these two distinct constructs were indeed different ways to define the same relationship, the second paragraph would say "^3 and ~3 are both ways of saying the third parent of the current commit, or the parent's parent's parent." Instead, they've defined the constructs as different once again.
They're not the same thing. Merge commits can have multiple parents, and HEAD^3 refers to the third parent. If HEAD is not a merge commit, then HEAD^3 doesn't refer to anything. HEAD~3, by contrast, refers to following the parent's parent's parent, independent of how many parents any of the commits in question had.
More specifically, following the first parent only. man git-rev-parse (surprisingly) has a nice explanation:
> A suffix ~<n> to a revision parameter means the commit object that is the <n>th generation ancestor of the named commit object, following only the first parents. I.e. <rev>~3 is equivalent to <rev>^^^ which is equivalent to <rev>^1^1^1.
"~3" however refres to grand-grand-parent to the current commit.
It's confusing because ^ looks like an up-arrow, so we think of it as a pointer to look n levels up. That's not the case.
The series made me also pick up vim, and I have not looked back since.
So, if we take the example from derefr [1], `git chekcout foo` lets you go to your own local branch `foo`. Then, `git reset --hard origin/foo` modifies the current local ref (`foo`) to be the same as `origin/foo`, and change the working directory accordingly.
Also git reset appears to be taking a commit as the final parameter, git reset [--soft | --mixed [-N] | --hard | --merge | --keep] [-q] [<commit>]
But why it is accepting `origin/mybranch`.
I am mostly rambling - feel free to ignore me.
That isn't a single argument, that's two arguments. If you look at docs for some commands which accept `origin mybranch`, git fetch for example, it says the first argument is `<repository>` which can either be a URL or a remote name. In your example it's a remote called `origin`. The 2nd argument is then a `<refspec>` which is a bit complex - it specifies what to fetch and where to fetch it. So in your example `mybranch` is shorthand for `mybranch:mybranch` (i.e. fetch `mybranch` from `origin` and update local `mybranch`). You can even do `git fetch origin mybranch:mybranch-local` which would fetch `mybranch` from `origin` and update `mybranch-local` with it.
It currently has 72 aliases and another 35 lines of defaults. My pride and joy is "git extract" which is a 700 character shell script that uncommit a single file from the latest commit while preserving the staged and uncommitted changes.
No one else in the world speaks my bizarre git dialog, but with my .gitconfig at my side, I feel like a git wizard. Take this file away from me, and I can barely commit a change. I have no regrets.
git checkout my_file HEAD^1
git add my_file
git commit --amend
The reason I wrote it is that I have a workflow where I build up a single commit gradually by adding changes once they are "done" and ready for the final code review. I think most people instead create a series of commits, rather than one, and then squash-merge them all at the end.
What I like about my approach is that at any time I have: committed changes that are complete, staged changes that still need a little cleanup (e.g. documentation or tests), and then unstaged changes that are what I'm working on right now. I may also edit the commit message as I go along.
It's not common, but I use the extract command when I've committed something but decide I want to revert it in part or in whole because I found a different way. Again, having all my committed changes in the top commit helps.
> $ git checkout merge-into-ours # current branch is "ours"
The current branch is now "merge-into-ours", not "ours". Except it is also "ours".
The next example adds to the confusion:
> $ git checkout theirs # current branch is "theirs"
At least, there should be a consistency and both branches should be either called "merge-into-ours" and "rebase-into-theirs", or "ours" and "theirs".
> Imagine that for some reason I just want to move commits F and G to be rebased on top of main. I think there’s probably some git workflow where this comes up a lot.
This comes up for me in a branch where I have some necessarly local changes which are permanently there and have to be maintained, and that branch experiences non-fast-forward changes from upstream.
Say we are up-to-date in this branch, plus our two local commits that are not in upstream.
We do a fetch. Upstream has rewritten 17 commits. So now we have 19 diverging local commits. We only care about two of them. We just want to accept the diverged upstream commits, and then rebase the two on top of that.
We do not want to rebase all 19 commits!
We can do:
git rebase HEAD^^ --onto origin/master
So HEAD^^ is the "upstream" here that we have explicitly specified. So following the documentation:- All changes made by commits in the current branch but that are not in <upstream> are saved to a temporary area.
[That's precisely our two local commits; those are the ones not in HEAD^^]
- The current branch is reset to <upstream>, or <newbase> if the --onto option was supplied.
[We did specify --onto, so the current branch goes weeee... to origin/master, the abruptly updated, non-fast-forward upstream that we want to catch up with.]
- The commits that were previously saved into the temporary area are then reapplied to the current branch, one by one, in order.
[That's the cherry-picking that we want: our two local commits.]
So end result is that our branch is now the same as origin/master (up-to-date) plus has the two needed local commits.
Further, I have a situation in which there are two git repos on the same machine which have a different version of a local commit. One of the repos (A) is an upstream for the other (B).
When B does a fetch, it gets the master branch from A with A's local commit, which is unwanted in B. B has its own flavor of that commit:
git rebase HEAD^ --onto origin/master^
takes care of it. We catch up with origin/master, but ignoring one commit, and on top of that we cherry pick ours.Discussion: https://news.ycombinator.com/item?id=37799082
I love Git, but with the huge caveat that it is relative love - relative to the universe of garbage software trying to perform complex tasks. I think it has a few glaring issues (the reset command and the overloaded pages of the official docs come to mind), but I also think 80% of the criticism aimed at the design or CLI of Git is undeserved or misapplied. E.g., here's my quick critique of some of the points under the "alien mental model" part:
> A commit is its entire worldline
> Commit content is both a snapshot and a patch
> Branches aren't quite branches, they're more like little bookmark go-karts
These are three versions of the same fallacy - technically correct, but only in the same sense that a text file is "both text and ones-and-zeroes", or "not text, but actually ones-and-zeroes". The author is pointing at different abstraction layers, some of which aren't even necessary parts of the user mental model. If you don't yet understand concepts like "abstraction layers" or "implementation details" (e.g. the difference between "a branch is a series of commits" and "a branch is represented by a pointer to a commit, which, by the definition of a tree, resolves deterministically to a series of commits"), then you will have this problem with any software that gives you power to work on a complex problem like version control.
> Merge conflicts are actually just difficult
If your version control system makes merge conflicts easy/rare in the absolute sense, then it is doing dangerous/cute "idiot-proofing" that will bite you at some point.
In summary, I think a good chunk of complaints about Git are actually just complaints about version control (i.e. Git is hard because version control is hard), or unnecessary combinations of different abstraction layers (which, to be fair, most software/documentation hides better from the user, but IMHO that is a bad thing). Git is an amazing piece of software (relatively speaking).
> +refs/heads/main:refs/remotes/origin/main thing.
>
> [remote "origin"]
> url = git@github.com:jvns/pandas-cookbook
> fetch = +refs/heads/main:refs/remotes/origin/main
>
> I don’t really know what this means, I’ve always just used
> whatever the default is when you do a git clone or git remote
> add, and I’ve never felt any motivation to learn about it or
> change it from the default.
This is a refspec[0], and it tells Git the relationship between local references and remote references, defined in the pattern [+]<src>:<dest>. So, in this case, it defines the local head of main to refer to the remote branch main on origin. You could, for example, define all branches on local as linked to all branches on the remote with +refs/heads/*:refs/remotes/origin/*.
An example of a real world usage of this was when we were migrating our repo between providers -- configuring the fetch field to link a different remote for different branch names kept things simple for our developers as we quietly moved branches to the new remote.
It also configures the scope of a git fetch command, iirc, so you can restrict your fetches to only scopes that you care about (maybe your team has a branch name prefix/namespace that you can specify, so when you git fetch you only get things that are relevant and not some other teams' branches you don't need).
[0]: https://git-scm.com/book/en/v2/Git-Internals-The-Refspec
One of the problems I have is figuring out how people get themselves into such situations. Like if someone submits a merge request and their branch looks like it's been merged into itself with one side rebased or something. When I ask them what they did they have invariably forgotten. If someone cooked a dish that was too salty I'd be able to tell them "try putting less salt in less time". I wish I could do the same with git.
Although upon typing this I'm thinking perhaps I could figure it out by looking at their reflog? But that involves accessing their computer or walking them through it which would probably confuse them even more.
This is wrong. 'git rebase main' doesn't affect main at all! If you are on main and check out a new branch, then add commit-one, then switch back to main and add commit-two, then switch back to the new branch and 'git rebase main', your new branch will have commit-one and commit-two, and main will still only have commit-two.
it's much more like simply 'git merge main' with some extra magic to avoid a separate merge commit.
Sorry if it’s already been mentioned but this fake git man page generator gives me a big stupid grin every time I use it: https://git-man-page-generator.lokaltog.net/ (I’m sure it’s been linked on HN before)
The documentation it generates just so plausible with the kind of verbs and nouns it uses, and how the help text is always descriptive, but never really helpful.
But that doesn’t mean it’s not confusing, both over- and under-documented and unpredictable.
To me it’s a bit like seeing someone work with highly specialized systems (think: CAD or a particle collider): who am I to say it “doesn’t work”, these people are designing cars and discovering bosons. But still I think: if someone could design these things again from scratch, they could likely be even better.
`git restore` though is slightly nicer than using checkout. I've incorporated its use into my workflow, even though I've previously used checkout for the same operation.
I'd suggest one edit in the first entry:
“heads” are “branches” --> a "head" is the end of any given "branch"
This makes it a little clearer when talking about detached head state, since switching to a commit on a branch can still leave you detached.
I was confused by that when I got to your detached head description, so I double-checked.
and throw "tip" in there, as in "HEAD is the tip of the current branch".
HEAD can point to a commit that is the end of a branch (typical state, HEAD is “attached”), but it can also point to one that isn’t (HEAD is “detached”).
This was always super confusing to me. I'm happy since using fork.dev that instead of "ours"/"theirs" it just names the branches. So much clearer. I know that's a GUI and not command line but no other GUIs I've seen have done so. 0
I find this one annoying because I generally don't want to run `git pull` -- I almost never `git pull`, I usually just `git fetch` and update my branches as necessary. I do wish there was a built-in shortcut for "try to fast-forward this branch"; I often just do a rebase or a merge, which will do the right thing for a fast-forward, but won't fail if a fast-forward is impossible. I can do `git merge --ff-only`, but I would like it if `git fastforward` or `git ff` was available instead because for me it is such a common operation.
To forestall the obvious -- I don't like to make custom aliases or commands (though I've done it in the past) because it makes it harder to migrate between environments.
git ff
Very useful to make a git alias for that. When git bitches at my about the unknown command on the work servers I know I can just use the full merge command.Despite knowing this is an option it's still hard for me to find decent documentation, but I think this is probably the best I found:
https://salferrarello.com/git-warning-pulling-without-specif...
- last: `git log -1 HEAD --stat`
- past: `git log --pretty=format:"%C(yellow)%h%Cred%d\\ %Creset%s%Cblue\\ [%cn]" --decorate --numstat`
- save: `!git add -A && git commit -m 'save'`
(Hat tip to those who helped me with these. I stand on the shoulders of giants.) commit -a
Tell the command to automatically stage files that have been modified and deleted, but new files you have not told Git about are not affected.
add -A
Update the index[...] This adds, modifies, and removes index entries to match the working tree.
So commit -a won't track new files, but add -A will.But maybe that's because I almost never use rebase, where it apparently is switched?
e.g. if you have your branch `feature-1` and want to rebase it on `main`, then you would do `git checkout feature-1; git rebase -i main`
Git will then switch to main (that's "ours" now) and then it will replay the changes from `feature-1` on top of it (that's "theirs" now) - like cherry-picking all your commits, but in sequence (and not actually merging them in `main`)
This is exactly the reason that git is hard. Its abstractions are so leaky that you'd think it's interface is designed to be a sieve. I really like the underlying model git uses, but the actual CLI does a terrible job of providing mechanisms to use it to the point where you have to pay far too much attention to the internal model to be able to avoid footguns. It's an indictment of a poor API when the best way to figure out how to do something isn't to search the docs but to figure out how to express the thing you want to do as an operation on the underlying model and then Google that to find the invocation that happens to map to that operation (and half the time, it's not even it's own subcommand; it's just some obscure flag to a grossly overloaded subcommand like `checkout`).
This breaks down during rebase, where the terminology gets reversed. The definitions are more accurately:
- "Ours" refers to "the branch whose HEAD will be the 1st parent of the merge commit" (during merge) or "the branch who will get commits applied on top of its HEAD" (during rebase)
- "Theirs" refers to "the branch whose HEAD will be the 2nd parent of the merge commit" (during merge) or "the branch whose commits will be applied on top of the other branch" (during rebase)
This gets a little more complicated during "octopus merges", where there are multiple "theirs" branches.
git restore --source=COMMIT PATH
I have learned and tried this before but still have the muscle memory of doing git checkout COMMIT PATHAlthough, the author seems a bit scared of reflog. I find it very helpful in a lot of situations, because it gives you the history of the state of your navigations around the commit tree - not the history of the tree itself.
(Disclaimer, I work at Perforce. Happy to answer any questions!).
You want to permanently destroy all work you've done since last commit (maybe it was a bad start and you want a fresh one).
The (new-ish) command you run is
git restore .> I think in the context of the merge commit ours/theirs discussion earlier, HEAD^ is “ours” and HEAD^^ is “theirs”.
Should HEAD^^ be HEAD^2 instead?
[1] https://git-scm.com/book/en/v2/Git-Tools-Reset-Demystified
The terminology and language of git is absurd.
True, in its current architecture it has no way to do it since it doesn't have a daemon or any background process which tracks user actions in real-time, all tracking is ad-hoc, only when git is invoked.
So it basically just side-steps the question, and internally a renamed file is recorded as a file which was deleted and another file which was created, "magically" in the same commit.
When you run merge or log, git tries to guess that a file was actually renamed based on similarity statistics between the text representation. With small, similar files, changed this leads to false conflicts.
I'm not talking theoretically here. I'm talking about something that has caused me, personally, and my organization, many developer hours and real money. I'm talking about something that has hurt productivity and confused me and my users. I'm talking about something that has caused me, just this week, to be alerted to go to work since this happened on a critical project and I'm as the "git expert" was deemed the only one with the know-how of how to deal with those conflicts.
I have spent 2 weeks during covid lockdown, 2 years ago, writing a complicated function which silently identifies those false renames and fixes them before calling git merge. This eliminated many, many support calls. Just last week I have fixed what I hope was the last bug with this function, a nasty corner case.
However this is specific to my situation, where the files are json files that were serialized from the database, and each one of them has a globally unique ID that users can't change. So I'm able to verify which files were renamed in which side and automatically make a "fixup" commit on each side so the 2 sides are as equal as possible (leaving real conflicts in place).
The situation isn't good since this function only works on a single branch merged from the remote. It doesn't work yet between branches. In fact I ought to begin working on changing it to work between branches as I am writing this very comment...
I have many, many other gripes about git, but many of them have been voiced on the comments in this page. But I generally mourn the fact that it is the de-facto source control solution in its current form. Since now its maintainers, quite justifiably, won't break backwards compatibility, and are mostly into fixing bugs and adding some small new features that are QOL improvements.
I still think that some of the bigger pain points could be addressed without breaking backwards compatibility. Similar to what was done with splitting the checkout command to switch and restore, while still keeping the original command.
E.g. for the rename problem above:
1) an optional daemon could be introduced that tracks renames in real-time and records them, both for display purposes in the log, and more importantly to avoid false conflicts during merges / rebases.
2) And/or, users could mark during merges that have conflicts, which files were actually renamed to other ones, in which point in history (perhaps by modifying a generated an input file like rebase does), to help git merge/rebase make more informed decisions instead of relying on statistical similarities.
> index
> A collection of files with stat information, whose contents are stored as objects. The index is a stored version of your working tree. Truth be told, it can also contain a second, and even a third version of a working tree, which are used when merging.
Correction: you can get unhelpful definitions of all the confusing terminology.
EDIT: I "let me google that for you"'d myself and found this [0] right away which explains it, I guess.
[0] https://unix.stackexchange.com/questions/306189/why-dont-man...
Instead of the original developer of the material trying to prognosticate what questions people would have, people asked questions and either other users or the developer could answer them.
$ tldr git cherry-pick
Apply the changes introduced by existing commits to the current branch.
To apply changes to another branch, first use `git checkout` to switch to the desired branch.
More information: <https://git-scm.com/docs/git-cherry-pick>.
- Apply a commit to the current branch:
git cherry-pick commit
- Apply a range of commits to the current branch (see also `git rebase --onto`):
git cherry-pick start_commit~..end_commit
- Apply multiple (non-sequential) commits to the current branch:
git cherry-pick commit_1 commit_2
- Add the changes of a commit to the working directory, without creating a commit:
git cherry-pick -n commitI imagine it as a competition between myself and an adversary with an uncooperative attitude, one who's prepared to act smarter, dumber, better informed and more ignorant than I am in order to to find gaps, inaccuracies, or ambiguities in the docs I write.
It doesn't hurt that it improves my understanding of whatever it is that I'm documenting -- and where it could be improved.
Now they make my eyes glaze over, probably because every command now has 57 options. I just google how to do the thing.
The things that make me most frustrated are when docs are missing, incomplete, or useless ("--foo: enables the foo option"). When it comes to man pages, at least in the good old days, I felt like I could always rely on them to cover all of the possible inputs, outputs, and errors, and to describe them, if not in the plainest language, in a way I could understand without being the author of the program.
For example, I don't have much bad to say about https://shrubbery.net/solaris9ab/SUNWaman/hman1/ls.1.html