Learn the workings of Git, not just the commands (2015)
developer.ibm.com
developer.ibm.com
> @wilshipley git gets easier once you get the basic idea that branches are homeomorphic endofunctors mapping submanifolds of a Hilbert space.
* https://twitter.com/agnoster/status/44636629423497217
(It's a 'spoof' on the Monad joke.)
Please let us discuss git in peace.
(Although I admit that this one was funny)
Because lots of people don't want to learn git. They want to bungle through it the same way they bungle through most of the software they touch. People who work with software professionally are very adept at bungling. Even through complicated professional software that other users would need training for. But git is resistant. Even rote memorization won't let you hide completely from learning git. There's a steep learning curve before you can half-ass it.
Depending on perspective this is either a damning indictment of git, or just another day at the office.
I've been helping people learn git for over a decade now. Unironically the easiest to teach are interns.
The thing is that something as trivial as version control should not require that much learning. Pretty much all version control systems other than git had already achieved that (by not providing such a vast "feature surface").
Of course this isn't exactly git's fault. It was developed specifically for Linux kernel developers, not (for instance) for artists on a game development team. The problem is that nowadays, people start to get the impression that git is the only version control system, and want to use it for things it wasn't designed for.
Most don't.
Most at GITHUB don't. It's why they use github.
Because it is.
> "git command line is confusing me"
Because it does.
Life being easier... maybe, I'm sure it would be for most projects. For some where git is required, likely not. I'd reckon MOST projects don't really need a DVCS; they have the much easier mental model and work quite well with a centralized "source of truth" that people can take from or give code to (coughgithubcough)
For example, ```git-dissect-tipdissects indexed local tips from all noted downstream upstreams, remotes commit graphs, etc.```
Anytime somebody is confused or stuck with Git, it's a great idea to helpfully send them some documentation links from there.
https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
It is surprisingly simple and efficient. I am starting to think this Linus guy knows what he is doing.
If the command and option names had just been sensible, it would be a lot easier to learn and use. Half the usability difficulty, maybe more, is just this.
Holy shit, yes.
`git branch` in gitspeak really means `git branch list` in human-readable CLI
But `git branch ${NAME}` does not mean `git branch list ${NAME}`, instead it means `git branch create ${NAME}`
There is no `git branch switch ${NAME}` but instead it's `git checkout ${NAME}`
There is no `git branch log` – that would be `git log`, whereas if you want what should be `git log` what you really want is `git log --all`
Don't get me started on whether modified files are added, staged, or indexed...
But most headache (and heartattack) inducing part of Git when I was a beginner was the "remote vs local repo" and additional complexities this concept introduced. Nothing more stressing for the new guy than making mistakes that screw up for others (mistakeful push, merge, mostly ) or if he think he did.
The article does a poor job (imo) regarding this. It skims on a 2 paragraph "Caveat" section how rebase(and later squeeze) can fuck local vs remote repo, but for most people I know, stress memories around this is what made them learn more about Git beyond than "branch & history keeping software & commands.
That and if you unfortunately end up in a project that uses subtree or submodules
Like LOTR, everyone has more than one name.
I think https://github.com/Byron/gitoxide also plans to provide a different CLI, but I don't think they've started working on it.
I bet if you created a great git desktop app with outstanding UI it would still have a few corners 99% of users would never touch or understand.
If that doesn't work you probably have done something really bad to begin with.
... yes? I've used all three of those commands in the past, although I don't remember what they do. and I'm hardly a git expert.
Anyway: “reset” changes the ref to which the HEAD branch points, and depending on options can also update the index and working tree to match as well. “Revert” creates a commit that is the reverse of an earlier commit. Restore rolls back uncommitted changes to a file.
This isn’t hard?
On the other hand, even among the professionals it has a reputation for being tricky and frustrating. Which is not the norm for most critical daily-use tools in this or any professional. So I'd also consider it hard in that sense.
git revert: create a commit that reverses a previous commit
git restore: you got me. I can't remember
Now, I generally do not remember many of the extra arguments that can be passed apart from 'reset --hard/--soft', 'git log --stat/--oneline --decorate', etc.
Personally I've found git complicated when I've worked at a place with bad 'git' discipline. When I've worked at companies with sane branching agreements, etc. I rarely face problems.
- reset: "blindly" load whatever commit you tell it into your working directory. This is the one if you add "--hard" will plow over your working directory in a way that cannot be undone. Actually I think without the "--hard" it won't make any changes to your working directory, it will treat whatever is in said directory as a change from the commit you told it to reset to. Dunno... have to look into it.
- revert: Creates a new commit that undoes whatever commit you pass in. Often times people mistakenly use this to "undo" a merge into a production branch that shouldn't have happened. The result is trouble when they want to push those changes back into production again.
- restore: I have no clue what this does.
All I really know is "--hard" is one of like two commands that you can do in git that you can't back out of. "git clean" is another one (I think).Yeah... conceptually Git is kinda easy but the command line is pretty nuts.
Git’s CLI is good as a unix-philosophy low-level tool that user-friendly tools can build on, and is a strength in supporting a wide range of different workflows.
It's suboptimal as a end-user tool for most workflows, but not bad enough for that people consistently build tooling rather than guides for particular workflows.
These are not toys, one should not expect to just jump into them and start doing work like you would not just jump into a Boeing 747 and just take off, fly and land. It having 100s of buttons is not a weakness.
It’d be interesting to see what a good redesign might look like. I do think some of the command names and command flags could be superficially updated/renamed/moved and provide a meaningful impact on git’s usability and learning curve. I’m curious though about whether the conceptual part of git is the primary source of difficulty, whether git is fundamentally a little hard to learn, because it’s fundamentally a little bit complex. If that’s the case, a rebuild might not help, there might not exist the kind of “sane ground” to build on that you hope for.
That said, this document in particular does not seem to be a very good way to get that understanding, it's a very "IBM documentation" way of presenting git. I wish I remembered what I read, because it's was a much easier read.
1) https://git-scm.com/book/en/v2/Git-Internals-Plumbing-and-Po...
Of course no one is stopping anyone interested in it from reading more about how it works up front.
I must admit that I'm at the point where I shrug, delete and clone the repo if I'm stuck. Some errors and messages were to cryptic to me.
I've used git for many years, and I still zip the repo folder, before doing a large merge, since 'fix the commit' can be a large effort with conflicts. Quickly renaming the repo root folder, then unzipping the old version can be quicker/safer, if you are not a git ninja and don't want to loose any work. There is probably proper command for it, but I sometimes get into a git mess that I cannot get out of. (reading stack overflow, trying get reset, git checkout, git reset --hard and so on)
Could you tell me what you think git is for and why you use it?
Though in this specific case of rewriting emails, there is a way in git to do it without rewriting history:
Personally, I git clone a fresh copy to do any advanced stuff in, and sometimes an extra just as a backup. Though that's not really much different than copying the folder. Especially if he has uncommitted files like IDE settings and whatnot he doesn't want to fix if things go really bad.
No, but I'm shocked git is so misunderstood that this would be considered a time to do this.
Keeping a backup of previous versions is git's raison d'être. If you can't trust git to do this, how can you trust "cp" to make a copy of a file? I would love to get a genuine answer to my question: what do you think git is for and why do you use it? This would help me enormously to understand where you are coming from.
Changing the git commit history (for example because you made a commit under the wrong name/email address) in particular can cause huge headaches.
Doing a git merge with tons of changes (think of clang-format white space changes in combination with various other commits) is another one.
This is not about the specific cases where your git repo can become a mess, but more about having a simple safety mechanism to always get you out of the mess.
All of us - at some point - were new to git and coming from other VCS systems meant there was a relatively steep learning curve.
I also understand git fear - some of the stuff is a little arcane, and organizations are often very dogmatic and loud about their usage and approaches to git. It leads to a lot of gatekeeping. Do we rebase here? Do we never rebase? Linear history? Rebase feature branches or merge directly into main? It makes the fear of the tool that much stronger.
`cp`, on the other hand, is pretty straightforward. There are no local approaches to using it, there's no real `cp` expert in-house at most places. It does what it says on the box without a lot of frills or options.
git is a complex tool. It becomes easier the longer you use it, like most things, but it can be very intimidating to new or intermediate users.
He's not new to it. He says he's been using it for over 10 years!
You shouldn't be. As you say, you shouldn't need to do this, but the fact that so many people do feel the need to do this speaks to poor UI design on git's part. There are, I think, two main reasons why people feel the need to do this.
The first is that git makes it really easy to rewrite history, without really offering much in the way of safety rails. You have but to be burned by this once to lose all trust in git whatsoever. And this is something where there are safer ways to rewrite history: Mercurial's concept of phases or changeset evolution is easily far better than git in this regard. Even exporting the excised revisions as a revset ("strip-backup") is far easier for me to undo than having to go into git reflog (especially because I don't need to race any `git gc` command--note that git is the only VCS that feels the need to have a garbage collector!).
The second issue is that git has a lot of different places where state can be hiding, and it quickly becomes unclear which of the various places a command is affecting. Is this going to update my working directory, the most recent commit, the staging area, or multiple of those copies? If I'm in the middle of an interactive rebase, do I need to `git commit` or `git commit --amend` or some other command to properly update "the" commit? Maybe if you're fluent in git, it's all obvious, but if you're not fluent, it's way too easy to accidentally do the wrong thing. And looking up git documentation doesn't help--it's the only tool I regularly use where reading the documentation actively leaves me more confused than before I consulted it.
It does, I just think they are not obvious enough because people haven't read an article like the one posted here.
Git is an append-only data store. You can't lose anything as long as you have committed it. Rebasing does rewrite history, but it doesn't delete the old history. It's still there. The reflog links to it. The old remote tracking branch links to it. You could even leave a tag there to link to it. There are many ways to get back.
If a small repertoire of git commands covers 95% of what a person needs git for, and recovering from backup covers the remaining 5%, why spend time and effort "git-ifying" that last 5%?
In the real world people have a ton of untracked files for their development environment, and uncommitted changes, but need an update from a coworkers branch that has a conflict with development, and a different conflict with your local changes. This can leave you in a weird state pretty quickly, and merge conflicts are a pain, especially if the conflicts are in a part of the code you're not very familiar with. Personally, I go full cowboy and work my way through it, but I am not surprised in the least when I hear that people make a backup so they can quickly restore.
One benefit of copying a directory instead of using git is that it will copy all of your untracked files, and uncommitted changes as is. Another benefit is that it works the same no matter what tooling you're using. I believe that some legacy code at my company is still on SVN and source safe. I also use open source projects that are developed with mercurial, bazaar, and fossil. If I were going to work on any of them, I would definitely be making backups the standard way instead of trusting my ability to use the tool properly.
Maybe it's not your intent, but you're coming off as a bit of a git yourself. He lacks your confidence.
But why? I'm trying to learn here. I'm know I'm autistic and don't understand anyone. It's clear that there's a huge difference in how I perceive git versus how a lot of people apparently perceive it. I'm trying to understand how other people perceive it but nobody will tell me!
git branch blah_branch_backup
Later, if your merge gets completely messed up, you can do:
git merge --abort
git reset --hard blah_branch_backup
That final command will restore the current branch's HEAD to point to the same commit as blah_branch_backup.
This same pattern is useful for other dangerous commands, like rebases or filter-branch.
When you're done, just delete blah_branch_backup.
One example: changing the commit email, somewhere on the middle of many commits. On github, I sometimes need to use the work email instead of personal email (due to https://github.com/apps/google-cla) Changing the email of an existing commit messed up my repo beyond repair. This is just one example from memory, it happens infrequently luckily. Thanks for the git tips :)
When you rewrite history, Git makes new commit objects that have the same topology as the original branch topology. It also updates the branch pointer to point at that new HEAD commit.
History rewriting commands essentially just fork the repo at some point in the past and create an alternate commit history. The original commits still exist, at least for a little while. Git will eventually garbage-collect them.
By creating a backup branch, you're "pinning" those old commits. With the backup branch in place, after you run a command that rewrites history, you will be able to look at the commit graph and see the point where the backup branch and the rewritten branch diverged.
I used to do what you do until I learned more about how Git works internally. And Git's really not super complicated at its core. I think much of Git's complexity comes from the minutiae of the CLI options. (For example, when you want to delete something, do you use `-d` or an explicit `remove`? It depends on the command you're running.)
But hey, if you have a workflow that works, then you do you.
Don’t forget to commit your changes if you’re going to `reset --hard` later!
Before you merge, the command is “git merge --abort”. After you merge, the command is “git reflog” to show your history, and then something like “git reset --hard HEAD@{1}” where the number in braces is the reflog entry you want to restore.
Reflog is the ultimate undo tool for work that’s been committed, it’s like having an infinite Ctrl-Z. If you use reflog to restore before your merge, you can even use it again to restore your mess, or to restore a different merge before the one you made a mess of.
The one thing you can’t recover using reflog is an accidentally dropped git stash. This is one reason I try to avoid stashes, and just use lots of one-off branches instead.
How it's done internally, as much as it's remarkable, is just an implementation detail. The fact that users are repeatedly being referred to docs/books, points to conceptual inconsistencies or mismatches in the intended application contexts.
As a user I operate with directories and files. Now version control also introduces a time-component, which is versions (and version of a version aka branch) and some kind of history/timeline/journal/log.
Why would a branch have a "head" and "base", isn't parent/child concept expressive already? Why "ref", if we've got "version" already? Why "stage", "index", "cache"? Why "cherry-pick" if we've got "patch"? Etc.
There should be a purely unix way to use git, as in "everything is a file". Sure, you can go to the .git folder and there are files in there. But the structure of these files does not directly represent the structure of the repository. What I want is to "mount" a git repository, so that I can explore its contents using cd, ls and cat. Also, by editing these files I implicitly make commits and create branches to the underlying structure. Would that even be possible?
EDIT: A disturbing feature that I don't see a way out is that in git there are "standard" and "history-rewriting" operations, and there is a clear distinction between the two. To what would that correspond in terms of pure file operations? Something about changing the default permissions?
cat `find . | grep commitmsg`
More interestingly, it would be able to grep around all previous versions of the files.> "history-rewriting" operations
Git is a Merkle hash tree. You can never rewrite history. Instead, what you think of as "history-rewriting" operations is just "non-fast-forward" changes to branches. Branches are really just named pointers -- {branch_name, commit_hash}. A change to branch from {$name, $commit0} to {$name, $commit1} is fast-forward IFF $commit0 appears in the linear history of $commit1. A non-fast-forward change to a branch is what you call a "history-rewriting" operation.
There's nothing wrong with "history-rewriting" -- I do it all the time. It's only ever bad when you do a non-fast-forward push to a branch that others track, and even then it's not the end of the world because usually those others can recover easily enough with `git rebase --onto`. Rewriting the history of a public branch is a problem, yes, because you want the history to be stable and reliable so you can use it to figure out problems, understand the history of the project, etc, and because making others have to `git rebase --onto` is impolite.
It's essential to understand that the only destructive operations in git are a) non-fast-forward branch changes, b) tag changes, c) gc/prune operations when you're like me and you're used to working in detached-HEAD mode.
If you're stuck on "Git has history rewriting, and that's bad!", then you've misunderstood everything about modern version control systems. You can -and I have- do the same sorts of destructive operations in version control systems going back to the early 90s (e.g., Teamware, CVS), to say nothing of newer ones like Mercurial, Fossil, and, of course, Git.
Inb4 "But Git encourages history rewriting" -- that is neither true nor a good argument.
Yeah, sure, why not? You could do it with fuse http://libfuse.github.io/doxygen/ There's even a git fuse filesystem already out there. I dunno if it works the way you want though.
The only other extremely important thing to understand about Git is that it has a way to name commits symbolically (branches and tags), and that those names can be changed, and that leads to understanding what a fast-forward and non-fast-forward change to a branch really is.