Oh Shit, Git
ohshitgit.com
ohshitgit.com
I would humbly suggest you avoid framing it that way, even if you believe it’s true. From someone who is only sometimes a cli advocate, my immediate assumption is that your opinion here may be formed out of naïveté and a bit of fear of the cli.
Please note I’m actually quite a fan of using various GUI tools for basic git workflows, and things like branching diagrams and even diffs are much better in GUI tools than in the console. Some of the tools are damn good and damn convenient most of the time. So you’ll get a huge amount of agreement from me about the benefits of sometimes using GUIs.
But git was designed around the command line, and there are no GUI tools (Kraken included) that have UI for everything or even most of what git can do. Advanced workflows and repo spelunking often require the command line. Using the cli is the only interface that can do everything, and it’s the advanced interface, so telling people who are already advanced (possibly more advanced than you) and already comfortable on the command line that they’ll find enlightenment in GUI tools isn’t always or even generally true. It’s not hard on the command line to see everyone else’s remote branches. Better to listen to them, and suggest GUI tools only when they express frustration about their workflow that you think a GUI can help with. Also perhaps better to ask the about CLI workflows and find out some of the benefits. One benefit is that CLI workflows are always usable over basic ssh connections.
And worth mentioning, the question’s partially moot for people using Github or other hosting services where some GUI tools are built-in. Using the CLI for all interactions with Github, and using the site for the visualizing diffs & branches is totally reasonable.
> I’ve worked on teams that just didn’t rebase at all and I think people are oblivious to the mess they make
I have to fully back you up and agree there! It’s unfortunate that rebase is a little scary and that some people just don’t like it. Rebasing local work before publishing it is the way git was designed to work, and it’s extremely helpful for creating a history that’s not insane. Squash commits are better than nothing I guess, but you can really tell and appreciate when people are comfortable with rebase and care whether their history is presentable.
I've found that using merge gives a readable trail of when something was merged, whether that be from a branch's original branch, or if you're merging into another branch.
Rebasing just seems to cause a lot more headache when something doesn't go perfectly correct in between commits.
Personally, I make a lot of mistakes. I commit too early when I missed some bugs or broken builds (eg build on release but not debug, rhel X but not rhel Y). My working area has unrelated changes. I forget to change branches...
So, I figure out what the destination looks like, do my best to keep things clean and small, and rewrite history before the PR as if I was one of those careful people.
My recent advancement has been to realize that reordering commits for an efficient fixup is much easier than splitting commits, so I'm better off doing things out of order. I also use worktrees to be able to ensure each commit is correct as checked out without stale state.
I think this part is really telling:
> And worth mentioning, the question’s partially moot for people using Github or other hosting services where some GUI tools are built-in. Using the CLI for all interactions with Github, and using the site for the visualizing diffs & branches is totally reasonable.
GitHub's branch view is uniquely awful, and doesn't try to call out branch history at all. Commit lists should either include a railroad diagram or only include the left-facing history. Anything else is just inexcusably broken.
In some sense that goes to the point that CLI can be better than some GUI tools. The ascii railroad diagram you get with git log might be preferable to what GitHub can do.
You mean merge conflicts? Rebase gives you an opportunity to deal with small merge conflicts arising from single commits. I much prefer that to a big merge commit dealing with all the conflicts simultaneously.
Also, rebase makes the merge conflicts disappear for future readers, making the included commits nicer to read.
Sometimes I think maybe we should have two different names for these commands.
I think an interactive rebase requires a good mental model of the changes to split or merge them correctly in addition to knowing the CLI. I find it quite hard at times.
I work on a project that has been around for ~8 years and there have always been ~5-10 people working at the same time on the codebase.
Nobody rebases, nor do they squash branches on merges into the branches. It's about as good as you'd imagine.
That's a very arbitrarily vague answer. It's very easy to hand-wave the misuse of a tool just by saying "you don't know it well enough". It's much harder to admit that maybe the tool is just difficult to use in the first place.
> It's much harder to admit that maybe the tool is just difficult to use in the first place.
That is unfortunately true. Git feels like a tool with several parts that were bolted on to solve problems its users (developers) faced. I didn't have that same difficulty with mercurial. I get this feeling that a pure patch based tool like pijul would match their workflow well and still be easy to learn.
I fall into the former camp, but others (who I respect) fall into the latter.
> I'd really like to know what mess people make when they don't rebase.
As a developer, we have the expectation that git will allow us to experiment with features, make mistakes, correct them, roll them back and improve. That stage is also so messy that you'd leave short commit messages that would make sense only to you. This is fine while developing. But this leaves many commits that are functionally broken, partial or rolled back later. That, along with vague commit messages and illogical commit order make it really hard for someone else to pick up a working commit from your branch and continue. Heck! I find it hard to choose a commit even from my own older repositories. That isn't the case with good projects like the kernel. You can pick any commit on the master and it will compile and work with all the features advertised up to that commit message. It makes a users' life easier.
> I've found that using merge gives a readable trail of when something was merged, whether that be from a branch's original branch, or if you're merging into another branch.
This is true. It's harder to achieve with rebasing. But it's possible with some work. I usually leave the original messy feature branch intact, and mark the rebased HEAD with a similar-worded tag or branch.
> Rebasing just seems to cause a lot more headache when something doesn't go perfectly correct in between commits.
Don't take any offense, but those are beginner blues. It happens in the early stage of learning rebases when you don't have a full grasp of what is going on. People evolve different strategies to overcome this once they are a bit more comfortable. My strategy is to create a temporary branch for rebasing (at the same commit as the feature branch) and do multiple rebases on it. I do only one or maximum two operations during each rebase. The result of each rebase is reviewed before doing the next round of rebase on the same branch. The original feature branch is left intact in case something goes wrong - though I never needed it.
Other people do rebasing occasionally while developing the branch. They do this after every two or three commits. They end up with a clean history to merge (fast-forward) to master, by the time they finish the branch. All of these operations can be simplified using helper tools like git-revise [1].
My absolute favourite method is to not use rebase at all. Craft the perfect commits as you develop. This can be achieved with a patch stack tool like quilt or stacked-git [2]. It allows you to move your changes to a patch (commit) of your choice. This is like having multiple staging indexes available. You can also split, combine or reorder patches. The patch stack evolves as you develop, but you end up with a perfect history to merge (fast-forward) at the end.
Objectively git kraken has more visual information density. To get the same information on the cli requires multiple commands and the user has to hold the information I'm their head between the commands.
Git kraken's buttons are also tightly correlated to commands and you can even bring up a window that shows what cli commands are being run under the hood.
I canceled my subscription and returned to the CLI + VS Codes native SCM UI because GitKraken is now useless in my work environment. Never mind the startup time is in minutes. And no, I'm not paying more for the ability to write a bug report and get it to them (their support is functionally non existent)
Great for a tiny project. Awful for companies. Not worth the money given how much more time it takes for me to do my work than use the CLI.
Sorry for the strong words, GitKraken is just a great example of bad design.
Can you give a precise example? The cli has support to using visual tools for diff/merge, so one doesn't need to use git via gui in order to use graphical interfaces where effective.
Git cli purists are probably cli purists in general, not just for git. And for good reason.
When you work with the cli/terminal, you aren't bound by what the creator of the GUI decided should be built.
> I often suggest git CLI purists to get something like git kraken and just use it as a visual dashboard. Watch what happens when you run git commands. You can see everyone else's remote branches and have a much better idea of what's going on than you can without it.
There are git commands for displaying whatever information you want, capable of drawing nice trees and whatever complete with everybody else's branches.
> It's so easy to create a new branch before I do anything that if I think something might go south, I just hard reset to my previous state (a branch I created right before the operation)
Duh! That's what git is for!
And when you say "It's so easy to create a new branch" I'm genuinely curious to know how you've been creating branches before you discovered the "state of the art GUI tools that change the game".
However I'm not sure Kraken is a good example. The GUIs for git are all similar, but Kraken has some unfortunate combination of being counter intuitive at times, and very easy to access undo and force buttons.
We have a lot of developers using all common GUI tools and the ones using Kraken are the only ones who not only regularly end up not only shot in the foot but force pushing the remnants upon their colleagues. For what it's worth, the few using Magit seem to have the least trouble but I suspect they may be more familiar with their tools. The ones using IntelliJ seems alright too, as long as they don't venture too far outside the familiar edit and push cycles.
Generally, any graphical git tool should probably be as closely integrated into the IDE as possible. That's where it's most useful.
I can't emphasize how pleasant those operations are in magit. I'm not sure if it (TUI) would be considered as textual or graphical. However, the interface model should lend well to a fully graphical implementation.
It is also with a GUI that I learned you could rebase on pull, instead if merging and getting an ugly "Merged remote tracking changes" commit.
I encourage everyone I know to do this.
Just last week, I found out that GitKraken doesn't support connecting to multiple Azure DevOps organizations within the same profile. So my team now needs to learn git CLI anyway, to efficiently migrate our 30-odd repositories.
It’s not there yet, but I kind of think git should be used the way I started using the save option when computers were less reliable. Type a sentence, save. Type another sentence, save.
Git has an extraordinary collection of foot guns and unwritten rules. I’ve been using it three years and often feel like I just get by (and also the devs who came up with the conventions in my org maybe could have chosen better)?
My theory is that people who like the CLI have a good (or just working...) short term memory.
I have a sample of 1 to prove my case: my SO, who has excellent memory (short and long term) and handles the git CLI just fine, and really thinks all git guis are like bicycle training wheels for kids.
I on the other hand use tortoise git almost exclusively (and get mocked by her that I'm still riding on training wheels...).
I actually use only Tortoise Git's git log view, all the important actions are available from there, instead of accessing them from their somewhat haphazard right-click menu. In this way it is very much like Git Kraken or SourceTree but more powerful, since it exposes (IMO) the most useful subset of git commands AND also shows in a single glance all the necessary information about the repo. I find it very easy to explain git branching / merging / rebasing / push/pull to new people using the tortoise git log view. Basically I'm using the same explanation as the one written by somewhereoutth below [0] but without referring to brambles... I also always stress using git reset --hard to get out of really confusing and bad situations, assuming you had committed first.
Three random complaints about git:
- The fact that it doesn't track renames at all, and instead just makes an educated guess based on comparing the file content. I'm not sure how this could be fixed, though, since git doesn't have a daemon running in the background keeping track of changes. This can lead to the dreaded "modify/delete" or "delete/delete" conflicts. Granted, this would happen less if people would merge/rebase their work branch often instead of accruing commits over time and not occasionally merging from trunk. But it does happen, and it's really unpleasant, it often means that git got confused and it takes a lot of digging and comparisons through history to find out why. Having a gui helps me a lot here but it still a heck of a detective work.
- The fact that there are many workflows and ways to do things and doesn't have a recommended way to do stuff. This can lead to arguments and conflicts, especially between people who have "opinions" (I'm definitely including myself in that group of people!).
This can happen especially with people who switch teams and where the new team works in a different workflow than what they are used to. Actually I've been part of just such an argument this week, it was unnecessary and unpleasant. Especially considering that at the end we both understood each other and agreed that my team's workflow is really not that different than her last team's workflow (for which she was responsible), and our team works under different set of circumstances and with different clients.
- Not specifically about git: Personally I never found it necessary to have a clean history, and I'm confused as to why people think this is important. I certainly don't require this from people in my team.
I review pull requests regularly, and my way of doing that is to run a diff between the start of the feature/bug fix branch and its tip, to see the set of changes introduced by this pull request. I usually don't care about all the "dirty" commits in the middle, although I will look through them if I don't understand something to see how something has developed.
But I never review pull requests solely through the pull request GUI (in github or azure). I fetch their branch locally and do what I just described through tortoise git. It usually takes me only about 30 seconds to generate the diff which lets me see the files changed. Reviewing the pull request itself usually takes far longer compared to this, especially with people who are worked with areas in the code they are unfamiliar with. I write all my notes offline in my favorite editor, and then copy them into the pull request web GUI. I also try to compile and run their local branch on my machine and run the relevant unit tests added by the author to verify they run.
I find the pull request GUI in repo hosting sites to be very ugly and cluttered, and it only shows the latest changes.
So my theory is that people who insist on having a clean history only know how to do review pull request through the web gui, which does show the diffs in the last commit and doesn't show the whole history. But this is probably me missing something or just my 5-person team being small compared to others.
The rename aspect is a problem in both cases and is related to how git works internally.
Dirty commits van be really useful for the above, do a commit to rename, then another to modify. That keeps the rename operation clean and modifications in their own diff.
For workflow, that is a team process and documentation problem. Every team should document the typical CLI commands for their workflow. Having no documentation around this is negligent, simply referencing some webpage is lazy, it should be spelled out for you
The bramble stems and branches are the commit history (which may join and well as split, unlike a bramble). The labels are stuck to particular bits of the bramble. Some labels are even stuck to other labels!
You can move the labels around, and you can glue extra bits of bramble to the tips of what you have. You can even hack about with the bramble, but this is not recommended. Label moving often happens automatically, e.g. when 'growing' a bramble tip.
It gets interesting when you compare two brambles (e.g. remote and local repos). You might determine that one bramble is identical to another, just with the labels moved. Or one bramble is the same as the other but with extra stems added. Or both brambles had a common ancestor bramble, but now they both have extra stems. Or perhaps they are completely irreconcilable.
Better visualisation tooling would help.
Here's a 2019 presentation of that talk.
Take the recursive merging stuff around minute 39. Do I need to know why git's model for merging is so much better and how it works its magic? I don't think so. It's an implementation detail.
Do I need to know how a gearbox works to drive stick shift?
I have never thought of git as a bramble (tree worked fine for me :)) but the thing I always tell people about too is the labels part. What I think is enough for people to realize is that it's just like say SVN or CVS or any other source control mechanism in that there's a tree (bramble) of commits and then every commit can just be pointed to by a label. I can move those labels around any which way I want. Everything else follows from that on a surface level that is the only thing required to work with git in most situations, including some advanced ones.
You don't need to know why certain operations in git are faster, better, have less conflicts or how they work internally. You just need to know what they do and when to apply them.
I don't need to know why I can't make my car start or switch gears without pushing down the clutch. I just need to know when to push the clutch, i.e. if I want to start the car, push it. When I want to switch gears push it (well, mostly, on cars you want to last a while still lol). Of course some people won't even be able to learn how to drive stick shift and can only ever drive automatic.
No, but it helps. If I treat my car's powertrain as a black box then it won't be intuitive that:
* I shouldn't slip the clutch to hold my car on an incline.
* Blipping the throttle gives me smoother downshifts.
* Double clutching lets me shift from 2nd to 1st while rolling.
* I can use the engine to slow my descent on longer hills and avoid brake fade
* I should park the car in gear for safety
* etc.
Whether it's better for a given person to memorise that list of facts or to understand the concepts behind them would depend on how much they drive. As a developer I've found it helpful to gain an understanding of the tools I use daily beyond "here's the commands to copy/paste when you want to do xyz".Like you say, you can just remember those things. In fact 4 out of those 6 were taught in driving school and you had to remember them. One I only learned because trucks still needed that (double clutching while shifting - not just your case, just in general) but cars didn't when I learned. I personally don't like copy and pasting commands like that but I see a lot of people doing that even for stuff that should be second nature because you need it all the time. I think - to stay in the analogy - for me this is the difference between knowing that I should engine break and how to do it when I want to slow down vs. having a piece of paper in the glove compartment that I pull out and check for what to do and how every time I approach a red light ;)
I'm also someone that likes to get an understanding of the things that I use and do every day. The thing is that there are so many things we use and do all the time that I think (almost) everyone just has to keep a certain level of abstraction away from many things, because there are just so many rabbit holes out there and it's not beneficial for most people to have explored every single rabbit hole all the way to the end (the 'gradient'). Git commit graphs are a DAG and lots of cool things can be done with DAGs, most of which I totally forgot about since I learned about and enjoyed them in university and have never needed them again at that level. It's good to know they exist and be able to dig in when needed/wanted.
The interesting part to me of your analogy is that it can demonstrate how having an understanding of what's going on under your layer of abstraction allows you to generalise.
To wring the last bit of life out of the gearbox analogy: if you tell a mechanically inclined driver that blipping the throttle will make their downshifts smoother, they'd hopefully understand how rev matching can be generalised to upshifts too. If you tell a "black box" driver then they probably wouldn't be able to do the same. Of course, the analogy falls apart a bit because understanding rev-matched upshifts isn't particularly useful :)
I've forgotten most of the "advanced" git knowledge I ever learned but I get a lot of value out of the (admittedly not very advanced) understanding that, as you said, commits are a DAG and that branches are just named pointers to nodes on the DAG. That understanding lets me generalise to, for example, backing up a branch (with git branch/tag) before doing a tricky rebase so I can restore it (with git reset --hard) if I need to undo.
I agree that exploring every rabbit hole to the end isn't beneficial - understanding the DAG is typically the only "advanced" git knowledge I need and I've only very rarely had to peel back more layers. I think this:
> It's good to know they exist and be able to dig in when needed/wanted.
is a good way of putting it. My ideal low-level understanding of most tools is knowing just enough that I know what to search for if I ever need to go deeper.
You can also do `git reflog` on branches to see what they used to point to :)
There was no handbrake.
I managed to work my way back down to the dealership and asked, "Where's the handbrake?"
It turned out that Subaru had decided to replace it with a button and an "automatic" "hill-holder" feature.
Also, if he's literally 3 inches, I would say the best approach is to slowly ease off the brakes until you actually touch the other car but do it so slowly that it's not an impact. All without even pushing the clutch. Then you don't even need to use the brake while pushing the gas pedal trick, because the guy behind you is your brake pedal ;)
Even off-road, I really only need to know to not ride it all the time.
I learned to drive stick in the Bay Area and the way everyone I know there drives stick on an incline is to use the parking brake with a second hand when letting off the brake and letting out the clutch.
Now, when you get good you can stop doing this for most, but for really steep streets (I now live in San Diego and we have a couple; Laurel being one) it’s still an excellent skill to have.
Imagine: Small towns, really narrow one way streets with foot traffic and underground parking. Getting out of some of those underground car parks is scary stuff!
You get onto a steep incline to get out of the car park but you come around a corner onto that. Cars might be coming down towards you at this point and they're sometimes hard to see. It's cramped too. So you can't just take it w/ speed to get up there. Also on the top you have pedestrians on your side, so you might need to stop on the incline, then when pedestrians have scurried away, go a little further but not directly onto the street until you can actually see if cars are coming. Half your car is still on the incline at that point.
Is there an extension or something for Git to have similar behavior to this?
I try and make a habit of squashing my commits so that git log is a narrative of development and each commit is a page number, not a sticky note.
It dominated the world despite the UX flaws, which suggests they really aren't that bad.
Yes, you need to follow a tutorial to use it effectively. Designing for beginning users is often detrementantal to advanced users, and I object to that being called "user friendly".
I'm not very vocal about git usually because I'm not an expert I have no big complaints, and I'm not that interested in arguing about it. I think there's a lot of us that are perfectly happy with git; we just make less noise.
At work, we switched from SVN to Mercurial, and from Mercurial to Git. Most people were happy to switch away from SVN, but few enjoyed the change from Mercurial to Git. Personally, it took me much longer to get used to Git than with Mercurial, even though Mercurial was my first DVCS. I now have a slight preference for git, but I am happy with either (just don't bring back SVN!).
Why Git came to dominate and not Mercurial? I am not sure, but I don't think technical reasons explain everything. Its association with Linux probably helped a lot.
> Why Git came to dominate and not Mercurial?
The answer is actually really simple and non-technical: Github. We take these feature rich online code hosting platforms for granted now, but Github was really the leading edge of the wave. It made it easy for people to work together on writing software so it had great takeup and started a Git snowball effect.
I was an early proponent of Mercurial over git. While they are similar, a few “minor” things made a big difference:
Mercurial distinguished between branches and heads, whereas git did not. This added extra complexity.
Git embraced checksums as identifiers while Mercurial provided local revision numbers, this obscured the mental model.
Git is faster to pronounce than either “Mercurial” (which has a tricky vowel in there) or “hg”.
don't make the mistake of comparing "git now" vs "svn then".
It has been described as distributed SVN. I really hated SVN, almost to to point of wanting to go back to CVS, so I didn't want anything that took any inspiration from it. But maybe, in reality, it is not that bad.
Yet, whenever I've worked on git with others, it's been on github (i.e. a centralised model). And I've worked in a decentralised way on svn on several occasions, simply by making my own local repositories and merging changes back to the parent repository when I'm done (which is effectively what git does too when working with remote repositories).
I feel that a lot of what ends up being 'bad' about svn really boils down to the fact that you need some good conventions to be honoured across the project in order to get things done (including using it as part of a decentralised workflow), but humans being humans take shortcuts and mess things up for everyone else. Which really means at the end of the day, the problem isn't the technology per se, but human relationships, manifesting as commit behaviours. Whereas git just imposes its highly opinionated model of doing things on you in order to ward off some of the more destructive human behaviours, which in a sense is good, but at the same time, it means that git can be too rigid, and svn effectively gets a bad rep for being potentially more flexible and scriptable. I've been in many situations where I had to get into an incredibly convoluted manual process to work around git's mental model to get it to work for me, when the equivalent in svn (for better or worse) would have been fairly straightforward.
(Disclaimer: This is just my personal experience from happily using both with no specific preference for one mental model over the other. If anything, I think I may probably prefer svn a bit more now. I'm probably completely wrong about all of it.)
If only that wasn’t the case we’d all be using a sensible source control system now instead of one where people repeatedly say things like “if you just learn the underlying model…”
GitHub for many years only had free public repos, not privet, BitBuckets USP was that you could have privet repositories for free.
GitHub "won" because of the social aspect around it and the tooling for open source projects - it never had more generous free levels.
(I like Python! But it's a scripting language, not a systems language. Also the "deployment story" is bananas.)
The worst systems are the Microsoft ones, but the one with the most complex interface is definitely Git.
I would argue it won because of github. I'm using Git because of that, but still prefer Mercurial in every way.
No it doesn't. It suggests the other features are so damn good that it overcame the horrible UX.
Especially if you'd have to lose github, thats the real force
If gh migrated to something else, then a lot of ppl would follow
If you prefer you can use mercurial as a git client in the same way.
Most of my git use stays local on my PC, uploading it to github or gitlab or whatever is a secondary concern for me.
I bet git produced by a no name Linus, in isolation not required by any major project, would have hardly been taken any major uptake.
I think parent's point is that that demographic is positively tiny, and there were plenty of other serviceable offerings around at the time.
The claim stands - git won despite its UX flaws, which does suggest the alternatives, while serviceable, weren't sufficiently fit for purpose.
I think perhaps you're over-estimating the network effect for VCS's -- having multiple version control binaries on your machine is low cost, and aliases can, up to a point, give you a consistent interface to them all.
The highly sophisticated demographic of kernel developers would not have been at the pub insisting that their friends drop svn and use git (exclusively!), anymore than they would have been trying to get any other projects they worked on to adopt bitkeeper previously.
The fact that bitkeeper was required for 'collaboration in Linux development' for so long supports this position.
There were perhaps 5k kernel contributors in 2005, would that be fair? There's perhaps 20k now, I guess.
github alone has 80+ million users.
It's possible, I suppose, that the former is predominantly the cause of the latter, but it seems hugely unlikely.
6 years ago Larry released BitKeeper under an Apache licence. I still don't anyone who's ever actually used that.
In software we often go for further complexity instead of less. I think because most of us who are the lead developers are often the most intelligent. And we often enjoy these complicated abstract models and they come easy to us. However in satisfying our own intellectual vanity we often don't see how many we leave behind. Which is good for our hourly rates, but less good for creating affordable and simple software.
Anyway one of the main practical advantages git had was decentralized repos meaning you were not dependent on an external server which meant git was often much faster if working with in daily tasks compared to centralized versioning systems
https://stackoverflow.com/questions/804115/when-do-you-use-g...
The idea of a directed graph of file system snapshots is pretty intuitive. Add in branches as pointers to locations in that graph. This is a fantastic model for source control.
However the operations that stage a potential update to the graph of snapshots is prettt confusing. The "index" is a terrible name that is overloaded with so many other non-git meanings, non of which really map to git usage.
That, and all the rest of the names are pretty hard to understand. Particularly reset, whose documentation is inscrutable without translation from git-speak into technical language, or at least a dictionary of what all those words that are used actually mean. And since reset is such a useful tool and has about eleventy different functions, it all becomes impossible to learn from the docs on your own.
The well-documented underlying model used to confuse me a lot - especially for operations like merging, rebasing (squash, reorder, ...), cherry picking, etc. They started to make sense only when I realized that git uses diffs/patches to propagate changes between unrelated commits. While the on-disk format is purely snapshot-based, the tool itself is a hybrid. As far as I can tell, this trips up a lot of others too.
Even the operations which do show will often require more thought to undo than a hypothetical "git undo" would. I know how to use the reflog but I often go out of my way to avoid it because "git branch tmp HEAD; git $POSSIBLE_MISTAKE; git reset ---hard tmp; git branch -D tmp" requires less effort than deciphering the reflog's output.
Undoing a push does require a different set of commands, but my point, to the question @amelius asked, is that you can undo both push and add, and whatever other mistake you’re thinking of, difficult or not.
Git, coincidentally, does have something equivalent to undo history: the reflog.
git reset HEAD~1
Usually works good enough for me. Idk if this is the proposed way to undo stuffI doubt that was the intention. Linux just needed a versioning system tailored to its needs, and that's exactly what Git is. Can't blame its creators that other people used it for scenarios it wasn't built for.
This is not one of those times. I agree there's a chance there are better word choices or feedback for some commands, but overall once you 'learn the language' it really is a lean, mean, well designed piece of software.
Folks that complain about the UI/UX don't realize it wasn't designed for less technical folks. It was designed for the folks who needed it.
It's success must at least partially prove that the UI/UX is not 1/10. Any real engineer will tell you, there are times where they wish they could do something better but the requirements and constraints left them making tough decisions, and that doesn't mean they aren't proud of their work.
It's success is also partially due to the fact that it is lean and mean, which allows it to be applicable to nearly all software projects of any flavor. So I don't understand why folks argue it could have been done better. If it was 'better', in my view it wouldn't have been successful. The success was driven by it's succinct design and Linus' take it or leave it attitude.
Technically accurate studio monitors don't sound as pleasing to the ear as good well tuned speakers. But they are exceptional at the job they were designed for. This is like that.
(FWIW, I don't consider git complicated and I'm far from a power user)
Git has enjoyed overwhelming success, which would seem to empirically indicate that it's done something right in terms of design
Perhaps the obviously wrong UX/UI isn't wrong? Perhaps UX/UI people aren't good at designing interfaces for experts?
It's open source, so why hasn't an alternate interface taken over?
For example, the master/main branch shift. Everything broke when that change was made, but it happened and it wasn't a big deal.
I'm not seeing the difficulty here. It seems that a more reasonable interpretation is that git has the type of interface that is hard to learn, but intuitive once learned.
For CLI, it's easy to create any number of command aliases and scriptlets to have the exact UI you want..
Projects like Firefox and nginx use it, although others like Vim and Python switched to GitHub. I don't know if Facebook is still using Mercurial internally, but for their public stuff they use GitHub.
I think that if GitHub had supported Mercurial like BitBucket or Google Code it would still be a lot more popular, but ah well...
I see this all the time, this "commits ON a branch" mentality.
My secret to understanding git is: branches are just adresses, labels, pointers, aliases to commits. A branch is just a label pointing to a commit.
So, you don't "remove (or add) a commit from a branch", you change where a branch point to.
Commits are tree nodes, they have a parent and they can have children. If you point a branch to a commit, you can now think of a branch as that commit and its parent commits.
So, `git checkout master` is just `git checkout <commit>`. But in a smarter way, as git is storing this as a special reference for you. If you do a `git checkout commit-ref` git will warn you that you are now working in a commit tree without a special name (detached).
Oh, I wrongly "commited to master". Ok, just point master back to the previous commit.
Oh, I want master to point to another point in time. Just find the commit reflecting that point in time and `git switch master; git reset <commit ref>`.
Reset should be called "point the current branch to this commit and make the working tree reflect the state of the code at that commit".
Well, once committed, the corresponding deltas to parent are part of the branch (that is a node of the branch graph). Thus, if user wants to have this commit in a different graph, then the deltas need to be regenerated and recommitted (then deleted from the wrong graph).
As for the branch name, in Git it's just a label for the graph leaf. However, some SCMs maintain branch name for all constituing nodes in the branch graph.
~/p/s/sigrok-cli> cat .git/refs/heads/master
7bdc46d6fcdcfa00dd29ac102f3128aeb44a479c
In principle all these objects are in the .git/objects folder. But the encoding is some binary thing for branches, and anyway it's often stored differently (I assume for space saving reasons). But git can explain what every object is. ~/p/s/sigrok-cli> git cat-file -p 7bdc46d6fcdcfa00dd29ac102f3128aeb44a479c
tree fa02da4030e5791695b8a7a5400ee2c2cc88ee46
parent 09fc39da486b8618d65b66ca96a760478c3e89e2
author Uwe Hermann <...> 1454100673 +0100
committer Uwe Hermann <...> 1454100673 +0100
NEWS: Update for upcoming 0.6.0 release.
~/p/s/sigrok-cli> git cat-file -p fa02da4030e5791695b8a7a5400ee2c2cc88ee46
100644 blob d1d8ffca429d2a4697a56fd7dedbeee3fc6d577d .gitignore
100644 blob 28c1b20e6d7d146fbc6ad0270ff5b09cfd057b26 AUTHORS
100644 blob 94a9ed024d3859793618152ea559a168bbcbb5e2 COPYING
100644 blob 04c29abe3dc42c8ecb8ec1b71093437d8c382ebd HACKING
100644 blob c6d6c8a7cff1a5b7d940e0c040ad1d4f5337f646 Makefile.am
... etc
The git software just updates the branch to create the illusion that it's not just a pointer to a commit. But really, it's literally a file with the sha1 hash of the commit.Not sure what you mean by deltas here, but if you are taking about changes, git don’t store changes, each commit references the state (the content) of the whole repository.
The diffs are just a representation calculated to you.
Like: `git show <commit B>` shows a diff, but actually what happens is that git calculates the difference between all repository’s files as they were (their contents) in commit A vs their contents in commit B.
Git does that in a very efficient way, but commit are actually snapshots.
The only time deltas come into play are in pack files, and those store objects in completely arbitrary order.
I've seen this a couple of times at $DAYJOB when someone doesn't understand how rebase works, smashes keys until it looks like they've achieved their goal, and then I have to ssh into their box and try to unravel the damage they've done, hoping that reflog has not been GC'd yet and there's something to revert to.
That's easier than trying to get git working right, lol. It feels like every git command is a PhD rabbit hole. All I know is that if I screw something in git, trying to fix it will just screw it up worse. Reset early and often and it usually works in the end. Lol, it's terrible...
But that doesn't work if you rewrote history in some way (I'm not sure why that's even possible). In that case your local git can get pretty messed up, some gui tools (IntelliJ) start to bug out and fail to diff properly, and it's easier to just start over.
After ten years of using git, I've more confused by it than ever. People here keep talking about mental models, but I've read a shit ton of docs on it and am still totally confused. Probably some of us just aren't smart enough lol.
To reset the current branch, you need to do two things:
$ git fetch
to fetch commits from the default remote and point remote branches accordingly, and $ git reset --hard origin/master
to reset HEAD to the remote branch. Substitute master for any other remote branch if you wish.That's it. I don't know if this is more difficult than re-cloning from scratch (especially if you're doing frontend and then have to reinstall node_modules and such…)
Same thing for node modules, lol. Delete the folder and try again. If that still doesn't work, delete the lock file and try again.
It's a terrible practice, I know, but it's the only thing that actually seems to work in my experience.
EDIT: Apparently origin/master points to a specific branch. origin/HEAD points to the top of the "default" remote branch, which is often, but not always, master. Or something like that. I don't know for sure.
First is that HEAD always points to the commit that you currently have checked out. So, for example, if you have the master branch checked out, HEAD will be whatever is the latest commit in master in your local copy of the repository.
Second is that to git there is no difference between the originating repository and your own. If you had your copy of the repo fully exposed, the remote machine that you cloned it from could push changes to your copy exactly the same way that you push changes to the remote machine.
That second part is important because it means that the remote machine you're pulling from also has its own HEAD pointer. When you reset your state to origin/HEAD, you're telling git to set your own HEAD to point to the same thing as the remote machine's HEAD. This is very likely to be the same asking it to reset to origin/master because HEAD is probably pointing to master on the remote machine.
The reason it's likely that HEAD=master is that in most circumstances the repo on the remote machine isn't manually being touched. When you first setup a repo there is only a single branch so that is what HEAD points to. If no one ever logs into the machine and executes a checkout command, that's what the server is going to continue to point to.
However since there's no guarantee that HEAD=master, you shouldn't rely on that and instead always use the actual branch name.
Even though I understand git really well it’s quicker to just blow it away and start again rather than accidentally make it worse realising you didn’t understand it as well as you thought you did.
Fortune favours the lazy.
[0]: https://github.com/blog/2019-how-to-undo-almost-anything-wit...
Pro tip - this is true for commits, but not for accidentally dropped stashes. This is why it’s better to commit first, branching if necessary, than to stash.
“If you mistakenly drop or clear stash entries, they cannot be recovered through the normal safety mechanisms. However, you can try the following incantation to get a list of stash entries that are still in your repository, but not reachable any more [git fsck…]”
https://git-scm.com/docs/git-stash
> Oh Shit! I accidentally committed to the wrong branch! [git reset … git stash …] A lot of people have suggested using cherry-pick for this situation too
Because of the above warning about stash, cherry-pick is indeed a bit safer and more easily recoverable if something goes wrong when moving to the other branch. This particular situation isn’t dire given the premise that you already committed, so the reflog can be used. The situation where it’s more important is if you’re sitting in the wrong branch, have uncommitted work, and git won’t let you switch branches first. In that case, committing first into the wrong branch is recommended over stashing and switching branches, even though it’s slightly more work. I’ve actually watched people mess this up and then get frustrated with the magic incantation fsck stuff and give up and spazz and nuke their repo instead, while shouting “wait! no no no no…” over their shoulder.
Stash is convenient sometimes, but never necessary, always less safe, and there are always commit flow alternatives. This is why I avoid it myself.
Stash is effectively equivalent to
git switch -c stash-branch-$i
git commit -am stashed
git checkout $currentbranch
so I use that instead. Only instead of an anonymous name, I pick something sensible in case I get distracted and return to the work later. Also, it means I don't ever need to commit to the wrong branch -- `switch -c` carries uncommitted changes along. Sometimes cherrypick isn't a fine enough granularity, and I'll use difftool to distribute changes between two branches.At least until someone does a foxtrot merge and then it never works again
My only complaint I have for git is it does not support $Id$ and other RCS variables. IIRC that was by design.
Also useful are git-fast-export and git-fast-import, if you really need to delve into the inner details of a commit. For example, I had three separate but related git repos that I needed to merge, so I created a new repo with separate branches to hold each repo, merged everything manually, committed that to a new branch, then used export/import to edit the commit to have the tips of the three other branches as its ancestors. Maybe there's a better way to do this with other git commands but I found it easier just to delve in and edit the commit data manually.
Similar story here: at a previous job we had a monorepo with a Rails app and Rails engines extending its appearance and behaviour and per customer.
At some point the architecture became problematic and we moved towards a shell app, a core engine, and extension engines depending on the core.
We refactored code to that end, and "forked" the original monorepo into multiple clones, one per component, then stripping the other components in each one, ending up with 1:1 repo/gem/component. This worked for a while, easing a lot of issues we had previously, allowing for proper dependency expression, independent development and releases... Everything was great and we lived happily ever after.
Then much later we hit a snag (I can't exactly recall what that snag was, IIRC it was not technical but organisational). So we looked at options and decided to merge into one single repo again. To that end we could do a big code drop, starting afresh, but (again I can't recall why) there was a need/requirement to keep at least some git history.
But at that point, "some" ends up ~== "all". So I devised a plan.
I git init'd a blank repo, added each one of the repos as separate remotes, and fetched each of them. Thus all git objects of all these repos were present. Then I checked out each one of these remote's master as a separate branch, created a subdirectory with the component name, moved every file for that checkout into that directory, and committed that. This way a) each project could live separately in the new monorepo and b) there would be no conflict for a merge.
Then came time for the merge. Two options: a) perform N merges subsequently or b) perform an octopus merge. a) just felt wrong and ugly, so I decided to try if I could work b) out, but I ended up not being able to achieve that with porcelain commands as git was being too smart and attempted to look into the history for some reason I can't recall which produced senseless conflicts (IIRC git merge isn't entirely assymetric)
So, since merge commits are merely commits with more than one parent I figured out I may be able to do that with plumbing commands instead. So the steps were:
- for each branch, check out content (but without moving the current HEAD, so, actually, export the git tree corresponding to a specific ref/sha)
- add all that to the index
- create a merge commit object with each branch's sha as parent
And it Just Worked, with the bonus that since up til the commit where we forked, each branch had common parents that were untouched, and commit history properly zipping up, by git's design, which is really a DAG of commit objects.
My first "production" introduction to git was really about someone polishing their resume and chanting magic words at me. Merges ... happen? Who decides whether Bob or Cindy's code is used here? It just happens okay.
Then I bought some books on git and was pretty unhappy with how arcane the naming was. How is "reflog" the correct and intuitive choice for "undo"? Was there something I wasn't understanding? No, I'm just supposed to accept that.
Thankfully, my current programming gig is simple enough that I don't have to look at git. Either eventually sanity will come to the naming of commands or something else will appear, just as it has for every other tool.
Actually your initial intuition was correct. There are things you obviously don't understand about git. "reflog" is short for reference log. As "git help reflog" will tell you:
Reference logs, or "reflogs", record when the tips of branches and other references were updated in the local repository.
It isn't arcane. It's perfectly logical choice for what the "git reflog" comand does, as explained in the man page: "This command manages the information recorded in the reflogs."Now if you don't know what a reflog is, you will be confused by this. But the solution to that is for you to learn what a reflog is. And yes, in order to use git well you need to have at least a rudimentary understanding of how it works.
I know that some people do find git hard to understand. Often that is because they want it to work a different way than it does. But git is a fairly complex tool designed to solve a very complex problem. It does an outstanding job of doing what it was designed to do. However, if you are not willing to invest the time in understanding git to a minimal level (and many people aren't), you will find it to be confusing.
There is no need for "sanity" to "come to the naming of commands" for git. The commands already have sane names. But if you don't know what those names mean, it will seem to you as if they are in a foreign language. Most of the confusion people have with git is due to them having an insufficient level of knowledge of how it works. Again, if you want to use git well, you have to gain that minimal knowledge . If you are not willing to do that, then you should either become willing or choose to use some other version control program that is more to your liking.
Of course everything has an obvious name if you have to learn everything about it. The point I am making is that the name ought to be obvious before you have to learn everything about it. "Undo" is a reasonable choice if you barely understand what git is for and "reflog" is not. Once you master a system and agree to all of its axioms and warts, it's all logical from the inside. That's true of almost any system, though.
There is a minimum level of knowledge you need to know to use git. It is not as straightforward as, say a simple text editor. The problem git is solving is more complex than that, and therefore understanding how to use it requires investing a significant amount of time in learning how it works.
If I was going to go back to my previous commit, I would use "git reset" to go back to the commit prior to the one I just committed. The only reason to use "git reflog" as part of that process is to see which commit was prior to the one you are working with now (in order to pass it to "git reset" to undo the changes).
> "Undo" is a reasonable choice if you barely understand what git is for and "reflog" is not.
But "git reflog" is not an undo mechanism. It can be used to determine a commit SHA to reset to (in order to 'undo' the last commit with "git reset"), but it is used for a bunch of other things also. It is appropriately named for what it does.
If you "barely understand what git is for" then the solution is to learn the minimum you need about git in order to use it effectively. I am not talking about mastering git, that's an entirely different topic. I'm talking about basic-level knowledge. Again, the required initial time investment for git is significant, but in my opinion it is worth it.
If you don't feel the investment is worth it, then you should use something else.
git reset HEAD^
Or if you also want to discard the changes you made to files in the repo at the same time: git reset HEAD^ --hard
But use the --hard option wisely. Be sure you really don't want to keep any of the changes you made (or that you have already saved them elsewhere before running it).Personally I use a UI for git which basically solves all of these problems. All the branches and commits are visible. If you want something somewhere, right click on it and you'll get all the available options. Nothing to memorize and everything is available!
My own favourite after testing out a few is SmartGit, but it's paid so not for everyone. There are lots of free git UIs out there as well, but what I like about SmartGit is that it's completely full featured - every obscure git command is available somewhere - so I have never ever need to use git command line when on my local machine, not even for these obscure things like resets, rebase, cherry pick, squashed commits, etc you name it. Also SmartGit is cross platform so I can use it anywhere
This is why it's good to put the linting in a git hook on commit or stage, so you literally can't commit without being formatted correctedly.
Also, unless you jump through a lot of hoops, linters generally run against what's on-disk instead of what's actually going into the commit. So checks that run at commit time discourage partial adds.
I still prefer to use Sublime Merge, I’ve tried almost every git gui out there and that’s the one that “stuck” and I rarely use the CLI. I’ve ever used and it takes so much pain out of coming up with all the crazy command lines and all the switches.
git branch NEWBRANCH
git checkout master
git revert —hard HEAD^
If you want to continue working on the new branch you do
git checkout NEWBRANCH
Source: git branch --help
git branch some-new-branch-name
git reset HEAD~ --hard
git checkout some-new-branch-name
Which is correct, assuming master is already checked out.0. prior state is that we're on master and have committed something that should have been on a branch
1. create a new branch that is identical to master (i.e., contains the commit) -- note this does NOT checkout the new branch (`git checkout -b some-new-branch-name` would do that)
2. reset current branch (master) to point at the commit before (i.e., strips the commit from master)
3. checkout the new branch to continue work there
At the end, master doesn't contain the commit anymore, and the new branch does. It's all correct.
https://docs.github.com/en/authentication/keeping-your-accou...
A guy in our team committed a big file unnecessarily into the repo which already had a year's history behind it, and it was only a month later when I found out. git rocket filter filtered it in a few seconds.
Since all team members had already gotten it, I ran it on all their repos to verify it's gone, and ran the usual git gc incantation [0] to clean their repo.
It's like a mechanic saying that a wrench is too complex.
Git is like CSS in the sense that for experienced users will sometimes forget the difficult learning process they went through, e.g. "oh use margin auto to centre the div"
I've been there, finding git too hard and giving up, but I've came back to it and eventually got past that uncomfortable learning bump.
Also, the fact that this is so complicated and that this document is as long as it is is strong evidence that git is just poorly designed from a UI standpoint.
Every Git novice should check out this awesome introductory video ;)
Thanks for posting it!
It is considerably more powerful than most teams need.
Take the following for example:
> Oh shit, I need to change the message on my last commit!
> git commit --amend
It's important to realize here that if you are simply trying to edit the last commit message, you *should not* have anything in your index (that is, staged). Otherwise those changes will be recorded in the amended commit! What Git does is essentially move all the changes recorded in the commit you are amending _into_ the index, and then run `git commit -m <amended-message>` ... so if you have files in there, those will get mixed up with the ones in the commit.
Here's another one:
> Oh shit, I accidentally committed to the wrong branch!
> A lot of people have suggested using `cherry-pick` for this situation too, so take your pick on whatever one makes the most sense to you!
Umm ... No! The solution proposed (with `git reset --soft`) and a cherry-pick are NOT the same! Not even close! You will produce two completely different histories.
This final one, given when this page was written, _may be_ understandably incorrect
> Oh shit, I need to undo my changes to a file!
> `git checkout [saved hash] -- path/to/file`
There is the introduction of a new command called `git-restore` (https://git-scm.com/docs/git-restore) that (thankfully) is named more appropriately—it "restores" a file. I wrote a thread on it on Twitter, so if you are curious perhaps this will help: https://twitter.com/looselytyped/status/1501934009370042371
*Shameless plug for my book*
My book, Head First Git, was published by O'Reilly this January. I posted a submission here on HN about it https://news.ycombinator.com/item?id=30072348 so if you want any details feel free to peruse that.
Some links:
- Amazon: https://www.amazon.com/Head-First-Git-Learners-Understanding...
- O'Reilly's online platform (Needs subscription): https://learning.oreilly.com/library/view/head-first-git/978...
- Companion website: https://i-love-git.com/
(Edited for formatting)
"people here" could be interpreted as "you people". Please remove it from your comment.
The name change was a complete non issue.
I’m frustrated if I take a day to work out some something that Jeff Dean mentioned to, uh, a friend, and it took me a day to work it out.
@bradfitz works on TailScale, there’s also a reason he’s a chapter in a book. @jwz is better, by so little you’d never notice, so his chapter is cooler.
What book are you in?
In a more measured tone: there is, in my opinion, a kind of creeping anti-intellectualism that’s been slowly-but-surely gaining ground on HN for the last 5-ish years.
We used to just really openly admire and respect iconic pros in this business, we used to openly acknowledge that some of the software work we talk about is pretty friggin elite and few of any of us will ever even work on some of it.
Whether it’s Linus or the TailScale pros or John Blow, I’ve seen people get gang-tackled by the “no one ever uses this LeetCode stuff in a real job” crowd repeatedly in the last week.
Knowing how to use “git reflog” to do surgery on a fucked-up repo isn’t necessary every day, but when you need it, you need it bad, and there’s nothing outdated about how knowing how the damned tools work.