Git is simply too hard
changelog.com
changelog.com
Like, the question: "How do I undo 'git add'" shouldn't have over ELEVEN THOUSAND upvotes on StackOverflow: https://stackoverflow.com/questions/348170/how-do-i-undo-git...
That's just bullshit
If the last thing you did was stage, do you want to unstage, or undo the last change outside the staging area?
If undoing the last thing you did is not compatible with the current state of the working copy, what happens?
If your last change was to discard or apply uncommitted changes, how can that be undone? If the answer is "do lots of hidden commits", how does that interact with undo?
I totally understand why those were chosen when the original code for git was written but I really think it's time to re-evaluate the CLI and do better.
Git should eliminate the following redundant naming: stage, cache, index all refer to the same thing.
If I want to diff between the working copy and the stage, why isn't it "git diff --stage", but "git dif --cached"?
A cache is something for fast lookup; the index isn't a cache; it's the permanent record of stuff in version control that will become a commit.
When you invalidate a cache entry, the semantics is unaffected, only speed of access. Removing from a git index is semantically significant.
Also I believe "git stage" exists too: https://git-scm.com/docs/git-stage
Because git commands are used in a bajillion scripts, it's very hard to remove "excessive" functionality from the old commands even when that would improve the orthogonality of the tools.
- Git stores snapshots, not changes.
- Snapshots go into something called an index. "Git add" means to add to an index.
- A new commit is prepared by creating a new index based on the index of the HEAD commit. This new index is the "stage". The index holds a complete copy of the files. When you stage only certain changes with "git add --patch", that is still the case: you're making a version of the file which has only those selected changes and that goes into the index. A built-in tool that is external to the git representation did the text manipulation to separate those changes.
- When you commit, the uncommitted index you have created as a stage becomes a commit.
"Consider the expression 1/3 + 25. In PL/I this expression has the value 5.33333333333. Why? One-third is computed to 15 digits of precision, 14 to the right of the decimal point. Then 25 is coerced to the same precision, losing the most significant digit 2! This does raise an error in PL/I, but the default is to ignore it."
When people ask me seemingly simple questions like "how do I delete a branch", I feel like I have to give them a lecture on the philosophy of git, which I then decide not to give, because it's going to make me sound like a smug neckbeard. However, the fact that such a lecture would actually be needed, is just a sign that centralized software has taken over the public mindset. Actually, the situation is sort of similar to when you try to explain people how to use IRC or Mastodon.
Simply not delivering said lectures has it's own downsides but seems to be the best strategy if you care primarily about how well liked you are.
Mind you this isn't everyone, it's maybe 50% of people and certain personality types really don't like receiving information in that format.
There are also those that think the same way and reciprocate in kind when you have a knowledge gap, I really appreciate those people. They also are better at communicating what they actually do know so you can reach a shared understanding with them faster.
Ergo git may be hard to it's clearly not hard enough for market forces to displace it.
What is your point exactly?
Nowadays even companies (that would in the past self-host) rely on SaaS and are thus swayed by network effects of whatever is popular.
I am still using svn for some projects (due to large file/subdirectory handling) and this is a pain every single day.
Easier alternatives to git already exist. However for a confluence of other factors (network effects/inertia, etc) git remains dominant.
All this means is git isn't "too hard". It's hard but that hardness isn't great enough to invite a competitor to displace it.
Replacing git won't come through offering something easier, likely some other major paradigm shift will be necessary.
There were two main factors contributing to it: Linux Kernel switching to Git, so open source people had to deal with it anyway, and an even bigger factor, GitHub offering a (free) streamlined UI for those most common GitHub operations (basically, making it not-hard). If you compare GitHub to Mercurial hosting web sites of the time, one was decidedly 90s/MySpace like, and another was modern, Facebook-like.
IMHO, without GitHub, git does not become the dominant VCS it is today.
Eg. one of the obvious missing features (for a distributed VCS) is nested history (i.e. in Bazaar/bzr, now brz, you'd keep "subcommits" when merging, without them overcrowding the main "bzr log").
Basically, while advertised as being distributed, you are forced to either rebase or squash-and-merge to keep a clean commit history, which is not how humans operate: basically, I want to have a gazillion tiny commits if it's a distributed system (iow, use it how I like it), but not have those show up for a single top-level change.
Also, internals do drive what UX you can have: this is evident even in web applications (eg. if you've got a particular backend REST API, you might not be able to achieve some UX workflows without reworking the backend too).
Good thing neither "pretty powerful" nor "incredibly powerful" have formal definitions ;)
> Also, internals do drive what UX you can have
Yes, but that's not what I'm talking about. Git just has bad UX. There's no reason why creating a branch can't be part of `git branch` rather than `git checkout`, but that was an opinionated decision that every user just has to deal with now. Or, why does `git stash pop` immediately stage changes if there's a conflict, but not otherwise? Yes, there's always a reason behind these choices, but they prioritize some concept of "purity" over UX. The result is a tool that has spawned a thousand "learn git" tutorials because it's so user-hostile.
The best way I learned git was by browsing the files and contents of those files in .git/. That way I had an idea of what's possible and what I wanted to do and could find the command to do it.
How many git commands do I commonly use?
status, log, whatchanged, diff, diff --cached
add, reset HEAD, reset --hard, checkout, branch, commit, commit --amend
fetch, push, push --force-with-lease
stash, stash pop, rebase -i, reflog (escape hatch)
I hear similar difficulties from some who learn databases from using ORM syntax and don't have a mental model that matches how relational sets compose or how SQL is primarily declarative rather than imperative.But, it does feel like a big upgrade over the past iterations of version control like svn.
IMO the git model is an incredibly high power-to-weight ratio for what professional software engineers need, I just can’t get behind efforts to water it down for easier onboarding. No one would suggest professional chefs replace their knives with preformed vegetable slicers from late night TV.
GitHub is just another git remote, does not centralise anything, and causes no impedance mismatches.
This. If new users had the git model itself explained to them, rather than tell them to do their work and learn git as they go (giving them blind commands to type here and there), things would be much better off.
It's even more complicated now with github and IDEs having a UI that abstracts over git's abstractions.
Undo the change: "The recommended approach is to make another change which undoes the original one and push it, using command AAA. If you want to un-push the change you made and pretend it never happened we can try it as well, but Git does not really expect published history to change, so but if someone else on your team has already pulled it, their machine will get pretty confused. It's like deleting email right from someone's inbox -- OK if they haven't seen it yet, but pretty confusing if they read it and already typing the reply"
Delete the branch: "Does that branch exist on a server?". Depending on the answer, you either delete remote branch + local tracking one, or just a local one.
The key idea is to have a workflow in mind, and steer users towards that workflow. Because sure, you can ask user if they want to "delete a local branch, and then just get it recreated when you pull again".. but that's terrible idea. Why would you want to do this _ever_ as part of any regular workflow?
There are certainly circumstances when you have to do advanced/unusual stuff, but there is no need to push that to novice users, unless you hate git and your goal is convince them that the git is hard.
> The key idea is to have a workflow in mind, and steer users towards that workflow. Because sure, you can ask user if they want to "delete a local branch, and then just get it recreated when you pull again".. but that's terrible idea. Why would you want to do this _ever_ as part of any regular workflow?
I've recently tried to guide someone to fix their bad branching strategy, and after a few of their bad attempts, just suggesting they remove their local branch and then refetch it from the published repo instead (just imagine how easy it is to figure out what commands-people-have-blindly-typed-in-from-the-internet have done, and then reverting that with git).
Going back to a known good state is the most common of all debugging tools (also part of more advanced tricks like bug bisection search — interestingly also present as `git bisect`), and with it being too easy to mess up your local repo, going back to a remote that's "clean" is sometimes the fastest approach.
That's a weird statement for a tool which has had `rebase` as the very basic command from the get go.
Git does expect history to change, it's publicised as a solution to a bunch of issues: basically, there is no getting around understanding the Git internals, and even once you do, it's hard to remember how GIT cli commands tie into those internals for things you seldom use because of their arcane naming.
> @wilshipley git gets easier once you get the basic idea that branches are homeomorphic endofunctors mapping submanifolds of a Hilbert space.
https://news.ycombinator.com/item?id=36179021
https://news.ycombinator.com/item?id=36178952
People in general are cultishly attracted towards instantiating centralized solutions to solve their problems. Simultaneously, they're blind to the long term negative effects that come from creating it, mainly:
- the developed overdependence on it for a particular service/function
- the societal mental lobotomization in terms of developing non-centralized solutions
- the incoming naivete of the next generation into assuming that it'll always exist & be there, and not planning for the worst case scenario
Referring back to Git, part of the reason could be that they assume a canonical "true" repository, which runs counter to how Git operates.
--------
Nevertheless, Git's UI could be cleaned up significantly: VS Code wise, Git Graph is a personal favorite of mine because of how well the plugin renders the commit graph for you.
https://marketplace.visualstudio.com/items?itemName=mhutchie...
The author mentions this near the bottom but there's already some great porcelains for git.
https://learngitbranching.js.org
Few more listed here:
[1]: https://ohmygit.org/
- not only saving changes to files as an ordered log, but being able to reorder those changes as if you'd made them at a different time (git rebase)
- merging multiple, conflicting changes to files made by different people (git merge)
- homing in on a specific change in which a bug or a feature was introduced (git bisect)
It's not unreasonable that such a powerful tool might come with a cost. Considering the time that it takes in other professions to learn the tools of the trade, I consider Git to be a pretty good deal!
Personally, I found my understanding of and proficiency in using Git improved dramatically when I decided to resist the temptation to delete the repository and start again whenever I had a 'merge conflict', a 'missing remote' or a 'detached head' (yikes). Instead, I tried to work through the problem step-by-step, learning about such things as ref logs and branch tracking and the distinction between the work tree and the index. It was fun, informative and didn't even take that long in my opinion :)
I think that for such a widely-applicable task as distributed version control, Git is a fantastic piece of software, and would recommend every software developer take the time to learn it, keeping in mind that only a few centuries ago, a stone-mason might spend years learning to use a chisel, or just in the last century learning how to use a powered lathe. I see Git as the programming equivalent of the lathe or chisel - equally fundamental to the trade.
For those who aren't programmers (and thus don't feel the need to study Git in such depth) but still want to make use of version control, git-annex[1] provides one nice way to put binary files like MS Word documents into a repository and synchronise them across devices.
All activities involving diffing, merging and conflicts and such are operations built on snapshots; they are external to the core git representation.
But it's not really about the number of subcommands, but rather how complex some of these subcommands are. e.g. git-commit(1) is 655 lines (and references other pages, too). hg's "help commit" is 59 lines, or 98 with --verbose (this also references some other pages, so it's not all you need, but it's indicative of the relative complexity).
> Multiple checkout in general is still experimental, and the support for submodules is incomplete. It is NOT recommended to make multiple checkouts of a superproject.
I'm not quite sure what is safe and what is risky.
Also, let me tell you from first hand experience: I disagree with you. Let me tell you a few things about myself, just to paint a picture... I managed to gain admittance to a special math high school class with like ten kids fighting for every seat. Later I got a math teacher masters. I have been a member of Mensa for a few years. I learned Z80 assembly when I was 12 years old. Prolog when I was 15. When git came ... it was an absolute impenetrable brick wall and let me emphasize, a directed acyclic graph is not exactly an alien concept to me. It was not until Charles Duan's git tutorial that I felt I am comfortable with it. https://www.cduan.com/technical/git/ All this to say: it is hard.
Auxiliary? I can’t speak for everyone, but VCS has to be up there as the most important development tool on my machine. I’d sooner uninstall my IDE than git.
No doubt it’s hard, but sitting down and making a concerted effort for a few hours does wonders. That’s not much of an ask for something most devs use every day.
> My brother in Christ
disingenuous start, this will be interesting...
> why should I put in the time and effort to learn an auxiliary tool to my job?
if you think git (or version control) is auxiliary tool then you have not worked on real software. I'm not saying this is you, but IF it's you (or anyone else reading): Writing side scripts and minor modifications on an existing project is not building real software, it's hobby coding, which is fine but you don't get to act like a software engineer. You're as much an engineer as someone cooking a random meal once a week is a chef.
> I learned Z80 assembly when I was 12 years old. Prolog when I was 15. When git came ... it was an absolute impenetrable brick wall and let me emphasize
Did you consider maybe that your "learning" of z80 asm or prolog was not complete? or possibly faux, false, fake, or gasp you're might not be as smart as you think? Or that when you were a kid you had way more time on your hands than now? The possibilities are infinite for an explanation (that is if what you're saying is true). "OMG Git is hard" is not one of the first 100 reasons that come to mind.
> It was not until Charles Duan's git tutorial that I felt I am comfortable with it
That's a good tutorial, but you never get comfortable with Git (or any other technical subject matter imho) until you work with it, and it will take practice and -sorry to say- humbleness in that you (or anyone else) can have hard time understanding something that many others find simple. It doesn't automatically mean it's a hard subject because you find it hard.
Git is _not_ anymore hard that any other abstraction. You work hard you learn, you don't you won't.
"whatever comes next should be closer to how humans think"
Hard no. I saw how people think and it is usually: this is what I think I want, but I ask for not that, but figure it out for me anyway.
I will admit that articles like that make me wonder if Altman is onto something saying that AI will end us all.
Waiting for the next jump forward. Please.
I would recommend the comment by edblarnes:
> Git is a brilliant data model that we hack into being a code repo. > It was designed from the model backwards, not the use cases inwards
Git can do a million things no one uses it for. It practice, teams develop and teach a happy-path workflow, and then for other things they drag in the git expert. But users do random walks through git operations, wasting everyone's time.
Design-wise, it's 80/20 again: most teams mostly work with gitflow-style workflows, with various rules -- AND most have some tiny variations they work in, so there's minimal capture of workflow.
As an aside, my gripe with git is text diffs and commit histories. It's simply too hard to understand changes and to review code when you have no control over the granularity of the change and are forced into per-commit granularity of comments (and even those are lost in squash).
At the very same time that we want to increase code velocity, teams have to impose e.g., 500-line limits on diffs to keep code reviews manageable. This slows change to tiny incremental steps. But they're only undigestible because the actual change is lost in the details, and commit comments have to decide between high-level changeset overviews or low-level file change descriptions. We need to zoom in and out to different levels of comments and changes.
IBM's 1990's (closed-source) Visual Age saved code as an AST, and could do great diffs. But now with LLVM and LSP, we're much, much better at sharing underlying infrastructure of languages.
My minimal rule for next-git would be that any automatable refactoring should appear as the one tiny change it is (renames, extracting methods, hoisting members in the hierarchy, etc.) in any language.
The reason we don't is not technical, because logical diffs are achievable. It's the coordination costs of shifting everyone from git to something else.
Given that, one option now is to support text conventions for commit comments that bridge comment levels and incorporate AST-level descriptors. They might be readable by humans, but most humans would use tooling to fly over and dive in.
The path to designing that is to work out the main and troubling use-cases, propose some formats, build in evolution/compatibility, and make some prototypes.
If people want to make a name for themselves, this is a way to do it that could transform the industry.
If a company wanted to make code understandable to AI, this is the transformative way to do it. If you squint, metadata over change is a model of intelligence. There's a lot of headroom in the space of building systems to feed AI instead of feeding human garbage into AI at sufficient volume to compensate for noise.
Keep text, as it is most universal to consume, and generate AST on the fly. Or serialize ASTs to xml and commit them. git is certainly flexible enough to work with non-text formats (a random example from internet: https://www.diffplug.com/features/git )
In particular, "any automatable refactoring should appear as the one tiny change it is" is trivial to add from the git side, just set "diffcmd" to whatever script does the logic. The hard part is actual logic to compare ASTs and derive if this is a "trivial rename" or "trivial rename except usage in file X is unchanged" or "trivial rename but I also made everything const".
Case in point, we use a 3rd party review tool at work (Reviewable) which overrides default github pull request diff screen and shows its own diff in a different format.
overcome it. if you can beat git, you can beat anything.
perhaps you’ll run into aws next. like git, it grew organically, has way too many knobs, and is easy to screw up. like git, it’s the best we’ve got.
as with all things, experiment in a safe context where failure is ok. learn which knobs are important and tape over the rest.
practice until you can do it asleep or drunk or both. maybe write some bash functions so you can forget some of what you learned if it doesn’t matter.
tldr; git good and have fun doing it!