It uses snapshots because that is efficient, but from a user perspective all common git operations look more like they are operating on diffs than snapshots.
When you cherry-pick a commit, the diff of the old and new commit will be similar, but the snapshots are normally completely different; that's the point. The same goes for rebasing, which is like applying a set of patches. Heck, if it goes wrong, you get merge conflicts, which wouldn't happen if you were just manipulating pointers between snapshots.
Git actually makes it quite difficult to manipulate commits as pointers to snapshots instead of diffs -- i doubt most git users even know about `git commit-tree`.
This is why i don't really get the "git is wrong and should work on diffs" crowd. If it did, the user experience would be 99% the same, unless you're manually editing your diffs before committing them.
Git does delta compression, so in fact it does usually store diffs on disk. That, however, is a technical detail that doesn't actually influence the user and can be 100% ignored as long as you don't mess with its internal files by hand.
What the user actually operates on in the repository are snapshots, with diffs being merely an intermediate representation useful for factoring, reading or distributing stuff. Git is good at confusing the user that it works on diffs, but the sooner you realize that it's not true, the easier it will be for you to work with Git.
And while "git cherry-pick" (and in turn, "git rebase" too) seems like an automated "git format-patch + git am" at first glance, it actually goes further and is using three-way merge for better conflict resolution. It works on snapshots, not diffs.
> Git does delta compression, so in fact it does usually store diffs on disk.
Yep. The comment you replied to originally was pointing out that snapshots and diffs are isomorphic, so i'm glad we all seem to agree.
> What the user actually operates on in the repository are snapshots
They don't, though. Git doesn't show you the tree IDs in normal operation, and you can't actually make commits that point to specific trees (snapshots) without unusual commands.
> it actually goes further and is using three-way merge for better conflict resolution. It works on snapshots, not diffs
It doesn't really matter what it's working on, the best mental model to understand cherry-picking and rebase is that of applying diffs. Can you (in general) even explain things like rebase and cherry-pick without the terminology of diffs? The git manual doesn't bother.
Didn't want to, consider "you" to be plural in my last comment, or replace it with "one". I honestly believe that reasoning about commits as "diffs" leads to nothing but confusion.
> They don't, though.
They do. The fact that to write a letter you type each character separately doesn't mean that what you're operating on in a text editor are one-char diffs. You're composing a single letter to save - just like in Git, where you're composing a single state of the repo to then commit (or in other words, to snapshot it). How exactly you compose that state (by using index, or commit-tree, or subtrees, or placing files directly in .git, or...) is irrelevant to the resulting repository graph - and that graph of snapshots is ultimately the data structure that you're conceptually operating on (regardless of how it's represented on the disk).
You work on commits, not trees. Commits are snapshots of your files. Trees and blobs are just how Git represents your files - almost an implementation detail. From the user PoV, Git could even be creating new directories in .git with copies of your whole working dir for each commit and nothing would change conceptually, it's irrelevant to the high-level mental model of a Git repository.
> Can you (in general) even explain things like rebase and cherry-pick without the terminology of diffs?
A diff is a result of an operation applied to two snapshots. Cherry-pick executes that operation and uses the result of it (at least conceptually). You can't think of it as operating directly on "commits as diffs", because then things like cherry-picking a merge commit wouldn't make any sense, while they still make perfect sense and are easy to explain with the "commit is a snapshot" mental model. Some things in Git calculate diffs between two commits and use that result in some way, but that doesn't change the model of the repository.
And because Git often shows you diffs for convenience, it's easy to develop a wrong mental model of the repository - a model that most people operate on, but which will bite you sooner or later. That's exactly why so many people end up being confused with Git.
IMO it makes just as much sense either way. When cherry-picking a merge commit you have to specify which parent to diff against, which could just as easily be explained in terms of diffs (i.e. which part of an n-way diff to apply).
> it's easy to develop a wrong mental model of the repository - a model that most people operate on, but which will bite you sooner or later. That's exactly why so many people end up being confused with Git
This doesn't match my experience (as "that guy that people go to for git help"). Perhaps you have a concrete example.
Just to be clear, i'm not advocating for teaching or believing that git works in a way that it doesn't, that would be silly. More that being able to think about it in different (equivalent) ways in different situations is helpful.
- cherry-pick creates "duplicated" commits
- you can checkout a commit, rather than a branch
- two commits in the same repo may be topologically unrelated to each other
- merge commit can contain changes unrelated to its parents (usually after making some by accident)
- you can git reset --soft
Those are just the ones that came to my head on a whim (and don't get me started on rebase-heavy workflows). I've had people telling me that it "finally clicked" when made aware that commit is a state rather than a change. Explaining what I just did to "fix" someone's repo is also often easier after a proper "commit as a snapshot" prelude. If people who used Git for years say "oh, that makes much more sense now" after being presented with basic Git concepts, it suggests that their mental model may have been somewhat flawed.
Once you internalize that commit is a snapshot, going from that to thinking about diffs between two commits is easy. The other way around is not so obvious - when checking out a commit from complicated topology, it may not be immediately clear how to get from one state to another by "applying diffs" even if it's technically equivalent, so simple operations will end up seeming like undecipherable magic to you simply because they don't fit your mental model very well.
The whole thread started with a notion of "particle-wave duality", which is technically true, but useless in practice when "diff between two arbitrary commits" is one of the simplest operations to think about in terms of graphs of snapshots.