Darcs, and pijul, are patch-based. That means that they think of the world as an ordering of patches. Patches aren't the same as commits: commit orderings, for example, are fixed, whereas patch orderings are computed. They can change when you e..g merge a "branch". Branching is similarly "simpler": a branch is just a collection of patches, not just a single commit with an implicit DAG attached to it.
Bitkeeper does have a weave data structure that more closely resembles patches. It's an encoded set of instructions for transforming one file from one state into another:
https://www.bitkeeper.org/src-notes/SCCSWEAVE.html
This data structure has a big advantage when computing annotations (blames): it's much faster than Mercurial's revlog (which in turn is faster than git's blob-tree-ref structure).
Is the difference between patches and git commits in a DAG really only a difference in internal representations or is there a user-facing difference?
Darcs' and pijuls' patches aren't glued, they only either commute or do not, and the conflict resolution mechanisms for non-commutative patches are different.
https://en.wikibooks.org/wiki/Understanding_Darcs/Patch_theo...
However, it would not be possible to implement a patch-based system (Pijul/darcs) based on git.
- One example is cherry-picking: in git, when you are on some branch A, and cherry-pick from another branch B, after the cherry-picking is done, if you try to cherry-pick from B again, you'll get conflicts. In Pijul and darcs, that comes for free.
- Another example is merge: merge between commits is provably wrong (https://tahoe-lafs.org/~zooko/badmerge/simple.html). With patches, this cannot happen.
> The difference between what svn does and what darcs does here is, contrary to popular belief, not that darcs makes a better or luckier guess as to where the line from c1 should go, but that darcs uses information that svn does not use -- namely the information contained in b1 -- to learn that the location has moved and precisely to where it has moved.
Darcs/Pijul understand which lines a patch is interested in, and if those lines get moved around in another branch, the patch "follows" them to the right place, rather than just finding/applying the shortest possible diff.
If you've heard someone say "never git pull, always git fetch and then either merge or rebase as appropriate", then they've noticed the difference between the two.
One confusing thing for git users is, git represents a number of commits (but not all types of commits) as patches.
Two differences:
- Cherry picking is possible, but when you cherry pick twice from the same branch, you get conflicts with commits, because cherry-picking change their identity. With patches, this works just as expected.
- Merging can be made associative with patches, not with commits. Concretely, in git, if Alice and Bob add lines to a file, even when there are no conflicts, Alice's new lines can be merged in the middle of parts added by Bob, even though she's never seen these parts. Even worse, there is no way to tell when this happens to you (git doesn't say). "Associativity" is the mathematical property that this never happens.
darcs had incredible cherry picking about ten years ago.
Instead of saying "get this commit" and solve merge conflicts manually, like git does, darcs would get one patch and every other that was necessary for it.
It effectively made cherry picking work as in "I want this feature from that branch" instead of "I want some code from that branch".
It was glorious, other than the little detail of occasionally exponential merge times...
The Darcs wiki has some cool graphics around cherry-picking merges that might answer your question better than a paragraph of prose can: http://darcs.net/Using/Model#merging-with-cherry-picking
The darcs/pijul world of repositories as loose ordered sets of patches makes cherry-picking the rule rather than the exception. A branch is mostly just the subset of patches you are interested in at a given moment. A "trunk" is just the superset of all possible patches. You can make interesting and easy usages of things like set intersections [1]: the intersection of the patches in two branches in darcs/pijul can be much more interesting than nearest common parent commit in the git DAG, and especially can be a lot more informative in the cases where things like bug fixes are cherry-picked across branches, which in git is a special bit of tree/commit surgery but in the darcs/pijul world that patch can be often the exact "same" in both branches.
[1] Aside, I love the concept of using intersection branches for consensus-oriented development (what releases to Production are the patches that every developer has pulled into their own working branch), which is a neat form of decentralized development that I think can only really be handled in the darcs/pijul model. (I have an ancient blog post on the idea of such a starfish development workflow.)
One thing you'll notice very quickly with Darcs (and presumably Pijul) is that the system always manages dependencies between patches. If you try to cherry-pick a single patch from a branch, you will get that patch and all the patches it depends on; you don't get the full linear history, you only get a subset.
In other words, you get the intuitive feeling that you're operating not on a log, but on a graph. Pulling one thread necessarily pulls other threads, and the whole graph rearranges itself to accomodate your changes. This has its downsides compared to the strictly-linear snapshot model, but the upside for most users is incredible. You can just commit and merge, and the system handles ordering for you.
Git was a major step down, UX-wise, when we switched from Darcs back in 2008, and it's still less user-friendly today. (It was also a major step up in some ways: Darcs, at the time, had a huge performance edge case where conflicts where sometimes effectively unresolvable because they took too much time to compute.)
Here is a very short video on a project called "Camp" that stalled out, nearly a decade ago. I think it very nicely explains how the user interface differs:
https://www.youtube.com/watch?v=iOGmwA5yBn0
The most important thing is that because there is no DAG, when you say "Darcs, pull this patch for me", like saying `git cherry-pick ABCDEF` -- the dependencies are automatically computed and pulled as well. You can sort of imagine it like if you had a git branch, and you ran 'cherry-pick' on one of the commits to your 'master' branch (because you wanted it). But, rather than pulling that one thing, 'cherry-pick' implicitly traversed the dependent patches and picks them all as well. But because there is no DAG, a dependency doesn't mean "parent commit". It means "the other patches that are mathematically required for this patch to work out". That means cherry-pick always works: you never have to calculate the dependencies yourself. To merge a patch is to implicitly merge all of its dependencies.
I've spent plenty of my time as an OSS maintainer dealing with merging multiple bug fixes from a development branch into stable branches. For example, a bug may already be fixed in HEAD when it's reported, but not STABLE, so you want to pull changes from HEAD into STABLE. Many times this requires multiple, carefully curated sequences of 'git cherry-pick' in order to correctly get the dependencies right. For example, the author may have made a small refactoring, then implemented the bugfix on top of that. Or it requires a complete reformulation or re-commit of a new change that matches the STABLE branch.
In a sense: this never happens with Darcs. If there's a bugfix, I say "Get me that bugfix patch". It always gets every dependent patch that is necessary, and never anything more. Every time. It always just works. Remember: no DAG. You aren't traversing parent commits. You are, in a sense, finding the transitive closure of "patches that cannot commute with this patch" (IIRC). That means: if a given patch does not commute with this patch, i.e. it is dependent, because we must apply them in a certain order, so there is a dependency -- then you also need that patch. And you need to apply that rule to that patch, and every patch it depends on, and so on and so forth (hence 'transitive closure')...
This allows a very powerful form of development, where features and bugfixes can coexist. But they do not necessarily need separate 'branches', so to speak. To merge a feature into a repository implicitly pulls its dependents, and the same with bugfixes. The net effect of this is that Darcs almost always gets merges correct, or it fails to do the merge at all. This kind of means that merges are sound ('kind of' because I don't know about an actual soundness proof, but the intuitive idea roughly is right): if Darcs pulls off the merge, then it's always correct, but it may not be able to always actually do that merge (perhaps not every merge is actually sensible, in the theoretical view of things, or perhaps the merge is sensible but the model doesn't allow it to handle that case).
The "fails to do so" is the tricky part, where Darcs 1 originally went exponential in some cases, though Darcs 2 mitigates this. It looks like Pijul will finally nail this problem dead, although admittedly I haven't looked over the theory.
Side note: Camp was originally envisioned to be the successor to Darcs, or at least the basis for "Darcs 3", using Coq to build formal proofs about the underlying patch theory to show it worked out correctly and avoided the harry bits that plagued Darcs 2. Unfortunately, it never panned out that way (due to time and lack of funding). The project was actually started by Ian Lynagh who worked at my current company before me and was one of the founders.
> Darcs almost always gets merges correct, or it fails to do the merge at all.
One of these is not like the other, which IMO is the problem with "magical" merging systems. Great when they work, f*cking hell nightmare when they don't.
I'd rather have something like git that works in normal usage all the time, and when it fails, is easy to fix. YMMV.
In contrast, Darcs and Pijul's merge are associative, and Pijul's merge is commutative. Even if you don't like maths, this means that they will always behave deterministically. This also means you can use them in scripts, although darcs might sometimes have performance problems (pretty bad ones, actually).
In git, you can get the following: https://tahoe-lafs.org/~zooko/badmerge/simple.html
In all honesty, given years of experience with Git, and fondly using Darcs as my first version control system: I still think merges are absolutely the one thing it beats Git at, hands down. When it works and it does its job, it always is correct. When it doesn't, you can bail it out. Not much different, but the "always is correct" and dependencies-being-implicit is what makes it good. Darcs could have saved me at least dozens of hours of hair pulling when doing STABLE merges I estimate... Git's still good. I wish it could do that, though...
Your note about git is interesting. In fact, Git is, in at least some cases, more magical than other VCSs in the merge department. You might just not be aware of it due to being so familiar. When I say "Darcs always gets the merge correct", I don't just mean it literally finishes with exit code 0, but also that the semantic model is, in some sense, more 'correct' or 'intuitive':
http://r6.ca/blog/20110416T204742Z.html
Darcs (and others) always get this 'merge associativity' case correct, where 'Base+A+B' where (+) is merge is associative (so it doesn't matter how you 'bundle' the changes or whatever). That means you have less edges to worry about. And to be fair, I don't think there's anything inherent about Git where this particular case can't be fixed. It's just a good example of why people are trying projects like Pijul/Darcs at all, so these things can be formalized and understood. The theory of patches is actually rather rich and helps formalize a lot of these notions of what a "merge" really is in an algebraic sense, how patches relate to one another, etc.
When you type darcs pull, darcs lets you pick and choose which patches (subject to the dependency constraints) you want to pull in. Those patches then get applied however darcs wants to apply them (again, subject to dependency constraints, of cousre). Because you do not necessarily need to pull all of them, you are always "cherry picking" by default.
This is different from cherry-picking in git, because when you cherry-pick in git you still have 2+ commits that exist in the context of a DAG; you're just transplanting the contents elsewhere to create a new commit in a different position in the DAG.
This sounds less and less like a tools/implementation thing and more like the default recommended/enforced workflow thing.
It might help to compare the "identity" structures of git commits versus darcs/pijul patches. In a pseudo-C, you can see a git commit as something like:
struct commit {
string author;
string description;
tree_id tree_snapshot;
commit_id[] parent_commits;
}
If any of those fields change (are amended), you have a new commit.This is a directed acyclic graph (DAG) because of that `parent_commits` link from one commit to the immediate previous parents (it can be multiple parents in the case of a merge commit). Git can only just move refs on a pull in the case of a "fast forward" when the remote branch is "simply" ahead of the current branch and all of its new commits "point to" the last commit in the current branch. Every other case it is a merge of the graph (via a merge commit with two or more parent commits).
(While git outputs a diff as the representation of the commit in places like `git show`, a commit doesn't store the diff but instead a link to a snapshot of the tree at the time of the commit.)
For something like darcs/pijul, the identifying information of a patch looks something more like:
struct patch {
string author;
string name;
change[] changes;
}
If any of these things are changed (amended) you have a different patch.This may seem like semantic quibbling in that the patch here actually contains the diffs as a part of its identity rather than a snapshot of a source tree, but that's not actually the important difference.
The important difference is that the context of the patch is no longer a part of its identity: there is no "parent patch" information, and the change structures don't directly refer to previous changes.
The reason that difference matters is because in the darcs/pijul models the context of the patch is more "metadata" about the patch than a direct part of the patch. Patches aren't "nailed" to a graph like a commit is, they "float in a basket" together. Darcs and pijul do the work to figure out which patches need to be in which order in a branch/repository.
This can be a nightmare to someone expecting a strict graph. Darcs and pijul can and will reorder history during a pull. You can see "newer" patches float down under "older" patches in the patch log as the systems work to build a stable sort of patches.
That movement, however, is also where the systems draw the most strength. That movement of the patches can be seen as a continual, rustling "cherry arranging" as the systems work to figure out the minimal set of previous changes that patch needs in order to exist.
If you cherry-pick a commit in git you copy the changes from that commit to the new branch (a new spot in the DAG) into a new commit with its own new identity. Down the line when you go to reintegrate/remerge the branches between the original branch and the cherry picked branch, git doesn't see the same commit/change and its merge can (in my experience, will) see conflicts in the exact same change made in different contexts.
When you cherry pick a darcs/pijul patch, you bring over the same exact patch and the system lets you know any other minimal dependencies that you need and brings them over as well. When you reintegrate/remerge the cherry picked branch, those exact same cherry-picked patches are already "in" the original branch and so don't necessarily need to be remerged/rearranged again.
You can duplicate git workflows on top of darcs/pijul, but it is very hard to duplicate some of the more interesting darcs/pijul workflows on top of git. Among other things, rebase/cherry-picking merge hell is a very real problem in the git ecosystem, whereas darcs/pijul almost seem like crazy smart magic in comparison when it comes to some of the scenarios where you might rebase or cherry-pick.
It might be something that won't entirely make sense until you try experimenting with it yourself: maybe, you might want to take darcs for a spin for a small project or two. I think you can feel a lot of the difference as you use it, especially as you start to push/pull between branches/repositories.
(Anecdotally, my workflows are quite different on darcs versus git, knowing that typically I could fix a bug discovered elsewhere in the code in the middle of a bigger project, without needing to branch, I would often just record that change into a tiny patch on its own right there on the spot, and generally know that if I needed to get just that one patch into another branch I could rely on darcs to cherry pick it for me later.)
Down the line when you go to reintegrate/remerge the
branches between the original branch and the cherry picked
branch, git doesn't see the same commit/change and its merge
can (in my experience, will) see conflicts in the exact same
change made in different contexts.
Finally I see some common ground: by tracking changes separately to commits, users won't see merge conflicts when the commits get moved.Sadly, git also recognises this on the user's behalf, so likely your experience was due to some other delightful quirk of the git UI.
edit: I'd also recommend not calling trees 'tree snapshots', because that will confuse people familiar with trees. Same for 'cherry-picking a commit', since from your description darcs seems to use 'cherry-picking' to mean 'fetch a set of changes from someone else', which maps to 'fetch a branch' in git-land. 'git cherry-pick' means, 'copy a single change from one local branch to another', so has almost no overlap.
You seem to think its "fetch a branch", but I'm trying to tell you that the `darcs pull` experience is a lot more like doing `git fetch && git cherry-pick origin/TIP --interactive` every time than `git pull`, but with a much, much better merge experience than that implies.
I'm not trying to confuse different concepts, I'm trying to show that the hard concept in the git case was the easy concept in the darcs case.
My point is, if you focus on the implementation differences, my understanding won't increase because we don't have the same mental model of how git works under the hood.
Can you explain a little more what you mean?
Even when using git am or send-email or whatever, yes, you're sending a patch -- but the way to apply that patch is to turn it into a commit and then cherry pick or rebase or merge or manually fix conflicts or whatever. In darcs and pijul, the model is _always_ that set of patches.
But, you are correct. Internally, git stores the full contents of files and computes the diffs on the fly.
For others' benefit, if you want to test for yourself, create a new repo, add a file and make a series of commits with changes. Git objects are compressed with DEFLATE (zlib) so gunzip and unzip won't work. I used https://github.com/jezell/zlibber because I was too lazy to write my own quick zlib wrapper. Then doing
for o in .git/objects/*/*; do cat "$o" | inflate ; echo ""; done
lists the contents of all the git objects. Notice that there are no changesets, only full copies of the file you modified at different states.This was surprising to me, since I had a very different model mentally. I still think the DAG of diffs is the better model mentally, but it is worth understanding that this is not what git is actually doing under the hood. It explains issues that arise doing rebases, cherry-picks, etc.
I now also understand the motivation behind Pijul. If I understand correctly, Pijul does use a collection of changesets as the underlying model. Like you say, that can be a critical difference.
But if you don't have A in the first place, then there's a real problem, how to take B-A and reconstruct B, without knowing what A is. You have to find A, and there may be multiple acceptable A's. It seems like the fundamental difference here would be not having a strict "parent" for any given patch. I can see why that would make some workflows a little nicer, not being forced to rebase, but I don't see any massive advantages -- as a user what does this really buy me? Does it enable some things like that are impossible with git? Or does it mainly make some advanced git workflows easier?
One difference is that by storing the patches only you can understand more clearly what the intended change was. When you store the whole file it is easy to compute the difference between A and B, but may be impossible to compute the correct differences between A, B, and C. By storing the whole file you now have to consider all the possible differences between them, not just the ones introduced by the commits you are trying to merge.
I would have to play around with it, but I know there are scenarios involving rebase, revert, and cherry-picking commits that can cause trouble in git that I now understand comes because of the fact that git is storing contents, not diffs.
One that I've run into regularly is cherry-picking commits from a dev branch into a master branch to hot-fix bug fixes directly into a prod release instead of waiting until dev gets merged as part of our regular process. If I had commit A on dev and cherry-pick it to master it creates a totally new commit A-1 that becomes part of the history of master. We lost the fact that A and A-1 represent the exact same changeset. Depending on what the changes are, and what further changes happen on dev afterwards, this can cause failed merges requiring manual resolution when dev does finally get merged into master.
I imagine that would not be a problem for Pijul.
Git certainly has some room to improve in the merge conflict department. I looked at the bad merge example you posted -- I suspect I've hit that before. It's rare, but yeah it's there.
I also frequently notice that git complains about merge conflicts, while the custom diff tool I use to resolve them says there's no conflict, and I don't actually have to do anything. Good reason to use a custom merge tool with git.
But, given all this, is this really all an outcome of patches vs snapshots, or is this just git's merge algorithm being suboptimal? Certainly git could selectively ignore the DAG when merging, couldn't it? Even after reading the other comments here, it still seems to me like git has more information when merging than the "patch-based" workflow of darcs & Pijul.
It seems to me like there's a language problem with trying to draw a distinction between patches and snapshots. Git is still storing and transferring patches at the tree level, even if it's not happening at the file level. Git does not store a commit as a zip snapshot of the entire tree, the commit is still only the changed files. It would be fair (but not standard or common) to call the overlay of changed files a "patch" or a "diff". People do still use git format-patch, and email git "patches" to each other. So it's inherently confusing & problematic to talk about git and say that it doesn't use patches.
What does make sense to me is the distinction of having a strict DAG vs not having one -- is that actually what people mean when they talk about snapshots vs patches? Am I tripping on it because I'm being too pedantic about what a "patch" is?
merging, rebasing, committing... all operate on refs. You might think you're transplanting changes (and you are), but the inputs are refs, and the outputs are refs. As you mentioned, refs are unambiguously snapshots.
If anyone is questioning this, just play with `git rebase -i [some old changeset id]` and you'll see that it's just an ordering of patches.
In fact, this is why rebase is an out-of-band tool that has odd effects on shared history - specifically because it's inverting git's model into something more like pijul's, and therefore isn't really native-to-git.
Assume you change file A, commit, then change file B, and commit. In git there is a dependency from the second to the first commit, because the state of the second is derived from the first one. However, the changes are unrelated because they are in different files. Pijul understands them as parallel unrelated changes (although with different time stamps).
(disclaimer: I infer that from using darcs many years ago. I never used Pijul)
In theory, the number of possible branches is much, much greater (arguably infinite) for git. Commits have their own metadata, so just amending creates a new commit. Patches are immutable, and there's no independent artifact like a commit that incorporates the current set of patches.
Furthermore, darcs (and I assume pijul) absolutely let you make multi-file patches.
Apart from the other answers that go into theories of patches vs. commits, I always found Darcs much more intuitive to actually use than Git. It has fewer commands that do more intuitive stuff. To record a patch, you do "record" instead of separate "add" and "commit" steps. To revert some changes you have not recorded and that you want to get rid of, you do "revert" instead of "checkout -- filename". The "diff" command works more intuitively than git's "diff", which you sometimes have to use as "diff --cached" due to its staging model.
Another thing is that every branch is a separate copy of the entire source tree in your file system. A drawback is that this can be viewed as wasteful, but the advantage is that it's much much easier to work in parallel branches at the same time since changing the branch is just doing "cd" instead of some boring dance of "stash" and "checkout". You also don't get problems due to files intended for one branch still lying around after a "checkout" to switch to another branch.
(I'm assuming that Pijul preserves all or at least most of these properties of Darcs. The docs aren't very exhaustive.)