What does Pijul do that Git can't? What does it do better?
What does Pijul do that Git can't? What does it do better?
Maybe you may also want your tool to serve your workflows and not the opposite.
This sounds like you claim that pijul serves all workflows well, which would contradict my own experience. I'm so used to think in branches / PRs (to display a history and showing lineages), that I find the current way of handling channels/patches of pijul alien. Right now I would need to break my workflow in my daily operation to use pijul.
Anyway, breaking the workflow when switching tools is generally expected. It's a moment to review your assertions.
I recently merged two functions from two different branches. It interleaved the lines because the functions were similar. Tooling can help this problem of course, also choosing different merge strategies i think is an option, but that in think is what Pijul is attempting to solve. To reduce conflicts when slicing and dicing past histories.
Git views history as a series of (notionally) atomic snapshots; programmers tend to view history as a series of diffs; Git tries to supply a compatibility layer to better support how we naturally think of history, but it's such a leaky abstraction.
Git's "problem" is that you can get identical content via different routes, but due to the way git fundamentally works, history is part of the snapshot and two snapshots with identical contents but different ancestry are not the same. Actually implementing a system that handles this properly is far from trivial, and git makes the tradeoff in favour of implementation simplicity.
The problem come when you merge and rebase your snapshots, possibly solving conflicts in the process. Then none of this snapshot thinking makes sense.
And actually, Git knows that well, since its default merge algorithm diffs the tips of branches with the youngest common ancestor. And rebase "replays" diffs (how would you "replay" snapshots?).
So I claim that very small changes are obviously thought of as diffs. And what is a large change but a composition of small changes? Pijul's model naturally represents the process of creating a change - you could slice up the diff as small as you liked without anything about the mental model changing, down to individual keystrokes if need be - whereas Git's model gets more and more unnatural when you do that.
GP cites rerere as evidence it is, and I'm inclined to agree - at least when I'm merging. During a few operations like a bisect, I probably view history as snapshots. But during merges and rebases, I definitely view it as diffs; I question whether someone can even explain these operations abstractly without resorting to a diff-based explanation. And I merge/rebase 100x more than I bisect.
E.g. if you merge two branches with the same patches but in different order, but the file contents are the same at the end, git will need to you to manually address the merges for each patch. In a patch based system, there is no conflict if two patches can be reordered with the same final result.
I have not used Pijul a lot, but I did use Darcs before Git. Some of the merges felt like magic.
Git is fine but it can get messy with complex branching and merging strategies. Patch based systems are intended to improve that.
That's not really true - only in the case where the changes are ambiguous.
I've worked on very hairy Git repos and haven't really felt blocked, even when some crazy merge conflicts happened.
I was surprised that the post numbers appear to be sequential, and that there would be ~30M. Over 15 years of HN, that's ~5K posts/day, and about 200 new posts/hour. Things accrue. I'd estimate that HN is well north of 300K readers/day ... https://news.ycombinator.com/item?id=9219581
edit - it looks like each comment has its own item id. So posts are both new items and comments. That makes more sense to me in terms of scale.
About 4 years for the next 10M from there to https://news.ycombinator.com/item?id=20000000
Approx. 2.5 years for the 10M after that to the current total number of posts and comments we have today.
That's pretty interesting to me.
I wonder if growth will continue, and for how long HN will be able to sustain its excellent signal to noise ratio.
- https://news.ycombinator.com/item?id=11111111
- https://news.ycombinator.com/item?id=12345678
I still have to look up the order of arguments for commands like git rebase, but overall, it makes complete sense.
I can really recommend "Pro Git" by Chacon and Straub. I got a printed version, but it is also available under the CC-NC-SA here: https://git-scm.com/book/en/v2
Specific things I've experienced that could be better in Git:
- Merge commits that aren't actually merges. I work on a game team w/ non-technical people who commit to the repo; periodically it'll happen that they get confused with a merge conflict, try to back everything out, and end up committing a "merge" which actually just completely drops one of the parents. Then some time later someone notices stuff is missing, and it's really annoying to go back and fix. This shouldn't even be possible! (And in Pijul it is not.)
- Inability to use VC while resolving conflicts. In git, being in a conflicted state is a "special" circumstance. Let's say you do some huge merge, and you have a gazillion conflicts. You can't gradually fixing these conflicts, committing as you go, nor can you collaborate with someone else to fix them. This is annoying and unnecessary. In Pijul having conflicts is a first-class state of the system, and you simply add more commits to resolve the conflicts.
- Bad "cherry-picking" supporting in git (e.g., inability to cleanly share bug fixes dev <-> stable). Let's say you make a bug fix on a dev branch which should also be applied to a stable release branch. If it's a single commit, you can cherry-pick; if it's a serious of commits, you can try to cherry-pick them individually, though this gets hairy. But in either case git doesn't actually "track" what happened: the cherry-picked commits are recorded as completely new commits that just happen to have the same content. Later on you're likely to get spurious conflicts. In Pijul (as I understand it, from reading) you just have the same commit included in two channels.
- Lackluster support for binary files. If you check in binaries, the repo gets huge and all operations slow down (and GitHub sends you angry emails threatening to delete your repo). So you use centralized things like git-lfs to get around that, but those are a little janky and not as nice as having everything in one VC. Not sure if Pijul does this better or not, but there's definitely room for improvement here.
I also suspect there are other advantages that would become clear working w/ a patch-based system, but since I haven't yet had the pleasure of doing so, I can't be sure. :-)
I'm old enough to remember when everyone used to use svn, and git was the crazy new upstart -- back then, very similar arguments were made against git as are made against patch-based VC now. Stuff like how distributed version control was too complex, and what did it really buy you if everyone was using a central repo anyway?
It's also interesting to note that distributed VCs had existed for a while, but they didn't break through until there was a really good implementation with a famous initial user (Linus and Linux). And git really took off once the excellent GitHub site was made.
For software to be good, besides strong theoretical foundations, you need to nail all kinds of nuts-and-bolts engineering and design issues. And for software to be succesful, you need to get the social and publicity factors right as well.
Will Pijul manage to fulfil all those? Who knows, I hope so!
And I think inevitably something will come along and replace Git, and probably it will be patch-based.
Everyone here points out that the big advantage to something like pijul is in how it does merges. Git has a set of extremely complicated merge algorithms (the most common up to today has been "recursive merge" which you'll see mentioned a lot in git console output, it's on the way to being replaced with ORT [Ostensibly Recursive's Twin] which is a from-scratch rewrite of the "recursive merge" algorithm with a better understanding of the problem space; there are a couple other algorithms [strategies in git parlance] that have more niche uses). When they work well, they do a brilliant job, but they are extremely complex algorithms and have to do a lot of work to figure out things like "what things got renamed/moved and where between these two branches?". Git mostly doesn't store any of that information at all and generally computes it all from scratch every time it is needed. (You'll notice, again, if you watch console output a lot how often "rename detection" alone often runs, sometimes on the same commits over and over.) But a patch language like pijul encodes things like renames/moves much more directly and doesn't need complex algorithms to guess when they occurred because it wants the user to encode that more directly (as possibly a better "higher level" representation of the user's intent: the user wanted a file renamed/moved and that was an action they took).
In theory then, the merge algorithm of a patch language like pijul (or darcs, its older relative) is simpler because it has a lot more higher level recorded constructs and a lot less "guesswork heuristics" to try to oracle user intent (sometimes long after the fact). (In practice we find such merge algorithms have nearly as much complexity in their own way. The patch algebras of pijul and darcs are both academically related to work done on OTs and CRDTs, all three approaches informing each other, which is in part why darcs was written in Haskell. [pijul is not, just as we have CRDT libraries in most languages today, we've come a long way in our understanding of the complexities of these tools since the early days of darcs.])
Though, I don't think merges on their own are the killer thing that makes an approach like pijul's truly better than git's. Again, git has achieved an incredible "local maximum" of "good enough". As complex as git's recursive and ORT merge strategies are, they are by far "good enough" (and hands down better than many predecessor VCSes) with a lot of smart engineering work behind them at this point. In many "day to day" uses cases you aren't going to notice a huge difference from git's well optimized "dumb" "guess what the user was thinking at the time" merge algorithms and pijul's smart "capture what the user was thinking at the time" approach. Most people don't see the horrors of things like git's rerere cache in their day-to-day git lives (and that's a great thing; I would not recommend learning git rerere if you can avoid it).
The killer feature is cherry picking. Git has git cherry-pick and it "exists" technically, but anyone who is smart has warned you away from ever using it, and especially not in day-to-day workflows. Cherry Picking is the concept of taking just one or two commits from the middle of a branch and applying them to a different branch without taking the other changes. In git this creates entirely new commits with their own very different identities that are entirely unrelated on the DAG (git object graph) in any way to the original commits. It's like a rebase, but much much worse because typically you rebase an entire branch and throwaway the originals, but if you are desperate enough to cherry pick you likely need to keep the originals as well safe in that other branch. In my experience in git this is one of the worst sources of bad merges when those two branches eventually (and often inevitably) merge. A lot of the "dumb" heuristics in git's merge algorithms get really dumb when the same changes were made in both branches.
On the flipside, cherry-picking is a "native" behavior of a patch algebra and how its merges work. Just like in CRDTs that are eventually consistent, a patch is often the "same" no matter what branch it is in (context it has of other patches/time it merges in on other devices in a CRDT perspective). You can generally pull an individual patch between branches with ease whether "beginning, middle, or end" of that branch, and it will just work. (And when you eventually merge branches back, though it is far less "inevitable" than in the git case, it knows the same patch was already applied once and doesn't have additional work to do.)
When I was using darcs heavily, I used a lot fewer branches than in my workflows in git. I knew I could pretty easily cherry pick changes between branches without worrying about where they were in "commit order", so often individual patches felt like entire branches (or tags) and you could often pick and choose specifically what you need at any time. Cherry picking was so common it was just assumed you could do it. (There are cases where you couldn't quite cherry pick what you want, where patch has a more direct dependency on an earlier patch that you didn't expect, but the darcs UI was really good about making it clear that pulling one patch would pull some others it relied on.)
Pijul should be the same, "natively and easily cherry picking", and that can be a killer feature that git doesn't have a lot of ways (today) to do better.
In many ways that "native cherry-picking" felt like a much better way to flow changes between branches. Cherry picking always seems like a key feature that developers need (often because management needs it: "can you get just Feature X into Production without Feature Y because Feature Y isn't ready yet?" but they were developed/integrated in the same branch), if not how they naturally think of changes flowing between branches, and I've seen too many teams already fall to "well git cherry-pick exists so it might help us here" and the ugly merge hell that approach leads to (including horrors like hand-holding the git rerere cache).