Pijul version control system
pijul.org
pijul.org
t's glaring flaw though is that it has historically been really slow as the repository grows in size. They've improved this since but it still can't compete with git and hg on speed. If it were faster though I'd use it in a heartbeat. Pijul looks to be a refinement of the darcs algorithms and thus much faster as a result. I'll definitely be watching them as they progress.
$ darcs get http://arc.liv.ac.uk/repos/darcs/sge &>/dev/null
$ cd sge
$ darcs log --count
4851
$ darcs show files | wc -l
4715
$ darcs show files | tr '\n' '\0' | wc -l --files0-from=- |& tail -n1
1343975 total
Rather than speed, it may be more important that the darcs theory has needed patching up.Least painful - presumably :-) But I have heard that if things don't go well it's much harder to fix up.
>28 October 2015: Discovered Rust, a functional language that is not only type-safe and fast, but also has threads and binds easily to C on Linux and Windows. Looks like a good target.
This is false, as demonstrated by Mercurial.
http://ahal.ca/blog/2014/new-mercurial-workflow/ http://ahal.ca/blog/2015/new-mercurial-workflow-part-2/ http://evolution.experimentalworks.net/doc/
That said, the problems with Mercurial Queues come from (1) a bolted-on command-line interface that doesn't fit in with the other hg commands and (2) poor integration with the rest of Mercurial. Which is why for most use cases evolve is considered the better option these days.
That said, I find the original claim, that users of snapshot systems have to know the internals, a bit idiotic. Subversion users definitely didn't, nor did most CVS users. Git seems to just be a very loud exception. Whether this is by necessity due to Git's UI, or whether it's simply that Git's data structures are more visible and discoverable, is a debate I'm not touching. (I'd note that there's a direct 1:1 correspondence for all core Git and Mercurial internal structures, so it's definitely not inherent in the snapshot model. There are even very nearly 1:1 analogs to older DVCSes, like BitKeeper and Monotone. Git users just tend to know more about the internals of their tooling.)
Mercurial tutorials aren't like that. It is a lot simpler to use so Mercurial tutorials tend to just tell you how to use it and don't mention the internel implementation details.
A comparison to any one of those would do wonders for the landing page.
For everyone else: In Git and Mercurial (and Bazaar, Monotone, Subversion, tla, Fossil, TFS2, and others I'm forgetting), the state is always represented conceptually as a snapshot of the current revision and a pointer to the snapshot of previous revision or revisions, so if you identify that you're at "snapshot 590", and I'm also at "snapshot 590", we're both at the same thing, and both have exactly the same history of how we got there.
In a patch-based system, like Pijul or Darcs, that's not true. The state is instead represented by the set of applied patches. These usually have an implicit ordering, but you and I can theoretically have the same patches applied and have a different ordering. (We also can briefly have different resolutions to merge conflicts that result from that, and Darcs' historical inability to deal efficiently with you and me resolving conflicts differently is a major reason it doesn't scale well.)
That sounds weird, but can have major benefits. For example, I can now sanely cherry-pick a bug fix from you regardless of the history of that bugfix--and, unlike in Git and Mercurial, Darcs/Pijul will know that the patch I cherry-picked from you is the exact same one that you have on your branch. No duplication, Cherry-Pick: messages, or merge collisions will result if I later pull in your whole branch (or vice-versa). You can also do things like have a patch that has all of your security settings, but always simply not push that when you push out to a public server. (N.B., giving this as an example, not best-practice.)
If Pijul works, and can provide Darcs' model in an efficient way, it'd be really awesome. I'm not sure whether it'd be better than Mercurial or Git's model at scale, but it's been really hard to answer that question historically when Darcs itself didn't scale well. Pijul could change that.
I give git training at work with some frequency. I now briefly discuss this sort of patch-based workflow, explicitly so I can tell people it's not how git works, because in my experience, it's what most people naively expect out of a version control system. Same problem when people try to work out workflows for git. I hypothesize that the reason for the profusion of git workflows is precisely that we really want a patch-based system. Or, in other words, it's not really weird; snapshot-based systems may be weird.
It never comes up in SVN or similarly klunky systems, because they're too incidentally complex to notice the underlying essential mismatch. Git and friends, to their credit, made the incidental complexity go away, but I think the essential complexity of their approach is still higher than it ought to be. I wish pijul all the best, because there is room for improvement here.
And may I suggest to the pijul team that they really, really want a Pijul-hub as soon as they think they're even remotely ready.
I give git training at work with some frequency. I now briefly discuss this
sort of patch-based workflow, explicitly so I can tell people it's not how
git works, because in my experience, it's what most people naively expect
out of a version control system. Same problem when people try to work out
workflows for git. I hypothesize that the reason for the profusion of git
workflows is precisely that we really want a patch-based system. Or, in
other words, it's not really weird; snapshot-based systems may be weird.
I actually agree. I advocated for DVCS-based workflows for a long time, but I think it speaks volumes that I "got" Darcs within about 20 minutes of first seeing it and playing with it, but I know for a fact that, somewhere in the Freenode archives, you can find me saying "The frak is Mercurial? The frak is Git? The frak is this?" as I tried to grok what on Earth they were doing--and this after having taught myself tla!That said, just because something is intuitive doesn't necessarily mean that something is best engineering discipline. I want a Darcs-like workflow to be the dominant one specifically because it's intuitive, but I'm very open to the fact that it might be a really, really crappy way to build a sane (forget performant) large code base.
Or it may be exactly what we've always wanted.
The simple fact is that I have no idea. This industry, for all the "science" in CS, has got to be the least results-based discipline I'm personally aware of. Pijul will at least permit anecdotal tests of how well patch-based systems scale as a workflow if it takes off, but I'm not holding my breath on someone doing a genuine A/B productivity/trade-off study.
The rationale behind SHA-1 being obsolete is that it will be used for password protection. The back end stores Hash(password+salt) and uses it to grant access. Finding any password_prime which causes a collision with password on Hash(p+salt) will brake the security of that scheme.
In this case, instead, the problem from the point of view of the attacker is to find repository_prime such that...
1. Is a syntactically correct program...
2. It is syntactically and semantically close enough to repository_orig that it will not be discovered by simple inspection.
3. The delta between repository_prime and repostory_orig causes a useful side effect in the program behavior (from the point of view of the attacker).
4. The delta between repository_prime and repostory_orig does not introduce other significant and unintended side effects that will trigger investigation beyond simple inspection from some legitimate maintainer.
I will say that this is a much higher bar to cross than merely "no collision shall be found". And if you are concerned that SHA-1 is not enough to protect against such attack, you probably have to consider that SHA-2 may not be enough either... You probably need to use half a dozen of cryptographically strong hash functions, preferably based on different principles, so that we ensure that at no possible delta can simultaneously fulfill all the 4 points above for every hash function.
And if you have reached that level of paranoia, you need to consider that the adversary will simply not bother and chose a different attack vector, like exploiting the bureaucracy of your commit process...
Just say no to SHA-1.
Indeed, this may be a higher bar to cross, but higher bars are usually being crossed. For years we known that RC4 had biases, but it was considered okay until we discovered that it was not okay after all. MD5 collisions led to rogue certificates.
It's not correct that SHA-2 is not enough to fix the issue. While they share some design structure, SHA-2 is not broken, while SHA-1 is. Also, the proposal for multiple hash functions isn't particularly good: collision resistance of many cascaded hash functions is not much better than the maximal resistance of one hash function used in the cascade.
It doesn't make sense to use a broken cryptographic hash function in a new project that does need cryptography. There are faster modern cryptographic hashes than SHA-1. And it doesn't make sense to use a broken cryptographic hash function in a new project that doesn't need cryptography. There are faster modern non-cryptographic hashes than SHA-1.
To be fair, I don't know why we even have a discussion about this if cryptographers say that SHA-1 shouldn't be used, and you're trying to defend this choice for exactly what reason?
this kind of bare statement benefits greatly from a citation
> collision resistance of many cascaded hash functions is not much better than the maximal resistance of one hash function used in the cascade
I don't think GP is suggesting including results of hashes in the contents to be hashed by subsequent hash functions, but rather calculating MD5(contents), SHA1(contents), SHA2(contents), FOO4(contents) etc, and then concatenating all the hashes together. Or am I misunderstanding and this kind of scheme is exactly the "cascading" that you're talking about ?
Seriously?
> I don't think GP is suggesting including results of hashes in the contents to be hashed by subsequent hash functions, but rather calculating MD5(contents), SHA1(contents), SHA2(contents), FOO4(contents) etc, and then concatenating all the hashes together. Or am I misunderstanding and this kind of scheme is exactly the "cascading" that you're talking about ?
Yes, this is cascading. See this paper: https://www.iacr.org/archive/crypto2004/31520306/multicollis...
Think about what happens if you were to fork a copy of something on github and then rebase to edit one of the commits. If I pulled from your copy and the pulled from the official repo I'd get a merge (and probably some conflicts) because git identifies commits by their sha and uses that to decide what needs to be merged.
But if you managed to change one of your commits without changing the sha, then I'd have a much harder time detecting it, since when I pulled from another repo later there would be no merge (even though things were actually different).
There are modern cryptographic hash functions, such as BLAKE2, that are faster than it, and I'm not even talking about non-cryptographic functions, which are many times faster.
The integrity of deltas doesn't matter much because if you have a process for maliciously inserting a falsified patch, you can generally also simply insert a genuine patch (where it's trivial to get even the most complicated cryptographic checksum right) that will be combined with the rest (because the state of the repository is the combination of the patches in it, and whether you replace one or add a different one is largely irrelevant in this context).
Examples of things Git doesn't do or do well:
* Shared mutable history (this is where Pijul/Darcs come into play).
* Scaling to large repository sizes [1], especially for monorepos.
* More intelligent merging (based on text history and/or semantics). Extracting the full text history of a file can be very expensive in Git (for example, see the performance problems of git blame).
* Long-lived branches (due to lack of meta information in Git's storage model, this is often a problem; for example, compare the conceptual complexity of bzr join vs. git subtree).
* Annotated history; Git does not have branches in the traditional sense, and this means that history is largely a soup of anonymous objects. Compare this with Fossil's model of propagating tags [2].
[1] See these two talks from Git Merge 2015 for example: http://git-merge.com/videos/git-at-google-dave-borowitz.html and http://git-merge.com/videos/scaling-git-at-twitter-wilhelm-b...
Also, it doesn't necessarily have to replace your existing VCS. Maybe a project like Pijul helps to develop the theory into practice that becomes an amazing new (backward compatible) backend for some future Git N+1 version, for instance.
Pijul stands to do the same thing for the patch-based model. Darcs has existed for a long time at this point, but it's not popular, because it's slow. Darcs had the free part from the beginning, but was slow; Pijul is aiming to finally be free and fast enough to allow adoption. If those things both happen, I'd expect a GitHub workalike to show up as well.
Being the first with an idea doesn't count for much if you can't execute on it in a way that others like. Darcs clearly has failed at that; Pijul is aiming to do better. We'll have to wait and see if it does.
Darcs (which predates git, by the way) represents a repository as a loosely ordered set of patches (just the changes, not the full state at any given time). Darcs works to know how to rearrange patches in time to create new states. The difference/benefit to this approach is best illustrated with cherry-picking: in Darcs it is very easy to state that I want the patches my fellow developer created for modules X and Y, but I don't think he's done with module Z yet so I'm going to skip those. The patches can be interleaved and recorded in any order and Darcs will do all the hard work of figuring out how to grab just the ones I want, and do so without changing the physical identity of those patches (they are still the same patches my coworker developed). When this works, and it works 98% of the time without an issue, it is a wonderful magic. (Contrast this to git cherry-pick which builds new commits and is likely to have future merge issues with the commits it was built from, even those commits are somewhat "the same".)
Pijul is an attempt at getting some of the same magic of a patch-oriented approach (and smarter patch merging) in a post-git world.
Also, it's definitely early to talk about switching to Pijul.
0:https://github.com/uutils/coreutils
1:https://github.com/mkottman/lua-git/blob/d64218a227fe26b1fc9...
I say they are novice because they are maintaining two parallel codebases, which is a mistake no experienced developer would make.
By the Pijul team forcing themselves to have two different implementations in lockstep, they will probably do a very good job avoiding that problem.
In order to force yourself to stick to a reliable standard repository format, maintain two codebases at once.
What?
How about just publish a standard and abide by it.
The approach is quite costly but if you care more about quality and creating an excellent standard than getting a product out of the door quickly, it's certainly worth it.
Also, from their website:
Why rewrite it in Scala?
The Scala implementation is not competing with the OCaml implementation, we do plan to maintain both in the future. Having several implementations of the core forced us to really think the formats through, and will allow others to write tools in the languages the like best, for their favorite platform (try to get and compile OCaml Pijul on Windows!).
darcs.reviewed$ darcs log --match 'author florent\.becker' --count
200
See also their FAQ.Is there something about the software that you feel could be improved, or are you just offended that "novices" write software as a hobby at all?