Darcs - Another open source version control system
darcs.net
darcs.net
You can imagine that I write a new filesystem for Linux, but I wait two years and the FS interface has changed drastically, the old patch is meaningless. If the patch is sufficiently independent textually, then Darcs lets you reorder it anywhere you see fit, even if you know that doesn't make sense. You have detailed knowledge of the system, Darcs doesn't.
By contrast, Git/SVN/Hg folk think about snapshots of development. So it's a different paradigm.
But the snapshot paradigm is almost always the most useful paradigm, since every snapshot in history (unless you rebase) is going to be one that was actually vetted by a real programmer. Reordering patches throws away the valued actual snapshots that you were working with in favor of automated reconstructions of what a snapshot might have looked like if you had been programming in that order.
Of course, you can always do the same dangerous things with Git and rebasing.
Except, cherry pick will only work for "in air" commits, and rebase requires you to review all the commits.
darcs will happily pull a set of related commits knowing the relations between them. You may have detailed knowledge of the system so in some cases you may do something better than what darcs does, but that does not mean you should be resolving merges and commit dependencies by hand all of the time.
Honestly, I have not used darcs in years, but that is the "right" thing I expect from a VCS, and it has no downsides per se[0].
[0]they may be in darcs implementation, or possibly in the theory, but definitely not in the feature
Honestly, I expect the differences between Darcs and Git in terms of rebase / interactive rebase / cherry pick / merge to be a matter of implementation. Both systems are fed the same input, both track enough of it, so both could theoretically provide the same output. (I think Darcs is a little better at it, but every merge should be done with a watchful eye in either system.)
But I get the feeling that the patch-centric model is encouraging users to play fast and loose with patch order. If you wantonly reorder patches, you might get a sequence of unusable code bases culminating in the current working version of your software. It's easy to move a patch that uses a feature to a point before the feature is introduced.
Basically, the claim that "Darcs automatically groups changes which depend on each other" is one that I can't take seriously. How can it know which changes depend on each other, if (for example) they're in different files? By comparison, I hear from the Git folk advice along the lines of, "Be careful with rebase, because it is a destructive operation".
And because Git is snapshot-centric, if the app crashes in testing you can look at the SHA-1 used for the build to check out the exact tree used for the build, even years later, if necessary. No tagging required.
Well, if you use git yes, since there is no check for this, in the normal darcs workflow either the patches are independent or you can't reorder them. So, "fix About" can be moved before "new login screen", but "using data in the new login screen" cannot.
> How can it know which changes depend on each other, if (for example) they're in different files?
Because _you_ told it so, by making them part of a single commit, or by creating a commit which depends on lines in both files affected by previous commits in those files, or by tagging the repo. Surely it's not a crazy idea to group related changes across files in a single commit :)
E.g.
# distinct commits
A1 "added foo() in foo.c"
B1 "added foo() in foo.h"
# depends on b1, if you cherry pick this you get B1, B2
B2 "added ifdef in foo.h"
# depends on all three, you get A1, B1, B2, C1
C1 "renaming foo() to bar()"
# depends on all four, you get A1, B1, B2, C1, C2
C2 "renaming foo.(c,h) to bar(.c,.h)"
Notice that yes, you can miss a commit, or include too much. In the worst case you end up with the same "I'll review the commits" workflow you use in git cherry-pick/merge/rebase.FWIW, I am not advocating darcs over git, just trying to shed some light.
# distinct commits
A "added foo() in foo.h and foo.c"
B "used foo() in bar.c"
Clearly, B depends on A. There's no textual dependency, however. How do I tell Darcs about that?the only way to reorder these in git is with rebase, but that is flagged as a dangerous tool, and with good reason.
if darcs can find a way to do this it would really be worth switching to, but if not then it would fall short in it's claim of safely being able to reorder patches. doing so behind the scene automatically would really worry me about code integrity.
It is quite common to have people come to #darcs asking how to reorder patches. The answer to that is "why would you want to do that?" To darcs, the order in which the patches are stored in the repository is a transparent implementation detail, and there's no command to change that. Darcs users are not constantly reordering patches, they're just enjoying a bit more freedom on pull. In general, if you don't use commands marked as unsafe such as unpull, each repository is honest about the order in which it pulled patches: darcs will simulate patch reordering to compute merge results, but won't change the order of patches in the repository.
There's one "patch theoretical" thing which git does better than darcs (and a lot of practical things), which is that it tracks more explicitely which stuff was pulled from where when. That information is neglected by darcs, which cares more about when stuff was originally written, and it should be tracked somewhere. Still, this seems easier to add to darcs than adding darcs' freedom of workflow to git, for which, as far as I can tell, you'd have to embrace the patch-centric view of darcs.
Consider for example, having two branches: "stable", "master" which have already diverged. Fixes go into "stable" and are regularly merged into "master".
Now imagine someone writes a fix F, which must be applied immediately to both "stable" and "master" and can't wait for the regular merge effort.
So with git, you just cherry-pick F into "stable" and "master".
Now, someone who uses "stable" discovers a big flaw in F, and reverts it. They assume their work will be merged into "master" as usual.
Then, a merge from "stable" to "master" happens, and the merge sees that the "stable" branch has 0 changes w.r.t F (commit+revert), whereas it has the F change in "master". Git's merge algorithm will consider this a trivial case, and do the wrong thing. The merge result will have the broken F application in "master", and there will be no conflict or any indication of a problem! This has been a huge problem for us.
So the result of this is that git forces you to either "merge" upstream, or "cherry-pick", but you really should not mix these two modes.
If you only use "merge", then the guy who merges must intimately understand everyone's changes in all cases of conflict. If he doesn't, others must help him resolve the conflicts on his working directory. Then, the merge result is one monolithic commit that makes understanding changes very difficult.
If you only use "cherry-pick", then you lose tracking of which changes have already been merged upstream, and which haven't. You start relying on the grep of "git log", which is unreliable, and you have no reliable tools to automate this. "git cherry" is unreliable and has plenty of false negatives/positives. Then there are of course problems with patch dependency, where you have to figure it all out on your own.
In short, the snapshot model is very problematic here. None of these problems would occur in the darcs model. Unfortunately, darcs does not scale. Whether that's inherent or not, I don't know. But if it did, I know I'd definitely prefer to use darcs in a large collaborative effort than git.
Surely this is something that the Linux project runs into constantly? I wonder if they have a solution?
In our case, we had the "bleeding edge" branch (master) and a maintenance branch. Fixes generally go into the maintenance branch and then are merged upwards. But we allowed "emergency patches" to be cherry-picked upwards (or downwards) too. The combination is the problem.
If we had stuck to just cherry-picks or just merges, we would not have this problem. But as I also explained, sticking to one of the approaches has severe drawbacks too.
Perhaps, but the same could be said for git: the original commit is also meaningless because so much has changed since. I don't see that as a compelling argument.
Also, darcs lets you explicitly set dependencies on patches for the cases when it can't automatically detect them. The key philosophical difference is that darcs doesn't use time as an implicit dependency.
We use darcs on greenfelt.net. We push patches to our development repo which is where the other developers pull from. When we are happy with the state of the feature we push the patches (manually, usually from 2 to 10 patches) to our test site for a quick sanity check and then push that live.
In practice this means that we have half finished features/ideas on our development repo that never get pushed (the oldest patch is from 2009 and is still perfectly relevant, I might add). It also means that bug fixes can be fast-tracked through (directly from development repo to the production server), when appropriate. This is all handled by darcs in a very straightforward manner (of all the DVCSes I've used, darcs has the absolute simplest/best UI--git has been copying it a lot lately, see "git add -p").
The other thing that I found that we do constantly with darcs that isn't really doable in git is "darcs unpull". If someone pushes a patch to the development repo that is flawed you can just unpull it and tell them to go fix it. We use darcs-notify so that the team is emailed when someone unpulls a patch from the main repo because it requires everyone on the team to unpull the patch. Of course, most of the time a patch gets unpulled it happens before anyone else has pulled it. We've done this even on fast-moving projects and it's less scary than it sounds, and allows you to keep the coding standards high.
Along those lines, I've noticed that the darcs push and pull UI encourages code review by asking you explicitly which patches you want to push/pull. At the prompt you press 'v' and it prints the patch up so you can see what it actually does. It's a seemingly small thing but I've watched other developers change their coding/patch styles dramatically after using darcs and watching the patches go by. I tend to use "darcs push" as a last minute quick code review to make sure there's nothing overtly stupid in my patched. I can't tell you the number of times I've control-C-ed a darcs push and amended to make the patches correct.
One thing I've noticed in using both darcs and git extensively is that once you've really embraced each one they encourage different kinds of commits. Darcs really likes light, independent patches (since conflicts are not darcs's strong suit) and git likes more atomic feature patches (since a commit that touches 15 files doesn't have any drawbacks).
Unlike in the relational database world, there is no accepted standard for interfacing with a versioned repository. It remains to be seen whether Darcs, like PostgreSQL, will eventually gather enough steam to build a substantial user base.
What killed darcs, besides performance issues with darcs 1, was "one repo one branch" mode, and tedious way of maintaining long lived forks, which wasn't that convenient as what git provides.
However, I'm glad that Darcs is still actively developed and used, since it has very nice theory and model behind it, and it still has some very useful concept which are quite unique for it.
Wasn't BitKeeper in use by the Linux kernel for several years before Darcs came out?
Therefore, proprietary tools don't count.
I like Bazaar and used Hg, but this is tool used to collaborate with others (unlike say an editor) so picking a tool that others know and use is important in a team. Darcs has an interesting approach, sure, but that is not enough for it to win mind-share, and I wouldn't spend time looking at bit because I will have a very hard time getting others to do the same.
Also the excitement and euphoria about learning and experimenting with distributed VCs is passed. I feel, for most they have stopped having the "wow" factor and became just a tool in the same category as patch, diff, less and tar.
For what it's worth, I was in the Monotone [1] camp against Darcs back in the day... :)
I bet that if any version control system is going to replace git, it'll look at least as much like Darcs as anything else.
I'm no expert by a long shot, but to me it appears that if we're going to try to finally make version control user-friendly, something inspired by Darcs' patch algebra may well be at the base of it. While Darcs is impractical for a whole set of use cases, it's very practical for a whole other set of use cases.
Hg has also quite some mindshare. I consider it the only DVCS, which still seriously competes with Git. Apart from that, there are niche groups: Haskell with Darcs, Ubuntu with Bazaar.
That's why I said "user friendliness". Even on a superficial level, Git is enormously lacking in this aspect. More fundamentally, the way moving changes around between branches works in Darcs is significantly simpler and "hey, it read my mind!" than in Git, where you need to understand the underlying object model and the tools very well to do this well. Sure, many other things are a disaster in Darcs, but I'm not saying Darcs will replace Git. I'm saying something inspired by it well.
People said this about Subversion not so long ago.
Hence, it's worth taking a little time to actually understand those features, even if you never use darcs. Plus the mental exercise and learning something new is valuable in and of itself.
I think it is good to periodically evaluate the tools used. And initially there was quite a bit of excitement around DVCS and everyone was playing with them (including me). Me and most have settled on one tool (git). And actually this was not my preferred tool of choice, I like Bazaar, but 2 things happened: 1) Bazaar has not won the popular vote (so chances are on a new team I would have to learn git anyway) and 2) DVCS as a technology has stopped being new and exciting.
For example at some point OO features and garbage collection was exciting and new. Now they are just another tool in the shed available for use and as I said, this area of tech for me has lost the "wow" factor.
I also realized I have a limited bandwidth both to learn and to remember things. I would rather devote this bandwidth to learn new languages or other technologies (databases) rather then learn more DVC systems.
Their way of doing things seems to require a 1000 times less code than current mainstream techniques. They can do a whole language stack in less than 2000 lines of self-implementing code for instance (not super-optimized, but fast enough). Compare with GCC or GHC. Even Lua is bigger.
Speaking of terseness, reminds me of Prolog. I remember thinking how beautiful the solutions are but how broken my brain is as I couldn't come up with them on my own. I could see the final result that someone else produced and understand it, but struggle for hours and hours myself.
I think at some point levels of meta and abstractions can go that high that only very few smart individuals can effectively create products, others more or less hit a wall.
When the teacher is a genius, the student needs not be bright.
Maybe for open source projects it makes sense to stick with the most popular tool but for private teams if you can see an advantage to another tool and it looks like it will be around for a while then there is no reason not to use it.
For a long time, darcs didn't have much of a developer community, possibly due to the choice of implementation language (Haskell). It also had a corner case that could cause a merge to take several days to complete (I suppose they solved it by now; it was rare, but if you hit it, you couldn't really keep using darcs).
And simmetrically, a few years later, david roundy (the original author of darcs) started working on iolaus[1] which is in fact a darcs-ish implementation on top of git.
[1] https://github.com/droundy/iolaus now abandoned I believe
On the other hand if you want to move from darcs to git , use the following - https://github.com/purcell/darcs-to-git
It's not ready for bi-directional incremental bridging, but for one shot conversions, it should be just fine.
The Darcs bridge uses the Darcs library to do things and understands Darcs better.
We tried using Darcs at work in 2006, but the early exponential merge issues made it impossible to use. (One time, we let Darcs try to merge "some changes" over the weekend, but it never completed.) Unfortunately we chose Subversion instead.
I loved it for small repos and my private work, but it just wasn't ready for larger code bases. Perhaps it is now, but in the meantime git got much, much better and I'm not looking to switch back.
Darcs got too popular too early and was hurt for it (although then again, its relative popularity at the time was also a good way for us to ferret out issues). We've now shrunk back to a more reasonable size (which makes it a bit harder to keep going, but takes off some of the pressure too, double-edged)
I think darcs suffered because it set high expectations. I remember being fascinated with the "theory of patches". The web site looked very promising. And it was written in Haskell, so it had to be correct, right? :-)
I liked the simplicity of the command-line interface of darcs. I just never looked around for 1M+ LOC projects using it on a daily basis. I should have been more careful.
I still use darcs for a number of small things and I'm happy with it. I'm happier with Git, mostly because of its practical approach and the freedom it offers ("text is just text" and "you own your history, do whatever you want to it").
Still, I would like to use git as my svn client, but it's just not a suitable replacement without externals handled nicely.
As an aside, I changed jobs to another place using SVN, and when they started moving to Git I changed jobs again. Still using SVN! :-)
It's made a lot of progress over the years, and works great for my needs, but we do still have serious bugs and performance problems we need to sort through; and yes still lacks a lot of critical tooling/infrastructure around it.
We love what we do, and think that we have something new to bring to the table (among other things, making it easy to pinpoint commits that you want to pull, delete, etc; DVCS'es may all do some cherry picking ala git add -p, but Darcs makes it possible to use it everywhere), but it may take us a very long time to get to stage where we can responsibly talk about it.
There lot's of work to do, lots. If you're looking into getting to some Haskell hacking, consider Darcs as a good side project to get into. We can use you.
I would like to see hard data and experience reports of where Darcs worked and didn't work for people, instead (so I'm loving this thread.) Here are a couple of repos I've worked with:
- hledger: 2477 patches, 251 files, 66M, 5 years of activity
- Zwiki 0.x: 1897 patches, 295 files, 12M, 5 years
- Darcs: 10171 patches, 731 files, 51M, 10 years
- GHC before it moved: 23413 patches, 1353 files, 162M, 5 years (+ 10 years of converted history)
Here's the darcsstat script that reports these numbers: https://gist.github.com/2653180
This is why darcs currently doesn’t always works so well for keeping a “local branch” of some source. The local deviations are likely to cause exponential conflicts after time. Local deviations must either be isolated in some way (kept in separate files) so they never conflict, or changes from upstream needs to be merged in “by hand” and recorded as local patches that doesn’t conflict with the local changes.
Can anyone help me reconcile why I would give darcs a try over git if I can't do local branching?
First, you can do "local branching" by cloning your repository. On a local filesystem, `darcs get` is fast (uses hard links). And instead of `darcs branch`, you use `cd` (personally, I don't see how `cd` is worse, do you have some pointers about that?).
Second, the Darcs way of managing patches gets you "spontaneous branches". http://wiki.darcs.net/SpontaneousBranches The main advantage lies in that you can select the patches you push to other repositories. Say you work on some feature. Suddenly you need to fix a bug. The way to do it is, fix the bug right away then send the relevant patches (easy if you name your patches sensibly). Your unfinished feature simply won't get send (unless of course there's an actual conflict).
I suspect there was less pressure for local branching because a lot of things you tend to want to do with branching, Darcs sort of trivially does with its easy cherry picking. But there are cases where cherry picking won't be enough and you want to have the separation, and it those cases, I can see where people would want to share working directories.
As one Darcs hacker among several (meaning I want it to be clear I'm not speaking for the rest of the team), the most important thing about Darcs is the friendly UI [and the friendly UI is made possible by the patch theory stuff]. And so my concern in trying to do Darcs branches is to do them in a way that keeps the UI very simple to learn and explore. We'll get there. Slowly.
The build for the project had passed on most platforms, but on one platform in particular (the one I was interested in), the build agent started failing on a particular day a month before. Being able to determine which commit a build started failing at has a great deal of real engineering value, but I couldn't figure out a way to do that with Darcs. The patches are ordered in the repository according to Darcs' own preference, and they're dated according to when they were committed locally, not when they were pushed.
I discussed this point with someone else and they said, to paraphrase, "usually, I know which of my patches has caused a problem, so this is a non-issue." It wasn't a non-issue in this case, of course, because I didn't know which of his patches had caused the issue, and no one else had cared to figure out why this third-tier platform wasn't building.
Is there any way at all to characterize what exactly a build machine is building at a particular time with Darcs? With perforce and subversion, there's a commit number. Git has a commit SHA hash. What can you do in a build agent to log exactly what is and isn't being built with Darcs?
We've made a lot of progress. You'll see people saying that their performance problems have gone away. But we have a long way to go.
Two things to fix: downloading repos is slow (latency, too many small files); and conflict merging can be slow under a few cases. The first one we have a fix for from a summer of code project, but needs more work. The second issue is a very deep, very serious, and will take us a seriously long time to think through.
I think darcs approach is a little different.