Is Git more than a version control system? Reimplementing CouchDB with Git+Bash...
ordecon.com
ordecon.com
1. I learned from Linus that git is a decentralized "content" database that gives identity to different versions of the same content and allows them to be compared and merged. (as opposed to a traditional VCS which is more of a "delta" database)
2. I learned from Damien that CouchDB is a decentralized "document" database that gives identity to different versions of the same document and allows them to be compared and merged.
I've been wondering just what the difference really is.
In addition, Git makes individual merges in a document, and when there's a conflict, it resorts to human intervention. Couchdb does not make merges inside a document, and will not do conflict resolution. Instead, using a brief set of rules, it picks one version of the document over another.
Large homogeneous systems that only provide one functionality are hard if not impossible to extend to make them do just what you want. Unless you want what they do in the first place, which makes the whole customization thing pointless ;-)
But that is self evident, or at least should be.
It's true that the road to hell is paved with good intentions :-)
I think a lot of us have had this one cooking in our minds for a while and its great to see that somebody got it Done.
Bravo.
First distributed VCS I heard of was darcs. I have not used it either, but from what I've read about the two, git has some real advantages in speed and robust-ness.
It's not the same thing. Using SVN as a document database is equivalent to using the filesystem as a document database. The only thing extra you get with SVN is a revision history.
Git actually is a content database. Its version control capabilities are built on top of that.
When you do a checkout on a traditional VCS you're telling the VCS to apply patch set x to the file system. When you do a git checkout you're telling git to load content x from the database into the filesystem.
As soon as I find some time I'll do a full write up for anyone that's interested.
Last night I was thinking about how much I hate existing file managers. How the metaphor was great when you had 2 meg hard drives and a few folders and maybe a couple hundred files tops, but how now it falls apart. There was a project called lifestreams at Yale ( http://cs-www.cs.yale.edu/homes/freeman/lifestreams.html ) that had some interesting ideas about allowing you to see your documents as a versioned timeline. Using a dcvs type system as a backend would get you a lot of that plus more, you could explore different ideas for files on different branches. I do however want more...
I want usable metadata like the BeFS had. Where any arbitrary metadata key/value paris could be attached to a 'file', For example, contacts in the BeFS were basically empty files with metadata attached. That metadata included name, address etc. Any application could uses the data, augment it etc. The file manager ( tracker ) could query the data to create live 'searches' that looked just like folders. You could add new types and what not. Very powerful when the filesystem is a database. BeOS ( and now haiku ) kept the traditional file manager around as well. Others have approached this idea, but haven't really gone after it.
What I thought of that I really wanted to see was something like a document-centric database like couchdb that is backed with dcvs like features that operates just like a regular filesystem, i mount it, i can drag and drop content into it, save into it, make it look like a 'normal' fs to existing applications, but all that path info etc is just saved searches on metadata that group information together and you can tag those searches with metadata themselves.
On [one](https://www.getdropbox.com/tour#3) of the pages from their tour, it shows how if you save a 10 MB PSD to Dropbox twice, it shows you a table with a row for the original file and another row for '400 extra bytes' or whatever.
> Dropbox is also smart with how it tracks changes to files. Every time you make a change, Dropbox only transfers the piece of the file that changed (also known as block-level or delta sync), making it easy to work with big files like Photoshop or Powerpoint documents.
Tracking changes in binary files (which cannot be merged in any reasonably generic fashion) is a fundamentally different issue than tracking changes in text, particularly source code. Git is designed to do the latter. While you can use it to track changes in binaries, merging doesn't make sense anymore, and hashing / scanning big binary files for changes is significantly slower. (A bunch of images generally won't matter, but I wouldn't use it to track, say, video, or large database dumps.)
It's just amazing what a good VCS can let you do.
It's worth picking any and just doing it, though.
* One style difference between the two. See Graydon Hoare's comment here (http://www.mail-archive.com/monotone-devel@nongnu.org/msg080...).
Don't have time to set it up at the moment, but I'd been wondering exactly the same thing.
I am using bazaar to do daily backups at some webapp that stores serialized data in files. I was thinking a little about syncing it between servers with bazaar if needed but you seem to go a few steps further. :)
I've been using rsync to keep certain directories of my various computers synchronized, but I think I'll go one step further soon enough: for my Linux computers, I'll keep some parts of my home directory in Git in order to maintain an identical user interface / configuration on those machines.